Skip to content

fix(composer): read a blank-padded composer as empty, not as unsent text - #36

Merged
huynhtandat223 merged 3 commits into
mainfrom
fm/fm-pair-delivery-false-negative-fix
Aug 10, 2026
Merged

huynhtandat223 merged 3 commits into
mainfrom
fm/fm-pair-delivery-false-negative-fix

Conversation

@huynhtandat223

Copy link
Copy Markdown
Owner

What broke

fm-pair-compose.sh releases both pair barriers through verified fm-send submission. The Claude driver visibly received PAIR READY, but fm-send exited 1 with delivery unconfirmed; verdict=pending. The composer treated that correctly under its contract and aborted before releasing the navigator, so ready was never published and psak-dms-kernel-canonicalization stalled before its plan gate with no navigator coverage.

Trigger, masking condition, symptom

  • Initiating trigger. Claude 2.1.226 pads its empty composer row with U+00A0 after the ❯ glyph. Verified bytes, byte-identical through both the Herdr ANSI read and a plain tmux capture:

    033 [ 0 m 033 [ 3 8 ; 2 ; 1 5 3 ; 1 5 3 ; 1 5 3 m 342 235 257 302 240 033 [ 0 m
    

    The blank is drawn at luminance 153, above the de-emphasis threshold, so fm_composer_strip_ghost correctly keeps it as real text. Every trim in the shared classifier and its callers is ASCII-only, so a genuinely empty Claude composer classified pending.

  • Masking condition. The composer verdict only decides delivery where no stronger signal exists. fm_backend_herdr_send_text_submit confirms from native agent state whenever the target is legibly idle before Enter, and only falls back to reading the composer when the target was already mid-turn. Every ordinary steer took the immune path. The paired-review barrier release takes the other one, because both roles sit in a blocking wait for the release file when it is sent.

  • Visible symptom. verdict=pending on a message that had already landed, and a pair aborted between its two releases.

The post-submit composer is the same padded-empty row plus a dim Press up to edit queued messages placeholder, which ghost stripping already removes correctly. One fault explains the pre-send read, the post-send read, and the retry exhaustion.

Proven path vs failing path

A known-good send and the failing send differ at exactly one branch: the pre-Enter agent-state baseline. idle goes to fm_backend_herdr_wait_for_working (never reads the composer, always worked). working goes to fm_backend_herdr_composer_state (reads the padded row, always failed). Pre-fix, an empty composer and a composer holding real text were indistinguishable - both pending.

The fix

Fold non-ASCII Unicode blanks to ASCII space in fm_composer_classify_content - the single fleet-wide owner the tmux, herdr, cmux and orca adapters all delegate to.

This can only move content that is entirely blank from pending to empty. Content carrying any visible character keeps its verdict, so unproven delivery still stops the pair and the away-mode injector still refuses to type over unsent input. Keying on the Unicode space-separator category rather than U+00A0 alone keeps this from recurring when a release picks a different blank.

No change to the pair protocol, and no weakening of delivery verification.

Per-harness evidence

Every installed harness measured 2026-08-10 against its own live binary, idle and holding typed-but-unsent text. Captures kept byte-exact under tests/fixtures/composer-captures/ (paths and account names replaced by same-length placeholders so column geometry is untouched) and replayed through the real Herdr composer reader. The pre-fix column is the same capture through the classifier as it stood before the fold:

harness version idle (pre-fix) idle (now) typed pads a Unicode blank
claude 2.1.226 pending empty pending yes, U+00A0
codex codex-cli 0.145.0 empty empty pending no
pi 0.84.1 empty empty pending no
opencode 1.17.20 unknown unknown unknown no
agy 1.1.11 unknown unknown unknown no

Exactly one verdict moves, and it is the fault this fix was written for. Claude is the only one of the five that pads its composer at all, so the other four are evidence that the fold changed nothing for them - not evidence that it repaired them. I am not claiming the fix "works for" a harness it never touched.

Live end-to-end tmux runs agreed with the table for claude, pi and codex. Codex required dismissing an update menu and a hooks-trust prompt first; I moved the selection off "Update now (runs npm install -g)" and chose "Continue without trusting" rather than confirming either default.

Against real claude 2.1.226 on herdr 0.7.3, through the production adapter, both directions:

### DIRECTION 1 - accepted while the worker is mid-turn MUST confirm
pre-send: agent_status=working composer_state=empty
send_text_submit VERDICT=empty
### DIRECTION 2 - a genuine swallow MUST still refuse
typed but NOT submitted: composer_state=pending

Direction 1's probe was observably delivered - it arrived in the receiving agent's turn - while the pre-fix adapter called the same send unconfirmed.

Unsupported seams, recorded not papered over

Two installed harnesses have composers this fleet cannot read on either backend. Neither is affected by the fold, and both refuse as unknown, which denies delivery confirmation and away-mode injection alike:

  • opencode 1.17.20 draws a composer row with a LEFT-only ┃ edge and a ╹▀▀▀ foot, matching neither the bordered shape (same glyph at both ends) nor a bare agent prompt glyph.
  • agy 1.1.11 draws a bare >, byte-identical to a dead shell prompt. The agy rule that would admit it on native identity requires that row to sit below all separator activity, and real agy draws a rule underneath its prompt - so the synthetic fixture that rule was written against does not match the shipping shape.

Reading either composer is a separate change with its own safety argument - admitting a bare > is exactly the dead-shell hazard the shared rule exists to prevent - and is deliberately not attempted here.

agy was also missing from the live guard's harness loop entirely, so the agy-specific prompt rule had never been exercised against the real binary. It is now in the loop.

A guard that ran agents in the tree

While developing the live guard I launched harnesses in the firstmate worktree to get past directory-trust prompts. During that window a real harness modified a tracked file (custom-skills/policy/SKILL.md) with no prompt ever submitted to it. The change was reverted and never reached a commit.

A guard that runs live agents must not run them in the tree it is testing. Launches now default to a scratch directory; FM_COMPOSER_BLANK_DRIFT_DIR takes an already-trusted disposable directory for fuller coverage and refuses this checkout. The honest cost is coverage: from a scratch directory only pi draws a composer on this machine, because the others gate on directory trust first. That is why the per-harness table above rests on captured real bytes replayed through the real reader, not on the guard alone.

An earlier draft of the guard also passed without the fix - it was reading Claude's blank startup screen, not a composer. It now records the untouched verdict and asserts it only after typed probe text becomes visible in that pane.

Coverage

  • tests/fm-composer-lib.test.sh - the fold, every declared blank, the dead-shell refusal, and the divergence assertion (padded-empty is empty while padded-real-text stays pending) so the case cannot go vacuous.
  • tests/fm-backend-herdr.test.sh - the busy-baseline submit path itself in both directions, plus all ten real-capture verdicts above.
  • tests/fm-composer-blank-drift-live-e2e.test.sh - opt-in real-harness guard (FM_COMPOSER_BLANK_DRIFT=1), registered in the live-harness-optin family; bin/fm-composer-lib.sh now selects that family.

Compatibility reviewed across all four consuming adapters (suites pass); zellij does not use the classifier. bin/fm-lint.sh clean, bin/fm-doc-audience-check.sh clean.

Notes for the reviewer, outside this change

  • tests/fm-backend-herdr.test.sh already fails on main at two same-labeled home workspaces with no launcher identity must refuse (expected exit 3, got 1). Confirmed pre-existing on a clean HEAD; untouched here.
  • fm-pair-compose.sh releases the driver and navigator in sequence and creates ready only after both. A failure on the second release still leaves a half-released pair needing manual repair. Out of scope - the proven fix does not require it - but worth its own decision.

🤖 Generated with Claude Code

Claude 2.1.226 pads its EMPTY composer row with U+00A0 after the `❯` prompt
glyph. The blank is drawn at luminance 153, so ghost stripping correctly keeps
it, and every trim in the shared classifier and its callers is ASCII-only - so
a genuinely empty Claude composer classified `pending`, as if it still held
unsubmitted text.

Nothing failed loudly, because the composer verdict only decides delivery where
no stronger signal exists. An ordinary steer to an idle worker is confirmed
from Herdr's native agent state and never reads the composer at all. The one
path with no stronger signal is Herdr's submit confirmation for a target that
was ALREADY mid-turn before Enter, and that is exactly the path a paired-review
barrier release takes, because both roles sit in a blocking wait when the
release is sent. fm-send reported `delivery unconfirmed; verdict=pending` for a
message the driver had already received, and the composer correctly refused to
release the navigator, leaving the pair without coverage before its plan gate.

Fold non-ASCII Unicode blanks to ASCII space in fm_composer_classify_content,
the one fleet-wide owner every adapter delegates to. This can only move content
that is ENTIRELY blank from `pending` to `empty`; content carrying any visible
character keeps its verdict, so a genuinely swallowed Enter still refuses and
the away-mode injector still declines to type over unsent input. Keying on the
Unicode space-separator category rather than on U+00A0 alone keeps the fix from
depending on which blank a given release happens to pick.

Verified against real claude 2.1.226 on herdr 0.7.3 through the production
adapter, both directions: a message accepted by a mid-turn worker now confirms
(and was observably delivered), while text typed but never submitted still
reports pending.

Regressions:
- tests/fm-composer-lib.test.sh pins the fold, every declared blank, and the
  divergence - padded-empty is empty while padded-real-text stays pending, so
  the case cannot go vacuous - plus the unchanged dead-shell refusal.
- tests/fm-backend-herdr.test.sh pins the busy-baseline submit path this fault
  actually reached, in both directions, on captures carrying the real bytes.
- tests/fm-composer-blank-drift-live-e2e.test.sh is the opt-in real-harness
  guard, since what a harness draws in an empty composer is a vendor surface a
  fixture can only replay from a previous release. It records the untouched
  verdict, then asserts it only after probe text typed into that pane becomes
  visible, so a blank startup screen or a trust dialog is reported unverified
  instead of passing. It fails on claude without this fix and passes with it.

Compatibility: the shared owner is consumed by the tmux, herdr, cmux and orca
adapters; all four suites pass, and zellij does not use it.
…ents in the tree

The blank fold needed per-harness proof, not just proof for the harness that
provoked it. All five installed harnesses were measured against their own live
binary, idle and holding typed-but-unsent text, and the byte-exact captures are
kept under tests/fixtures/composer-captures/ and replayed through the real
Herdr composer reader.

Measured against the pre-fix classifier, exactly one of the ten verdicts moves:

  harness   version    idle(pre)  idle(now)  typed     pads a Unicode blank
  claude    2.1.226    pending    empty      pending   yes, U+00A0
  codex     0.145.0    empty      empty      pending   no
  pi        0.84.1     empty      empty      pending   no
  opencode  1.17.20    unknown    unknown    unknown   no
  agy       1.1.11     unknown    unknown    unknown   no

Claude is the only one of the five that pads its composer at all, so the other
four are evidence that the fold changed nothing for them, not evidence that it
repaired them. Recording that distinction matters more than a green row.

Two seams are unsupported on BOTH backends and are recorded rather than papered
over. opencode draws a composer row with a left-only edge glyph, matching
neither the bordered shape nor a bare agent prompt. agy draws a bare `>`,
byte-identical to a dead shell prompt, and the rule that would admit it on
native identity requires that row to sit below all separator activity while
real agy draws a rule underneath its prompt. Both refuse as `unknown`, which
denies delivery confirmation and away-mode injection alike, and neither is
affected by the fold. Reading either composer is a separate change with its own
safety argument and is not attempted here.

agy was also missing from the live guard's harness loop entirely, so the
agy-specific prompt rule had never been exercised against the real binary.

The guard no longer launches harnesses inside the checkout. During this work a
real harness launched in the firstmate worktree modified a tracked file with no
prompt ever submitted to it; the change was reverted and did not reach a commit.
A guard that runs live agents must not run them in the tree it is testing, so
launches now default to a scratch directory, and FM_COMPOSER_BLANK_DRIFT_DIR
takes an already-trusted disposable directory for fuller coverage and refuses
this checkout. The honest cost is coverage: from a scratch directory only pi
draws a composer here, because the others gate on directory trust first.
The gotmp fixture builds a fake FM_HOME and symlinks each library
bin/fm-teardown.sh sources into it. The orchestrator work (#32) added
bin/fm-supervisor-lib.sh, made teardown source it, and left the fixture's
symlink list untouched, so teardown died inside the test on a missing sibling:

  .../bin/fm-supervisor-lib.sh: No such file or directory

Real teardown was never affected - the file exists in a real bin/. Only the
fake root was incomplete, and this is fork-only: upstream has no
fm-supervisor-lib.sh and no source line for it, which is why upstream stays
green here.

fm-supervisor-lib.sh sources fm-primary-scope-lib.sh in turn, so both are
needed; symlinking only the first moves the same failure one library along.

This failure has been red on main since #32 landed. A permanently red CI stops
being able to report the next real regression, which is the actual cost.
@huynhtandat223
huynhtandat223 merged commit 87b2251 into main Aug 10, 2026
22 of 24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant