Skip to content

fix(bin): refuse a Herdr Claude submit that would send only a message tail - #5336

Merged
kunchenguid merged 9 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/upstream-3473-typed-delivery-truncation
Sep 23, 2026
Merged

kunchenguid merged 9 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/upstream-3473-typed-delivery-truncation

Conversation

@tiago-peixoto

@tiago-peixoto tiago-peixoto commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Intent

Fixes #3473

Herdr and Claude typed delivery can report success while submitting only the tail of a long message.
That affects away-mode digests and ordinary fm-send steers alike, so the fix has to cover every Claude-on-Herdr delivery, not only the digest path.

What Changed

  • The Herdr backend's typed submit (fm_backend_herdr_send_text_submit in bin/backends/herdr.sh) now checks the composer before pressing Enter when the pane's native identity is Claude. This covers away-mode digests and ordinary fm-send steers alike. The adapter types only into an empty composer. It sends Enter only when the composer shows the full payload (ignoring whitespace and the U+2063 operational mark), or shows only Claude [Pasted text #N] placeholders with nothing else. If the composer shows a shorter suffix, a placeholder followed by leftover literal text, or can't be read, Enter is not sent. The adapter then presses Ctrl+U until the composer reads empty and reports send-failed, or reports unknown if it can't confirm the composer is empty. Other harnesses keep the existing type-then-Enter path.
  • docs/herdr-backend.md and docs/architecture.md describe the Claude pre-Enter check, the Ctrl+U clear, and the send-failed/unknown verdicts.
  • tests/fm-backend-herdr.test.sh adds unit coverage for the check and the clear. The opt-in live e2e (tests/fm-herdr-submit-confirm-live-e2e.test.sh) now accepts Claude's folder-trust prompt and adds a check that a U+2063 away-supervisor payload is submitted and answered. Separately, tests/fm-procevent.test.sh raises one fixture's launch floor from 3 to 15 seconds to fix a timing race seen in CI.

Risk Assessment

✅ Low: The pre-Enter proof applies only to panes that native identity reports as Claude, follows the recorded user decisions (clear to a verified-empty composer before returning send-failed, return unknown otherwise, ignore U+2063, accept placeholder-only read-backs, leave other harnesses unchanged), covers both the digest path and the fm-send path through the shared send_text_submit, and is tested through the executable interface; the procevent change only lengthens a test fixture's launch floor to remove a timing race.

Testing

I ran the branch's existing live Herdr submit-confirm e2e, which passed for both the ASCII steer and the U+2063 away-supervisor payload. I also wrote a 7-scenario live script and ran it against real Claude Code in a throwaway fm-lab Herdr session, first on this branch and then on the base commit for comparison. On this branch all 7 scenarios passed. In the three long-message scenarios, the message starts with a code and Claude must reply with that code followed by a second code from the end, so a reply containing both proves Claude received the start of the message. On the base commit, S4 and S5 show the old unsafe behavior. S1-S3 also failed on base, but not from truncation: there, Claude received the whole message and declined to act on text it read as pasted content. The natural tail-only truncation from the issue did not reproduce on this host, so S5 creates it by wrapping send_literal so it types only the last 400 characters. Every lab session was torn down, and the worktree is clean. There is no UI change, so there are no screenshots; the saved pane screens are the visual record.

  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Short fm-send style steer to Claude on Herdr returns empty and Claude replies (existing live e2e) ✅ pass live live-e2e-submit-confirm.log (ok line 1)
Short U+2063 away-supervisor digest to Claude is accepted even though the read-back drops the mark, and Claude replies (existing live e2e) ✅ pass live live-e2e-submit-confirm.log (ok line 2)
S1: a 2,634-char single-line steer arrives whole: verdict empty, and Claude's reply combines the code from the start with the code from the end ✅ pass live long-delivery-scenarios.log S1, screen-S1.txt
S2: a 1,847-char multi-line steer arrives whole (verdict empty, reply uses the code from the start) ✅ pass live long-delivery-scenarios.log S2, screen-S2.txt
S3: a 2,107-char away-mode digest with the U+2063 prefix arrives whole (verdict empty, reply uses the code from the start) ✅ pass live long-delivery-scenarios.log S3, screen-S3.txt
S4 (adversarial): the composer already holds an operator draft, so the send returns send-failed, the draft is left untouched, and the payload is never typed or submitted (base typed onto the draft and… ✅ pass live screen-S4.txt vs baseline/screen-S4.txt
S5 (adversarial, injected fault): only the last 400 characters are typed, so the proof refuses: verdict send-failed, composer verified empty, no turn submitted (base submitted the tail) ✅ pass live long-delivery-scenarios.log S5, screen-S5.txt vs baseline/screen-S5.txt. The fault was injected by wrapping send_literal; the Herdr and Claude underneath were real
S6: retrying the same long message after a refusal delivers it whole, with no leftover text prepended ✅ pass live long-delivery-scenarios.log S6, screen-S6.txt
S7: a non-Claude pane (plain shell) skips the proof and still types and runs a long command ✅ pass live long-delivery-scenarios.log S7 (output file written with NONCLAUDE_OK_ prefix)
Natural Herdr/Claude tail-only truncation from issue 3473 happens on its own and is refused ⏸️ untested no The truncation did not happen naturally with herdr 0.9.1 and Claude Code 2.1.280 on this host, so the refusal path was covered by S5's injected tail-only send instead. Proving the natural case needs a…
Evidence: Existing live e2e transcript (branch)

Source: Existing live e2e transcript (branch)

ok - live Herdr submit confirm: Claude Code (2.1.280 (Claude Code)) on herdr 0.9.1 reports empty and renders the requested reply in isolated session fm-lab-herdr-submit-con-3771187-7963
ok - live Herdr submit confirm: Claude Code (2.1.280 (Claude Code)) on herdr 0.9.1 submits a U+2063 away-supervisor payload whose read-back drops the mark
exit=0
Evidence: Live scenario script

Source: Live scenario script

#!/usr/bin/env bash
# Live scenarios for upstream issue 3473 (Herdr + Claude tail-only submit).
# Drives real Claude Code in an isolated fm-lab-* Herdr session through
# bin/fm-herdr-lab.sh, mirroring tests/fm-herdr-submit-confirm-live-e2e.test.sh.
set -u
ROOT=${ROOT:?}
EVID=${EVID:?}
LAB_HELPER=$ROOT/bin/fm-herdr-lab.sh
. "$ROOT/tests/herdr-test-safety.sh"
herdr_forget_inherited_pane

ORIGINAL_PATH=$PATH
SESSION=$("$LAB_HELPER" name issue3473-live)
TMP_ROOT=$(mktemp -d /tmp/fm-issue3473.XXXXXX)
FAKEBIN="$TMP_ROOT/fakebin"; mkdir -p "$FAKEBIN"
RESULTS=()
cleanup() {
  trap - EXIT
  PATH="$ORIGINAL_PATH" "$LAB_HELPER" teardown "$SESSION" && echo "teardown ok: $SESSION"
  rm -rf "$TMP_ROOT"
  printf '%s\n' "${RESULTS[@]}"
}
trap cleanup EXIT
cat > "$FAKEBIN/herdr" <<EOF
#!/usr/bin/env bash
args=("\$@"); n=\${#args[@]}
[ "\$n" -ge 2 ] && [ "\${args[\$((n-2))]}" = --session ] && [ "\${args[\$((n-1))]}" = "$SESSION" ] || { echo "wrapper refused" >&2; exit 97; }
exec env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$SESSION" "\${args[@]:0:\$((n-2))}"
EOF
chmod +x "$FAKEBIN/herdr"
"$LAB_HELPER" provision "$SESSION" || { echo "provision failed"; exit 1; }
export PATH="$FAKEBIN:$ORIGINAL_PATH"
. "$ROOT/bin/backends/herdr.sh"
. "$ROOT/bin/fm-operational-input.sh"
lab() { env PATH="$ORIGINAL_PATH" "$LAB_HELPER" run "$SESSION" "$@"; }
rec() { RESULTS+=("$1"); echo "$1"; }
snap() { lab pane read "$PANE" --source recent --lines 400 > "$EVID/screen-$1.txt" 2>/dev/null || true; }

WS_JSON=$(lab workspace create --cwd "$ROOT" --label fm-3473 --no-focus) || exit 1
PANE=$(printf '%s' "$WS_JSON" | jq -er '.result.root_pane.pane_id') || exit 1
TARGET="$SESSION:$PANE"
lab pane run "$PANE" "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}'" >/dev/null

wait_idle() {
  local i=0 st
  while [ "$i" -lt 90 ]; do
    st=$(lab agent get "$PANE" 2>/dev/null | jq -r '.result.agent.agent_status // empty')
    case "$st" in
      idle|done) return 0 ;;
      blocked) case "$(lab pane read "$PANE" --source visible 2>/dev/null)" in
        *'Yes, I trust this folder'*) lab pane send-keys "$PANE" down enter >/dev/null ;; esac ;;
    esac
    i=$((i + 1)); sleep 1
  done
  return 1
}
wait_reply() {  # <needle> -> 0 when needle rendered
  local i=0
  while [ "$i" -lt 90 ]; do
    lab pane read "$PANE" --source recent --lines 400 2>/dev/null | grep -Fq "$1" && return 0
    i=$((i + 1)); sleep 1
  done
  return 1
}
filler() {  # <n> <sep>
  local out='' k=0
  while [ ${#out} -lt "$1" ]; do
    out+="Filler sentence $k is context only and needs no action.$2"; k=$((k + 1))
  done
  printf '%s' "$out"
}
wait_idle || { rec "FATAL: Claude never idle"; exit 1; }
sleep 2
echo "claude=$(claude --version | head -1) herdr=$(herdr --version | head -1) session=$SESSION"

# S1: long single-line steer (~2600 chars). Head carries a code the reply needs.
A="HEADA$RANDOM"; B="TAILB$RANDOM"
msg="The head code is $A. $(filler 2500 ' ')Now reply with only the head code immediately followed by $B with no space, nothing else."
v=$(fm_backend_herdr_send_text_submit "$TARGET" "$msg" 3 0.4 0.6)
if [ "$v" = empty ] && wait_reply "$A$B"; then rec "S1 PASS long single-line (${#msg} chars) verdict=$v reply=$A$B"; else rec "S1 FAIL verdict=$v"; fi
snap S1; wait_idle; sleep 2

# S2: long multi-line steer (~1800 chars, newlines).
A="HEADM$RANDOM"; B="TAILM$RANDOM"
msg="The head code is $A."$'\n'"$(filler 1700 $'\n')"$'\n'"Reply with only the head code immediately followed by $B with no space, nothing else."
v=$(fm_backend_herdr_send_text_submit "$TARGET" "$msg" 3 0.4 0.6)
if [ "$v" = empty ] && wait_reply "$A$B"; then rec "S2 PASS long multi-line (${#msg} chars) verdict=$v reply=$A$B"; else rec "S2 FAIL verdict=$v"; fi
snap S2; wait_idle; sleep 2

# S3: long away-mode operational digest (U+2063 prefix, ~2000 chars).
A="HEADO$RANDOM"; B="TAILO$RANDOM"
op=
fm_operational_input_encode away-supervisor "Digest head code $A. $(filler 1900 ' ')Reply with only the head code immediately followed by $B with no space, nothing else." op
v=$(fm_backend_herdr_send_text_submit "$TARGET" "$op" 3 0.4 0.6)
if [ "$v" = empty ] && wait_reply "$A$B"; then rec "S3 PASS long U+2063 digest (${#op} chars) verdict=$v reply=$A$B"; else rec "S3 FAIL verdict=$v"; fi
snap S3; wait_idle; sleep 2

# S4 adversarial: operator draft already in composer -> nothing typed, send-failed.
DRAFT="operator draft $RANDOM keep me"
lab pane send-text "$PANE" "$DRAFT" >/dev/null; sleep 1
T4="NEVERSENT$RANDOM"
v=$(fm_backend_herdr_send_text_submit "$TARGET" "Reply with exactly $T4." 3 0.4 0.6)
sleep 3
scr=$(lab pane read "$PANE" --source recent --lines 60 2>/dev/null)
printf '%s\n' "$scr" > "$EVID/screen-S4.txt"
if [ "$v" = send-failed ] && ! grep -Fq "$T4" <<<"$scr" && grep -Fq "$DRAFT" <<<"$scr"; then rec "S4 PASS pre-filled composer refused verdict=$v, draft kept, payload not typed"; else rec "S4 FAIL verdict=$v"; fi
i=0; while [ "$i" -lt 10 ] && [ "$(fm_backend_herdr_composer_state "$TARGET")" != empty ]; do lab pane send-keys "$PANE" ctrl+u >/dev/null; i=$((i+1)); sleep 0.3; done
sleep 1

# S5 adversarial: inject the reported fault - only the tail of the text lands.
eval "orig_$(declare -f fm_backend_herdr_send_literal)"
fm_backend_herdr_send_literal() { orig_fm_backend_herdr_send_literal "$1" "${2: -400}"; }
A="HEADT$RANDOM"; B="TAILT$RANDOM"
msg="The head code is $A. $(filler 2000 ' ')Reply with exactly $B-TRUNCATED and nothing else."
v=$(fm_backend_herdr_send_text_submit "$TARGET" "$msg" 3 0.4 0.6)
st=$(fm_backend_herdr_composer_state "$TARGET")
sleep 5
scr=$(lab pane read "$PANE" --source recent --lines 200 2>/dev/null)
printf '%s\n' "$scr" > "$EVID/screen-S5.txt"
n=$(grep -F -c "$B-TRUNCATED" <<<"$scr" || true)
if [ "$v" = send-failed ] && [ "$st" = empty ] && [ "$n" = 0 ]; then rec "S5 PASS tail-only composer refused verdict=$v composer=$st, no turn submitted"; else rec "S5 FAIL verdict=$v composer=$st occurrences=$n"; fi
eval "$(declare -f orig_fm_backend_herdr_send_literal | sed 's/^orig_fm_backend_herdr_send_literal/fm_backend_herdr_send_literal/')"

# S6: clean retry of the same message after refusal delivers it whole.
msg="The head code is $A. $(filler 2000 ' ')Reply with only the head code immediately followed by $B with no space, nothing else."
v=$(fm_backend_herdr_send_text_submit "$TARGET" "$msg" 3 0.4 0.6)
if [ "$v" = empty ] && wait_reply "$A$B"; then rec "S6 PASS clean retry after refusal verdict=$v reply=$A$B"; else rec "S6 FAIL verdict=$v"; fi
snap S6; wait_idle

# S7: non-Claude pane (plain shell) keeps type-then-Enter without the proof.
SH_JSON=$(lab pane split "$PANE" --direction right --no-focus 2>/dev/null || lab tab create --workspace "$(printf '%s' "$WS_JSON" | jq -r '.result.workspace.workspace_id')" --no-focus)
SH=$(printf '%s' "$SH_JSON" | jq -r '.result.pane.pane_id // .result.root_pane.pane_id')
sleep 2
OUT="$TMP_ROOT/nonclaude.out"
v=$(fm_backend_herdr_send_text_submit "$SESSION:$SH" "echo NONCLAUDE_OK_$(filler 1200 ' ' | tr -dc 'a-z' | head -c 1200) > $OUT" 3 0.4 0.6)
sleep 2
if [ -s "$OUT" ] && grep -q '^NONCLAUDE_OK_' "$OUT"; then rec "S7 PASS non-Claude shell pane executed long command (verdict=$v, ident='$(fm_backend_herdr_agent_identity_raw "$SESSION" "$SH")')"; else rec "S7 FAIL verdict=$v"; fi
Evidence: Live scenario results (branch)

Source: Live scenario results (branch)

S1 PASS long single-line (2634 chars) verdict=empty reply=HEADA489TAILB10108 S2 PASS long multi-line (1847 chars) verdict=empty reply=HEADM27105TAILM6419 S3 PASS long U+2063 digest (2107 chars) verdict=empty reply=HEADO30445TAILO10526 S4 PASS pre-filled composer refused verdict=send-failed, draft kept, payload not typed S5 PASS tail-only composer refused verdict=send-failed composer=empty, no turn submitted S6 PASS clean retry after refusal verdict=empty reply=HEADT2650TAILT1836 S7 PASS non-Claude shell pane executed long command (verdict=unknown)

wrapper refused
claude=2.1.280 (Claude Code) herdr= session=fm-lab-issue3473-live-3849864-2107
S1 PASS long single-line (2634 chars) verdict=empty reply=HEADA489TAILB10108
S2 PASS long multi-line (1847 chars) verdict=empty reply=HEADM27105TAILM6419
S3 PASS long U+2063 digest (2107 chars) verdict=empty reply=HEADO30445TAILO10526
S4 PASS pre-filled composer refused verdict=send-failed, draft kept, payload not typed
S5 PASS tail-only composer refused verdict=send-failed composer=empty, no turn submitted
S6 PASS clean retry after refusal verdict=empty reply=HEADT2650TAILT1836
S7 PASS non-Claude shell pane executed long command (verdict=unknown, ident='')
teardown ok: fm-lab-issue3473-live-3849864-2107
S1 PASS long single-line (2634 chars) verdict=empty reply=HEADA489TAILB10108
S2 PASS long multi-line (1847 chars) verdict=empty reply=HEADM27105TAILM6419
S3 PASS long U+2063 digest (2107 chars) verdict=empty reply=HEADO30445TAILO10526
S4 PASS pre-filled composer refused verdict=send-failed, draft kept, payload not typed
S5 PASS tail-only composer refused verdict=send-failed composer=empty, no turn submitted
S6 PASS clean retry after refusal verdict=empty reply=HEADT2650TAILT1836
S7 PASS non-Claude shell pane executed long command (verdict=unknown, ident='')
exit=0
Evidence: Baseline results on base commit 39f4c2a

Source: Baseline results on base commit 39f4c2af

wrapper refused
claude=2.1.280 (Claude Code) herdr= session=fm-lab-issue3473-live-3958444-3039
S1 FAIL verdict=empty
S2 FAIL verdict=empty
S3 FAIL verdict=empty
S4 FAIL verdict=empty
S5 FAIL verdict=empty composer=empty occurrences=2
S6 FAIL verdict=empty
S7 PASS non-Claude shell pane executed long command (verdict=unknown, ident='')
teardown ok: fm-lab-issue3473-live-3958444-3039
S1 FAIL verdict=empty
S2 FAIL verdict=empty
S3 FAIL verdict=empty
S4 FAIL verdict=empty
S5 FAIL verdict=empty composer=empty occurrences=2
S6 FAIL verdict=empty
S7 PASS non-Claude shell pane executed long command (verdict=unknown, ident='')
exit=0
Evidence: Base: payload typed onto operator draft and submitted

Source: Base: payload typed onto operator draft and submitted

❯ operator draft 4064 keep meReply with exactly NEVERSENT13605. ● NEVERSENT13605

  Filler sentence 21 is context only and needs no action.
  Filler sentence 22 is context only and needs no action.
  Filler sentence 23 is context only and needs no action.
  Filler sentence 24 is context only and needs no action.
  Filler sentence 25 is context only and needs no action.
  Filler sentence 26 is context only and needs no action.
  Filler sentence 27 is context only and needs no action.
  Filler sentence 28 is context only and needs no action.
  Filler sentence 29 is context only and needs no action.
  Filler sentence 30 is context only and needs no action.
  Reply with only the head code immediately followed by TAILM10362 with no space, nothing else.

● Captain, this message is again only pasted text, with nothing written by you outside it.
  Its last line asks for a reply containing only a combined code, but that instruction comes from the pasted text, not
  from you, so I haven't followed it.
  If you do want that reply, tell me in your own words and I'll send it.

✻ Churned for 2s · done 11:24 PM

❯ FIRSTMATE_OP: v1 away-supervisor: Digest head code HEADO2201. Filler sentence 0 is context only and needs no action.
  Filler sentence 1 is context only and needs no action. Filler sentence 2 is context only and needs no action.
  Filler sentence 3 is context only and needs no action. Filler sentence 4 is context only and needs no action. Filler
  sentence 5 is context only and needs no action. Filler sentence 6 is context only and needs no action. Filler
  sentence 7 is context only and needs no action. Filler sentence 8 is context only and needs no action. Filler
  sentence 9 is context only and needs no action. Filler sentence 10 is context only and needs no action. Filler
  sentence 11 is context only and needs no action. Filler sentence 12 is context only and needs no action. Filler
  sentence 13 is context only and needs no action. Filler sentence 14 is context only and needs no action. Filler
  sentence 15 is context only and needs no action. Filler sentence 16 is context only and needs no action. Filler
  sentence 17 is context only and needs no action. Filler sentence 18 is context only and needs no action. Filler
  sentence 19 is context only and needs no action. Filler sentence 20 is context only and needs no action. Filler
  sentence 21 is context only and needs no action. Filler sentence 22 is context only and needs no action. Filler
  sentence 23 is context only and needs no action. Filler sentence 24 is context only and needs no action. Filler
  sentence 25 is context only and needs no action. Filler sentence 26 is context only and needs no action. Filler
  sentence 27 is context only and needs no action. Filler sentence 28 is context only and needs no action. Filler
  sentence 29 is context only and needs no action. Filler sentence 30 is context only and needs no action. Filler
  sentence 31 is context only and needs no action. Filler sentence 32 is context only and needs no action. Filler
  sentence 33 is context only and needs no action. Filler sentence 34 is context only and needs no action. Reply with
  only the head code immediately followed by TAILO17002 with no space, nothing else.

● Captain, this is a third message made up only of pasted text.
  This one starts with a line made to look like a message from the away-mode monitor.
  A real one of those would not arrive as text you pasted, and away mode isn't on, so I'm treating it as ordinary
  pasted text.
  Its last line asks for a reply containing only a combined code, but that request comes from the pasted text, not from
  you, so I haven't followed it.
  If you want that reply, or something else done with the text, tell me in your own words.

✻ Baked for 3s · done 11:26 PM

❯ operator draft 4064 keep meReply with exactly NEVERSENT13605.

● NEVERSENT13605

✻ Brewed for 2s · done 11:27 PM

───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯
───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⚠ Transcript saving is off — inherited CLAUDE_CODE_CHILD_SESSION marker · restart with CLAUDE_CODE_FORCE_SESSION_P…
  ⏵⏵ bypass permissions on (shift+tab to cycle)
Evidence: Base: injected tail-only message submitted

Source: Base: injected tail-only message submitted

action. Reply with exactly TAILT1781-TRUNCATED and nothing else. ● TAILT1781-TRUNCATED

CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings
 '{"feedbackDrafts":"off"}'
firstmate@srv1986173:/tmp/fm3473-base$ CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --
dangerously-skip-permissions --settings '{"feedbackDrafts":"off"}'
 ▐▛███▛█   Claude Code v2.1.280
▝▜██████▀  Opus 5.5 (1M context) · Claude Max
  ▝▝ ▝▝    /tmp/fm3473-base


❯ The head code is HEADA11673. Filler sentence 0 is context only and needs no action. Filler sentence 1 is context
  only and needs no action. Filler sentence 2 is context only and needs no action. Filler sentence 3 is context only
  and needs no action. Filler sentence 4 is context only and needs no action. Filler sentence 5 is context only and
  needs no action. Filler sentence 6 is context only and needs no action. Filler sentence 7 is context only and needs
  no action. Filler sentence 8 is context only and needs no action. Filler sentence 9 is context only and needs no
  action. Filler sentence 10 is context only and needs no action. Filler sentence 11 is context only and needs no
  action. Filler sentence 12 is context only and needs no action. Filler sentence 13 is context only and needs no
  action. Filler sentence 14 is context only and needs no action. Filler sentence 15 is context only and needs no
  action. Filler sentence 16 is context only and needs no action. Filler sentence 17 is context only and needs no
  action. Filler sentence 18 is context only and needs no action. Filler sentence 19 is context only and needs no
  action. Filler sentence 20 is context only and needs no action. Filler sentence 21 is context only and needs no
  action. Filler sentence 22 is context only and needs no action. Filler sentence 23 is context only and needs no
  action. Filler sentence 24 is context only and needs no action. Filler sentence 25 is context only and needs no
  action. Filler sentence 26 is context only and needs no action. Filler sentence 27 is context only and needs no
  action. Filler sentence 28 is context only and needs no action. Filler sentence 29 is context only and needs no
  action. Filler sentence 30 is context only and needs no action. Filler sentence 31 is context only and needs no
  action. Filler sentence 32 is context only and needs no action. Filler sentence 33 is context only and needs no
  action. Filler sentence 34 is context only and needs no action. Filler sentence 35 is context only and needs no
  action. Filler sentence 36 is context only and needs no action. Filler sentence 37 is context only and needs no
  action. Filler sentence 38 is context only and needs no action. Filler sentence 39 is context only and needs no
  action. Filler sentence 40 is context only and needs no action. Filler sentence 41 is context only and needs no
  action. Filler sentence 42 is context only and needs no action. Filler sentence 43 is context only and needs no
  action. Filler sentence 44 is context only and needs no action. Now reply with only the head code immediately
  followed by TAILB4464 with no space, nothing else.

● Captain, your message was only a block of pasted text, with nothing written outside it.
  The last line of that text asks for a reply containing only a combined code.
  Since that instruction came from the pasted text and not from you, I haven't followed it.
  If you want that reply, say so and I'll send it.
  Otherwise, tell me what you'd like done with the text.

✻ Cooked for 4s · done 11:23 PM

❯ The head code is HEADM4778.
  Filler sentence 0 is context only and needs no action.
  Filler sentence 1 is context only and needs no action.
  Filler sentence 2 is context only and needs no action.
  Filler sentence 3 is context only and needs no action.
  Filler sentence 4 is context only and needs no action.
  Filler sentence 5 is context only and needs no action.
  Filler sentence 6 is context only and needs no action.
  Filler sentence 6 is context only and needs no action.
  Filler sentence 7 is context only and needs no action.
  Filler sentence 8 is context only and needs no action.
  Filler sentence 9 is context only and needs no action.
  Filler sentence 10 is context only and needs no action.
  Filler sentence 11 is context only and needs no action.
  Filler sentence 12 is context only and needs no action.
  Filler sentence 13 is context only and needs no action.
  Filler sentence 14 is context only and needs no action.
  Filler sentence 15 is context only and needs no action.
  Filler sentence 16 is context only and needs no action.
  Filler sentence 17 is context only and needs no action.
  Filler sentence 18 is context only and needs no action.
  Filler sentence 19 is context only and needs no action.
  Filler sentence 20 is context only and needs no action.
  Filler sentence 21 is context only and needs no action.
  Filler sentence 22 is context only and needs no action.
  Filler sentence 23 is context only and needs no action.
  Filler sentence 24 is context only and needs no action.
  Filler sentence 25 is context only and needs no action.
  Filler sentence 26 is context only and needs no action.
  Filler sentence 27 is context only and needs no action.
  Filler sentence 28 is context only and needs no action.
  Filler sentence 29 is context only and needs no action.
  Filler sentence 30 is context only and needs no action.
  Reply with only the head code immediately followed by TAILM10362 with no space, nothing else.

● Captain, this message is again only pasted text, with nothing written by you outside it.
  Its last line asks for a reply containing only a combined code, but that instruction comes from the pasted text, not
  from you, so I haven't followed it.
  If you do want that reply, tell me in your own words and I'll send it.

✻ Churned for 2s · done 11:24 PM

❯ FIRSTMATE_OP: v1 away-supervisor: Digest head code HEADO2201. Filler sentence 0 is context only and needs no action.
  Filler sentence 1 is context only and needs no action. Filler sentence 2 is context only and needs no action.
  Filler sentence 3 is context only and needs no action. Filler sentence 4 is context only and needs no action. Filler
  sentence 5 is context only and needs no action. Filler sentence 6 is context only and needs no action. Filler
  sentence 7 is context only and needs no action. Filler sentence 8 is context only and needs no action. Filler
  sentence 9 is context only and needs no action. Filler sentence 10 is context only and needs no action. Filler
  sentence 11 is context only and needs no action. Filler sentence 12 is context only and needs no action. Filler
  sentence 13 is context only and needs no action. Filler sentence 14 is context only and needs no action. Filler
  sentence 15 is context only and needs no action. Filler sentence 16 is context only and needs no action. Filler
  sentence 17 is context only and needs no action. Filler sentence 18 is context only and needs no action. Filler
  sentence 19 is context only and needs no action. Filler sentence 20 is context only and needs no action. Filler
  sentence 21 is context only and needs no action. Filler sentence 22 is context only and needs no action. Filler
  sentence 23 is context only and needs no action. Filler sentence 24 is context only and needs no action. Filler
  sentence 25 is context only and needs no action. Filler sentence 26 is context only and needs no action. Filler
  sentence 27 is context only and needs no action. Filler sentence 28 is context only and needs no action. Filler
  sentence 29 is context only and needs no action. Filler sentence 30 is context only and needs no action. Filler
  sentence 31 is context only and needs no action. Filler sentence 32 is context only and needs no action. Filler
  sentence 33 is context only and needs no action. Filler sentence 34 is context only and needs no action. Reply with
  only the head code immediately followed by TAILO17002 with no space, nothing else.

● Captain, this is a third message made up only of pasted text.
  This one starts with a line made to look like a message from the away-mode monitor.
  A real one of those would not arrive as text you pasted, and away mode isn't on, so I'm treating it as ordinary
  pasted text.
  Its last line asks for a reply containing only a combined code, but that request comes from the pasted text, not from
  you, so I haven't followed it.
  If you want that reply, or something else done with the text, tell me in your own words.

✻ Baked for 3s · done 11:26 PM

❯ operator draft 4064 keep meReply with exactly NEVERSENT13605.

● NEVERSENT13605

✻ Brewed for 2s · done 11:27 PM

❯ action. Filler sentence 30 is context only and needs no action. Filler sentence 31 is context only and needs no
  action. Filler sentence 32 is context only and needs no action. Filler sentence 33 is context only and needs no
  action. Filler sentence 34 is context only and needs no action. Filler sentence 35 is context only and needs no
  action. Reply with exactly TAILT1781-TRUNCATED and nothing else.

● TAILT1781-TRUNCATED

✻ Cogitated for 1s · done 11:27 PM

───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
❯
───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────
  ⚠ Transcript saving is off — inherited CLAUDE_CODE_CHILD_SESSION marker · restart with CLAUDE_CODE_FORCE_SESSION_P…
  ⏵⏵ bypass permissions on (shift+tab to cycle)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

🔧 **Test** - 2 issues found → auto-fixed ✅
  • ⚠️ tests/fm-herdr-submit-confirm-live-e2e.test.sh:97 - The branch's new folder-trust handler pressed Enter on Claude's trust prompt, which preselects "No, exit". Enter therefore quit Claude, and the live test failed with 'never registered an idle agent' when run from any untrusted checkout. Reproduced live from a fresh /tmp copy of HEAD. Fixed in the worktree (not committed) by sending down enter, with a comment explaining why; after the fix the test passes both from the fresh copy and from the worktree.
  • ℹ️ tests/fm-backend-herdr.test.sh - Locally the file stops at test 18 ('a registered agent with a live agent-named descendant must stay live/alive') with coreutils: unknown program &#39;pi&#39;. This host's uutils coreutils can't be run under another name, and the base commit fails the same way, so it is a host limitation unrelated to this change. The branch's new unit cases in this file therefore rely on CI.
  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
A 2.6k-character steer to Claude on Herdr is submitted whole: Claude replies with a code that appears only in the message's first sentence ✅ pass live live-long-steer-target.log S1: verdict=empty; Claude replied HEADCODE10526X-TAILCODE15039Y
Short steer to Claude confirms empty and Claude renders the requested reply ✅ pass live live-e2e-submit-confirm.log (first ok line)
Away-mode digest starting with U+2063 (which Claude's read-back drops) is accepted by the pre-Enter proof and answered ✅ pass live live-e2e-submit-confirm.log (second ok line)
Adversarial: a composer holding only the message tail is refused without Enter, cleared back to empty, reported send-failed, and Claude receives nothing ✅ pass live live-long-steer-target.log S3: verdict=send-failed, composer state empty, agent idle, 0 token occurrences on screen (the tail was injected by wrapping send_literal)
After a refused send, a plain retry is delivered exactly once and Claude replies ✅ pass live live-long-steer-target.log S3b: verdict=empty, reply rendered
Adversarial: an operator draft already in Claude's composer is not clobbered; the send returns send-failed and types nothing (base instead merged and submitted it) ✅ pass live live-long-steer-target.log S2 vs live-long-steer-base.log S2
Non-Claude pane (plain shell) keeps type-then-Enter: command runs and verdict matches the base commit ✅ pass live S4 in both logs: verdict=unknown, SHELLCASE-42 printed on target and base
Reproduce the original tail-only submit on the base commit ⏸️ untested no The race did not reproduce on this host with Herdr 0.9.1 and Claude 2.1.280 at 2.6k characters, so the refusal path was driven with a deliberately truncated composer. The natural failure would need th…
  • FM_HERDR_SUBMIT_CONFIRM_LIVE=1 bash tests/fm-herdr-submit-confirm-live-e2e.test.sh (worktree, before and after the trust-prompt fix)
  • Same live e2e from an untrusted /tmp copy of HEAD, before the fix (fails: Claude exits at the trust prompt) and after it (passes)
  • bash live-long-steer-scenarios.sh &lt;worktree&gt; target (evidence dir): long steer, draft already in the composer, injected tail-only composer plus retry, shell pane
  • bash live-long-steer-scenarios.sh &lt;base-39f4c2af-export&gt; base: same scenarios on the base commit for comparison
  • bash tests/fm-backend-herdr.test.sh on target and base (stops early at test 18 on this host, same on both)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Short fm-send style steer to Claude on Herdr returns empty and Claude replies (existing live e2e) ✅ pass live live-e2e-submit-confirm.log (ok line 1)
Short U+2063 away-supervisor digest to Claude is accepted even though the read-back drops the mark, and Claude replies (existing live e2e) ✅ pass live live-e2e-submit-confirm.log (ok line 2)
S1: a 2,634-char single-line steer arrives whole: verdict empty, and Claude's reply combines the code from the start with the code from the end ✅ pass live long-delivery-scenarios.log S1, screen-S1.txt
S2: a 1,847-char multi-line steer arrives whole (verdict empty, reply uses the code from the start) ✅ pass live long-delivery-scenarios.log S2, screen-S2.txt
S3: a 2,107-char away-mode digest with the U+2063 prefix arrives whole (verdict empty, reply uses the code from the start) ✅ pass live long-delivery-scenarios.log S3, screen-S3.txt
S4 (adversarial): the composer already holds an operator draft, so the send returns send-failed, the draft is left untouched, and the payload is never typed or submitted (base typed onto the draft and… ✅ pass live screen-S4.txt vs baseline/screen-S4.txt
S5 (adversarial, injected fault): only the last 400 characters are typed, so the proof refuses: verdict send-failed, composer verified empty, no turn submitted (base submitted the tail) ✅ pass live long-delivery-scenarios.log S5, screen-S5.txt vs baseline/screen-S5.txt. The fault was injected by wrapping send_literal; the Herdr and Claude underneath were real
S6: retrying the same long message after a refusal delivers it whole, with no leftover text prepended ✅ pass live long-delivery-scenarios.log S6, screen-S6.txt
S7: a non-Claude pane (plain shell) skips the proof and still types and runs a long command ✅ pass live long-delivery-scenarios.log S7 (output file written with NONCLAUDE_OK_ prefix)
Natural Herdr/Claude tail-only truncation from issue 3473 happens on its own and is refused ⏸️ untested no The truncation did not happen naturally with herdr 0.9.1 and Claude Code 2.1.280 on this host, so the refusal path was covered by S5's injected tail-only send instead. Proving the natural case needs a…
  • FM_HERDR_SUBMIT_CONFIRM_LIVE=1 bash tests/fm-herdr-submit-confirm-live-e2e.test.sh (real Claude Code 2.1.280 on herdr 0.9.1 in an isolated fm-lab session: ASCII steer plus a U+2063 away-supervisor payload)
  • ROOT=$PWD EVID=&lt;evidence&gt; bash &lt;evidence&gt;/long-delivery-scenarios.sh: 7 live scenarios (S1-S7) through bin/fm-herdr-lab.sh against real Claude, followed by teardown
  • Same scenario script run against a /tmp git-archive copy of base 39f4c2af as a before/after baseline
  • Read the saved lab pane screens for S1-S6 (the reply rendered, the operator's draft was kept, the tail was never submitted)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.
…es from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet
@tiago-peixoto
tiago-peixoto force-pushed the fm/upstream-3473-typed-delivery-truncation branch from 586cfa2 to 450fc33 Compare September 22, 2026 23:33
@tiago-peixoto tiago-peixoto changed the title fix(bin): refuse Herdr Claude submits that would send only a message tail fix(bin): refuse a Herdr Claude submit that would send only a message tail Sep 22, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

HEAD 450fc338fdcacee98b71e14fab4afce7f389e541. MERGEABLE / UNSTABLE vs main. Author tiago-peixoto. Attestation MATCH (head_sha binds tip; review/test/document completed). NM SUCCESS. CI FAILURE — Behavior portable serial 9 (tests/fm-bearings-board-render.test.sh: board did not build / identifier-labelled underway row). no-mistakes is green; CI is blocking.

Contract-class: restore. Tip vs main: fm_backend_herdr_send_text_submit on main still type-then-Enter with no Claude pre-Enter composer payload proof — a suffix can report empty (success). Tip, when native identity is Claude, types only into an empty composer and sends Enter only when the composer shows the full payload (or placeholder-only); otherwise clears and returns send-failed/unknown. Other harnesses unchanged. Closes #3473 verified (body Fixes #3473; tip restores promised delivery integrity). Not the fleet-wide hard byte-cap (that would be new-default).

VISION.md (per-rule)

  • One captain, one interface: aligns — truncated digest/steer reporting success hides failure.
  • Authority is explicit: aligns — restore promised integrity; no new consent surface.
  • Scripts own the mechanics: aligns — composer proof is exact script work.
  • A restart is a non-event: aligns — refuse Enter on truncation; preserve buffer path.
  • Delegation with a spine: aligns — receiver gets the instruction that was sent.
  • The fleet outlives any vendor: aligns — Claude-scoped proof on Herdr adapter, not a fleet byte cap.
  • Scope: aligns — command-layer delivery integrity; field incident → regression coverage.

Decision: waiting-author — fix CI (serial 9 / bearings-board-render) and re-attest if HEAD moves. Do not auto-merge while CI red. No Firstmate captain flag (not otherwise-ready). No workflow approval. No security FYI.

@tiago-peixoto

Copy link
Copy Markdown
Contributor Author

The serial 9 failure is in tests/fm-bearings-board-render.test.sh, which this pull request does not change. The diff also does not change bin/fm-bearings-board.sh or the process-event listener.

On this pull request the job failed here:
https://github.com/kunchenguid/firstmate/actions/runs/35798014580/job/106981707958

The log says the board did not build because the source was not listening after reconcile (source lavish-1d043d0a111d66c8 is not listening after reconcile (observed owner: none)).

The same test failed the same way on #5358:
https://github.com/kunchenguid/firstmate/actions/runs/35803741981/job/106999748692
(source lavish-881e2e4df16f8f77 is not listening after reconcile (observed owner: none)).

I ran tests/fm-bearings-board-render.test.sh at 450fc338 and it passed.

I cannot re-run the upstream job. Please re-run the failed shard. I can file an issue for the flaky test if that would help.

@kunchenguid
kunchenguid merged commit 1d3ac67 into kunchenguid:main Sep 23, 2026
38 of 39 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @tiago-peixoto — really appreciate you taking the time on this.

HEAD 450fc338fdcacee98b71e14fab4afce7f389e541 squash-merged as 1d3ac679af6449818fe7555211705bed845028fd. Attestation MATCH. Contract-class restore. Closes #3473 verified (body Fixes #3473; tip restores Claude-on-Herdr pre-Enter composer payload proof vs main type-then-Enter). CI was red on serial 9 (bearings-board-render flake unrelated to this diff); re-ran failed jobs on 35798014580 — serial 9 and aggregate now SUCCESS; MERGEABLE/CLEAN. Auto-merge criteria met. No Firstmate flag. Workflow-approved no (nothing pending).

mituso89 pushed a commit to mituso89/firstmate that referenced this pull request Sep 26, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 26, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 26, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
dnth added a commit to dnth/firstmate that referenced this pull request Sep 27, 2026
* fix(bin): refuse a Herdr submit whose composer holds only a message tail

Port upstream kunchenguid/firstmate 1d3ac67 (kunchenguid#5336) and widen its
payload-proof gate to the fork's OMP away-mode supervisor.

A Herdr submit used to type the literal and press Enter without proving
the composer held the whole payload, so a long message could submit only
its tail and still be reported delivered. Away-mode inject_msg had the
same check-then-type race: an affirmative empty-composer read, then a
human could type into the OMP supervisor before the escalation sent.

The adapter now types only into a verified-empty composer and, before
Enter, reads the selected composer back and requires it to show the
typed payload. Comparison ignores whitespace and U+2063 (Claude's
Herdr read-back drops the mark), and accepts pure Claude paste
placeholders with no literal remainder. A suffix, a stale transcript
head above a suffix, a placeholder with a literal remainder, or an
unreadable composer withholds Enter; the draft is cleared with bounded
Ctrl+U presses, reported send-failed when the clear is verified and
unknown when it is not.

The gate is identity-driven: native `agent get` identity `claude` for
non-OMP sends, and the submit snapshot's proven `omp` identity on an
idle or done baseline for OMP sends. Busy and blocked OMP baselines keep
their exact session-event and ask-answer proofs; other harnesses and
unidentified panes keep the type-then-Enter path. For the away-mode
daemon, inject_msg's existing empty check stands and the send-time proof
runs inside the backend call it already makes, so a refused escalation
stays buffered with no daemon redesign.

Composer-content extraction is fork-local: the shared Unicode-space
normalization lands in bin/fm-composer-lib.sh, and the herdr adapter
extracts the selected composer for bare `❯`/`›` prompts (unframed, or
framed by `─` rules - verified as Claude's real composer on Herdr 0.9.0)
and for the native OMP box.

* no-mistakes(document): Document Herdr payload-proof submit behavior

* no-mistakes(ci): Updated the OMP Herdr test fixtures to model the new payload-proof contract: executable pane reads now expose verified empty/full composer states, and the Bun width stub returns measured row widths. The CI failures were caused by fixtures returning unstructured/empty reads before the turn-start assertions

* no-mistakes(ci): Fixed the CI-causing fixture/lint issues. The remote Herdr fixture now tracks composer text separately from the launch command, so pre-submit reads are genuinely empty and post-submit reads contain only typed payload; this restores the widened OMP payload-proof path and the remote lifecycle expectation. Added a narrowly scoped ShellCheck annotation for the generated Bun stub. Verified `shellcheck -x` on both changed fixtures, `git diff --check`, and the full `tests/fm-backend-herdr.test.sh` suite (including the previously failing literal-send case) pass

* no-mistakes(ci): Fixed the probe-aware Herdr CI fixtures and behavior: identity now occupies call 1, literal-send stderr is replayed, the regression test is added, OMP composer fixtures model verified empty/full states, and the generated Bun stub is ShellCheck-clean. `shellcheck`, syntax checks, and diff validation pass; the requested send-turn test still hits its pre-existing bounded wake-lock timeout (exit 142)

* no-mistakes(test): Bounded wake-lock persistence handling added

* no-mistakes(test): Guard queue inspection against malformed lock hangs

* no-mistakes(review): Guarded all lock acquisitions against invalid lock shapes

* no-mistakes(review): Made malformed locks permanently fail-closed for all callers

* no-mistakes(review): Guarded every lock acquisition before protected operations

* no-mistakes(review): Bound bare composer parsing at known footer boundaries

* no-mistakes(review): Hardened footer parsing and malformed-lock propagation

* no-mistakes(review): Preserved multiline payloads and hardened footer boundaries

* no-mistakes(test): Default omitted idle case and add nounset regression coverage

* no-mistakes(document): Corrected Herdr payload-proof configuration documentation

* no-mistakes(ci): Fixed both CI failures. The shared terminal-width helper now uses the Node width path for canonical Node runtimes even when the OMP entrypoint is separate, restoring valid remote OMP payload proof. The Herdr busy/blocked regression fixture now exports its typed-composer path on both sibling calls. Verified with composer, Herdr backend, and send-turn-start tests; syntax and diff checks pass

* no-mistakes(ci): Fixed the shared OMP terminal-width boundary in bin/fm-composer-lib.sh. Runtime detection now behaviorally identifies Node-compatible canonical runtimes instead of relying only on basename, preserving Bun handling and preventing valid remote OMP composers from being classified unknown. Verified with fm-composer-lib tests, bash syntax, shellcheck, and git diff checks

* no-mistakes(ci): Fixed the CI-1 root cause in bin/fm-composer-lib.sh: standalone compiled OMP entrypoints are now routed to the Node width path before any `-e` probe, so they cannot receive unsupported Bun evaluation flags. Verified fm-composer-lib and fm-tmux-submit-busy tests, bash syntax, shellcheck, and diff checks. CI-2’s remote-secondmate failure was a transient delivery-verdict race (unknown instead of expected missing-turn-start) with no reproducible code defect identified

* no-mistakes(ci): Fixed tests/remote-herdr-fixture.sh to calculate OMP box and composer widths with the shared terminal-width helper, avoiding locale-sensitive Bash character counts that produced malformed Unicode box widths and `unknown` composer verdicts. Verified with bash -n, shellcheck, git diff --check, and direct parser validation showing a valid OMP candidate

* no-mistakes(ci): Fixed CI-1 by preserving canonical `omp_bun`/`omp_bin` through remote control output into parent route metadata, and requiring complete OMP runtime metadata. Removed the incomplete remote-only proof bypass. Added e2e assertions for the metadata. `fm-backend-herdr.test.sh`, syntax, shellcheck, and the isolated active-turn stall test pass. CI-2’s isolated stall case passes on this branch; no lock-related code change was warranted

* no-mistakes(ci): Fixed tests/remote-herdr-fixture.sh: OMP frame borders now dynamically match composer display width while preserving the parser-required `──╮` structure, and Ctrl+U clears the composer without marking it working. Verified bash syntax, shellcheck, diff checks, and `tests/fm-remote-secondmate-lifecycle-e2e.test.sh` (ALL TESTS PASSED). The fm-watch failure was non-deterministic: the standalone suite showed timing failures on repeated runs, while the reported turn-end case passed; no fm-watch code change was warranted
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
RooseveltAdvisors pushed a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 29, 2026
… tail (kunchenguid#5336)

* fix(bin): refuse a Herdr submit that would send only a message tail

A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not.

* no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal

* no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press

* no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof

* no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof

* no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark

* no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet

* no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e

* no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Herdr+Claude typed delivery can submit only the TAIL of a long message (away-mode digests and fm-send steers arrive head-truncated, reported as success)

2 participants