Skip to content

Telegram captain-comms, shipped into the repo - #1

Merged
Alchemy86 merged 321 commits into
mainfrom
fm/fm-telegram-ship-t8
Aug 23, 2026
Merged

Alchemy86 merged 321 commits into
mainfrom
fm/fm-telegram-ship-t8

Conversation

@Alchemy86

Copy link
Copy Markdown
Owner

Moves the Telegram captain-comms machinery out of the gitignored state/ directory and into tracked bin/, so a configured clone works without hand-holding.

Includes the four guarantees measured against real failures tonight: only firstmate contacts the captain (crew sessions are refused), no message is retired unanswered, no message is answered twice, and the '...' acknowledgement does not count as a reply.

Also carries an upstream catch-up (16 commits) picked up during the work.

Validation note, stated plainly: the no-mistakes pipeline could not complete. Six consecutive review attempts stalled before emitting a findings table, each time after 20+ minutes of no filesystem activity and no live process. Every finding that any partial run did surface was fixed by hand rather than lost. Directly verified instead: fm-telegram 50/50, fm-pi-watch-extension 33/33, fm-supervision-instructions 10/10, shellcheck clean.

Ballestar and others added 30 commits July 18, 2026 03:39
…nsion (kunchenguid#205)

* fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn

Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4
(notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom
in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(brief): harden fm-brief regression coverage for parse and scaffolds

Tighten bash -n checking, pin literal backtick rendering in the no-mistakes
DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the

Co-authored-by: Cursor <cursoragent@cursor.com>
kunchenguid#166 apostrophe regression cannot return unnoticed.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#693)

* fix: make watcher supervision continuous

* no-mistakes(review): Bound watcher retries and log attached signals

* no-mistakes(review): Add bounded successor-recovery wake fallbacks

* no-mistakes(review): Prevent overlapping successor-arm retries

* no-mistakes(review): Resume supervision after late arm closes

* no-mistakes(review): Bind OpenCode recovery to attempted arm

* no-mistakes(test): Synchronize peer beacon regression fixture

* no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures

* no-mistakes(document): Captain: document watcher successor protocol behavior

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix: always fetch PR head for review diffs

Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded
pr_head= so reviewers never hold a merge over a "missing" fix that already
landed on the remote PR. Recorded SHA is offline fallback only; local branch
is last resort with a warning. Store the tip under refs/fm-review/ so a later
base-branch fetch cannot clobber the compare tip via FETCH_HEAD.

* no-mistakes(test): Isolate session-start nudge tests from gate state

* no-mistakes(document): Correct review-diff documentation
…and skills (kunchenguid#736)

* docs: resolve five contract contradictions

* no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): reverify grok exit command

* no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for exited paused crew

* no-mistakes(review): Gate pause suppression on confirmed agent death

* no-mistakes(test): Fixed stale pause cadence

* no-mistakes(document): Document dead-agent hold cadence
…id#744)

* fix(supervision): distinguish ordinary wakes from repair

* no-mistakes(review): Make passive guard follow-ups recovery-only

* no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes

* no-mistakes(review): fix x-poll claim error deduplication

* no-mistakes(review): separate claim diagnostics from relay recovery

* no-mistakes(document): Document X-mode once-only mention wakes
…enguid#747)

* feat(wake): enrich drained signal context

* no-mistakes(review): Bound wake enrichment reads

* no-mistakes(document): Document wake-drain annotations

* docs(wake): explain at-least-once drain boundary

* no-mistakes(review): Prevent symlink races in wake annotations

* no-mistakes(review): Exercise wake symlink race regression

* test: document intentional AFK marker subprocesses

* fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces

* no-mistakes(review): Harden Herdr projection creation and spawn serialization

* no-mistakes(review): Captain, disarm Herdr cleanup before launch submission

* no-mistakes(test): Correct stale Orca metadata failure fixture

* no-mistakes(document): Document Herdr presentation projection accurately
…nchenguid#775)

* fix(send): treat opencode busy-queued composer state as submitted

When fm-send sends a message to a BUSY opencode crewmate on the tmux
backend, opencode accepts the Enter and queues the message for the next
turn, but leaves the typed text visible in the composer row.  The
submit-verification loop sees a pending composer, exhausts retries, and
reports a false "Enter swallowed" failure while the message is actually
delivered.

Fix: after Enter retries are exhausted and the composer still shows
pending, check fm_pane_is_busy.  If the pane is busy (agent mid-turn,
footer shows "esc interrupt"), the harness queued the message, so
return "empty" (accepted).  On an idle pane, keep returning "pending"
(genuine swallow detection preserved).

Regression tests cover four scenarios:
- busy pane + pending composer -> empty (message queued)
- idle pane + pending composer -> pending (genuine swallow)
- busy pane + composer clears on first Enter -> empty
- idle pane + composer clears on first Enter -> empty (existing path)

* docs: document busy-queued Enter exception across backend docs and skills

Add explanatory comments and backend documentation for the
busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn
but keeps typed text in composer until the turn ends):

- bin/fm-tmux-lib.sh: document the busy-aware fallback in the
  file header and above fm_tmux_submit_enter_core
- .agents/skills/afk/SKILL.md: daemon-facing policy note
- .agents/skills/harness-adapters/SKILL.md: harness-specific fact
- docs/tmux-backend.md: submit-acknowledgement section with the
  busy-queue exception
- docs/herdr-backend.md: record the known gap
- docs/architecture.md: cross-reference in the daemon section

* test(tmux): fix SC2181 and make busy-submit test executable
…unchenguid#765)

* fix(spawn): require two stable reads before accepting worktree path

The treehouse-get worktree-detection loop in fm-spawn.sh accepted the
first pane_current_path read that differed from the project path, but
on some tmux/WSL setups a brand-new window transiently reports a
stale-but-real path before the pane actually settles into the
worktree. Since that stale path is itself a real, distinct git
checkout, it also passes validate_spawn_worktree's isolation check,
so the loop silently recorded the wrong worktree in state/<id>.meta
(and, for claude harness spawns, installed the turn-end hook there
too).

Require two consecutive polls to agree on the same non-project path
before accepting it, using the existing inter-poll sleep as the
confirmation gap so an already-settled pane isn't slowed down by an
extra cycle.

* fix(tests): drop unused CASE_DIR read in worktree-settle test

ShellCheck SC2034: CASE_DIR is split out of the case record but
never referenced; discard it with _ instead.

---------

Co-authored-by: Freudator86 <tim@allesknut.de>
…anges (kunchenguid#752)

* fix(watcher): stabilize Linux process identity

* no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state

Stop free-text tokens like "merged" from promoting nonterminal working: lines
to captain-relevant, so AFK no longer permanently suppresses idle recovery.
Defend wedge aging independently for nonterminal progress verbs, bind
no-mistakes current-state attribution to code identity (not branch alone),
and mark setup-complete as nonterminal in the ship brief scaffold.

* no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution

* no-mistakes(document): Document current-code-bound run attribution

* no-mistakes(test): Wait for stable Herdr shell readiness

* no-mistakes(test): Make Herdr and watcher readiness tests deterministic

* no-mistakes(test): Make tmux capture and watcher lifecycle deterministic

* no-mistakes(document): Document corrected supervision contracts
* fix: allow safe teardown during watcher recovery

* no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code

* no-mistakes(document): Sync continuity-gate docs to allow teardown recovery

* test: mark dynamic teardown fixture literal

* no-mistakes(document): docs: add teardown to continuity gate allow list
…nguid#790)

* feat(herdr): order presentation worker spaces

* fix(herdr): preserve focus during projected cleanup

* no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs

* no-mistakes(review): Serialize Herdr aborts with guarded focus regressions

* no-mistakes(review): Fall back flat when Herdr serialization is unavailable

* no-mistakes(test): Stabilize watcher startup and AFK handoff tests

* no-mistakes(document): Correct Herdr ordering and focus documentation
…#809)

* Send literal config reread after inherited config push

When declared inherited config changes under an already-running secondmate,
build a per-home instruction from validated destination post-write bytes and
deliver it on the routed secondmate path. Unchanged config sends nothing;
ABSENT represents removal; captain-shared is never inlined. Covers mid-session
config-push and the locked bootstrap convergence path without hardening spawn
against deliberate runtime choice.

* no-mistakes(review): Fix config reread framing, partial propagation, and respawn order

* no-mistakes(review): Send config rereads via durable single-line pointers

* no-mistakes(review): Make failed config rereads retryable

* no-mistakes(review): Make config reread retries generation-safe

* no-mistakes(review): Make config rereads durable and ordered

* no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode

* no-mistakes(review): Retain write retries and quarantine stale respawn generations

* no-mistakes(review): Preserve exact config reread retries and delivery order

* no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning

* no-mistakes(document): Consolidated config-reread documentation
* feat(watch): follow GitLab merge requests to merge

The merge watch only understood GitHub pull requests, so a task whose
deliverable is a GitLab merge request was never followed to merge.

Generalize the stored poll identity from owner/repository to a
provider-tagged provider/url/host/path/number record. GitLab runs mostly on
self-hosted instances and its projects nest under groups at no fixed depth,
so the host and the full project path are data in the record rather than
constants, and every consumer rebuilds the URL from those parts and refuses
any record that does not reconstruct it exactly.

The GitLab state is read with plain glab, matching the GitHub path's use of
plain gh, so an upstream checkout needs no extra tooling. Two things about
glab were established by running it rather than assumed, because a wrong
invocation here fails silently into a permanent "not merged":

- glab has no field selector, and its JSON would need a JSON processor that
  firstmate does not require, so the state is read from glab's own field
  output. Only an exact "merged" wakes firstmate, so a changed format stays
  silent instead of reporting a merge.
- glab cannot take a merge request URL the way gh can, because that form
  resolves through the current git repository and the watcher has none. It
  is addressed by project URL and merge request number instead.

An absent glab produces no wake rather than a false merge, and arming
refuses with a clear message since that is the one point where a missing
CLI can still be reported. A GitLab task records no pr_head, which both
consumers already treat as optional. The merge path still addresses GitHub
only and refuses a merge request URL rather than sending it to the wrong
forge.

The record version moves to v2, and the existing non-executing migration
rebuilds an already-armed watch from its recorded URL, so no watch is lost
by upgrading.

docs/gitlab-merge-watch.md records the evidence, taken against the public
fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture.

* no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation

* no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs
…uid#821)

* feat(herdr): correct all-home child presentation topology

Inherit the presentation opt-in to secondmate homes, label new projected
spaces with the approved corner format, insert each child under its owning
parent under one session-scoped lock, and keep flat non-destructive fallback.

* no-mistakes(review): Exclude secondmates from Herdr presentation projection

* no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering

* no-mistakes(review): Use adjacency-only Herdr child ownership

* no-mistakes(review): Reject foreign legacy projections safely

* no-mistakes(review): Validate Herdr session sockets before projection

* no-mistakes(test): Fix Herdr teardown fixture session socket metadata

* fix(herdr): canonicalize presentation lock socket paths

Always resolve the session socket parent directory so symlink parents
such as /tmp -> /private/tmp cannot split the shared cross-home lock
identity. Refuse relative socket paths. Clarify lock-unavailable warnings.

* no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing

* no-mistakes(document): Document all-home Herdr child topology

* no-mistakes(lint): Quote fallback provenance string for ShellCheck
* fix(no-mistakes): drop full-suite local Test override

Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad
tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a
focused contract test so the override cannot silently return.

* no-mistakes(lint): Make CI contract assertion ShellCheck-clean
* feat(test): add canonical timed suite runner and honest CI timeout

Introduce bin/fm-test-run.sh as the single serial owner for selecting
one script, a family, a conservative changed-file set, or the explicit
complete suite, with per-script timing markers and a JSON artifact.
Wire CI Behavior through the runner, raise the hang-tripwire timeout to
25 minutes, and document entry points without restoring a full-suite
local no-mistakes Test command.

* no-mistakes(review): Captain: fix changed selection and empty summaries

* no-mistakes(review): Captain: fail closed on unmapped changed sources

* no-mistakes(document): Document canonical timed test entry points
* fix: disclose main-home orphan and unstructured inventory gaps

Main Bearings could report an empty fleet while structured in-flight rows
lacked meta or current backlog rows were free-form. Emit main_inventory from
the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next
gate, and keep meta as the only live Underway source.

* no-mistakes(document): Document Bearings inventory-integrity projection

* no-mistakes: apply CI fixes
* feat: add concurrent test isolation proof for Phase 2

Prove an audited portable candidate set passes under concurrent
workers with private mode-0700 temp roots, without enabling
production CI sharding or fm-test-run --jobs.

* no-mistakes(review): Pin isolation proof to audited candidate manifest
* feat(secondmate): parent-owned guards for missed status reports

Marked parent-to-secondmate requests now create a durable pending-reply
expectation with a privacy-safe correlation id before delivery. Transport
success never resolves it; only a correlated parent status or document
pointer does. After a completed turn with no report, the parent sends one
recovery repost and escalates once if that turn is also missed, without
scraping the secondmate conversation or looping.

* no-mistakes(review): Deduplicate wrong-home pending-reply sightings

* no-mistakes(review): Harden pending-reply recovery and escalation guards

* no-mistakes(review): Bound pending-reply backend polling

* no-mistakes(review): Cache pending-reply status scans

* no-mistakes(review): Protect undelivered pending-reply records from scans

* no-mistakes(review): Close pending-reply delivery durability gaps

* no-mistakes(review): Separate pending-reply transport outcomes

* no-mistakes(review): Escalate stalled pending-reply deliveries once

* no-mistakes(review): Resolve attempted deliveries from correlated reports

* no-mistakes(review): Resolve late reports after delivery escalation

* no-mistakes(document): Document pending-reply grace and ownership

* no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings
* feat: add required pinned Herdr CI lane

Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and
SHA-256 pins, run the real-herdr-gated family serially through
fm-test-run with hard-fail on herdr-not-found, and keep portable
Behavior free of claimed Herdr coverage.

* no-mistakes(document): Consolidate real-Herdr CI documentation ownership

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…guid#841)

* feat: shard portable CI tests after isolation proof

Balance the Phase 2 proven-isolated set into two LPT portable parallel
lanes from Phase 1 timing evidence, keep stateful work in a required
portable serial lane, exclude real Herdr to its dedicated required lane,
and prove complete inventory coverage with a deterministic guard.
Add bounded local --jobs only for the proven set, per-lane timing plus
aggregate artifacts, and reduce the interim portable hang tripwire now
that the serial remainder owns the long wall-clock path.

* no-mistakes(review): Captain, fix CI contracts and completion-order worker scheduling

* no-mistakes(review): Captain, preserve stderr gate-skip detection in parallel tests

* no-mistakes(document): Document portable sharding and timing aggregation

* no-mistakes: apply CI fixes
)

* feat: fence primary-session delegation outside the fleet

A firstmate primary that delegates through Claude Code's built-in
delegation tools creates work with no state/<id>.meta. Because
fm-supervision-lib.sh counts *.meta and fm-turnend-guard.sh exits
silently at zero, such work does not merely go unsupervised: it makes
the whole guard stack structurally inert, and it dies with the primary
session. On 2026-07-22 that cost two workers mid-flight and left
supervision down for 73 minutes unnoticed.

Layer 1, the primary fix: a permissions.deny list in
.claude/settings.json removes the 18 delegation, scheduling, worktree,
and task-tracking tools from the model's schema, so they are never
offered. This is removal rather than interception, so there is no call
to intercept and no fail-open path. The list is flat and in one file so
its width stays reviewable; the captain owns that width.

Layer 2, bin/fm-subagent-pretool-check.sh: a deny list is fail-open
against tools that do not exist yet, and permissions.allow is a
pre-approval list rather than an availability list, so there is no
fail-closed allowlist to use instead. This backstop classifies the tool
NAME by shape rather than against a fixed list, so a delegation tool
that ships before the deny list is updated is still refused. It excludes
mcp__* names and observe-or-stop operations, scopes itself to a genuine
primary home via the shared fm_primary_scope_matches predicate so a
crewmate's task worktree is unaffected, and offers one deliberate
FM_ALLOW_SUBAGENT=1 escape hatch that must be set at launch.

Verified live against Claude Code 2.1.217, including a deny-key A/B with
a nonsense-name control, layer 2 denying an un-denied Workflow call, the
same call allowed in a linked worktree, and the escape hatch. Corrects a
prior finding: both Task and Agent work as deny keys, so both are
pinned. Codex 0.144.1 verified to expose no delegation tool; grok,
opencode, and pi are inspected and documented as not wired because those
binaries are absent from this host and the repo requires live validation
before trusting a harness hook. Evidence in docs/subagent-guard.md.

* no-mistakes(review): Ship scoped Claude delegation guard

* no-mistakes(test): Ship Claude delegation deny list

* no-mistakes(document): Clarify PreToolUse guard ownership

* no-mistakes(lint): Keep Claude deny list local
Reproduction: portable-parallel-2 completed successfully without tasks-axi while fm-decision-hold-lifecycle emitted a gate skip in 30 ms. The pre-shard lane installed tasks-axi and exercised the test fully. Installing tasks-axi is the smallest counterfactual and makes the representative shard execute the test with gate_skip=false in about 20 seconds. Both parallel jobs receive symmetric setup, while the exact 91-test inventory and coverage guard remain unchanged.
* feat: make dispatch profiles quota aware

* no-mistakes(review): Fix quota window and Grok product scoping

* no-mistakes(document): Document implicit quota-aware dispatch accurately
Inthuson and others added 29 commits August 21, 2026 11:21
…lared pause (kunchenguid#2748)

* fix(bin): give a captain hold the same bounded pause cadence as a declared pause

Two supervisors read a finished task's last status line and disagreed about which
declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh
suppresses its inactive-outcome scan only on `captain-held`, while the away-mode
daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so
the two verbs are mutually exclusive and no finished task waiting on a person
could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the
900s outcome scan and immediately produced five possible-wedge escalations in one
batch, because the 240s wedge detector no longer saw a pause verb.

fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined
question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This
extends that same answer to the paths still asking the narrower one:

- bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to
  move together. classify_stale returns the pause action, reconcile_pause_tracking
  and migrate_watcher_pause_markers record and migrate the marker, and
  housekeeping defers the wedge and then re-surfaces the recheck. Changing only
  the stale-persistence gate would defer the escalation while
  reconcile_pause_tracking recorded nothing, so the wedge marker would persist and
  the sweep would `continue` past it forever: quiet, but never re-surfacing.
- bin/fm-watch.sh's secondmate stale gate, whose downstream owner
  pause_state_class already treats both declarations identically.
- bin/fm-push-transition-lib.sh's absorb, where either declaration already names
  the human the transition would report and the wait is already durably recorded.

Quieting alone would be half a fix, so the bounded re-surface had to reach a hold
too. A hold has no current-state mapping, unlike `paused`, so authoritative crew
state reports it as unknown and pause_state_class received `none`. An ordinary
crew recovers pause classification from that state through confirmed agent death,
which proves no live decision gate is being silenced. A secondmate's endpoint
liveness is deliberately never read there, because an idle mate is healthy by
design, so that confirmation is unavailable by construction and cannot be
required: without recovering the classification for a mate, every caller silenced
a held mate outright and its hold would rot invisibly. That promotion is bounded
by the declared-wait guard at the top of the function, so it can only reclassify a
task that already declared a wait and shows no positive working evidence.

Two narrow `status_is_paused` calls are deliberately left alone.
bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a
wedge path; reporting a hold as `paused` would erase the distinction
status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held`
close is a verified durable transfer and a `resolved` close claims outright
settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no
change because that function's own case list already returns non-relevant for
`captain-held`.

bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its
guard exists because a finished task's crew state still reports done from a
higher-priority source than the log, and a declared pause needs no such guard: the
scan only reports done or failed, and nothing else reaches its record path.
Widening it would change a separate subsystem's reporting contract, which this
defect does not require.

Coverage extends the existing colocated patterns for these predicates and asserts
both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker
converting to pause tracking with no escalation, the bounded re-surface with its
window reset, and the boundary case where an answered hold stops claiming the
cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on
the same bounded cadence without being labeled a wedge.
tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one
of these fails on the pre-fix code except the answered-hold boundary case, which
is there to pin that the quieting was not widened too far.

The `paused:` workaround appended to those 11 tasks is live supervision state and
is untouched here. It can be retired once this lands.

* no-mistakes(review): name the captain in a held task's bounded recheck

* no-mistakes(document): extend declared-wait supervision docs to captain-held holds
…guid#2758)

* fix(lint): name the installer when ShellCheck or actionlint is missing

A missing actionlint exited 127 like a bare command-not-found. Fail with
exit 1 and point at the pinned installer, matching the missing-ShellCheck
path, without weakening the version pin.

* test: isolate kimi and muse detection from inherited Cursor markers

Harness detection checks CURSOR_AGENT before ancestry, so these
markerless-adapter cases failed when the suite itself ran under Cursor.
Clear the verified markers the same way the secondmate harness tests already do.

* no-mistakes(document): Document Muse Cursor marker cleanup
…lled but inert (kunchenguid#2684)

* feat(checks): report tool updates that are available or installed but inert

Firstmate had no way to notice that tooling this home depends on needs an
update, and no way at all to notice the worse case: an update that installed
correctly and then did nothing.

That second case is why this exists. A tool that self-installs into
~/.local/bin while a version manager keeps its own older copy earlier on PATH
looks completely up to date to anything that asks only "is a newer version
published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0
copy stayed earlier on PATH, so every Herdr command failed on a protocol
mismatch and firstmate could not read its own fleet.

bin/fm-tool-update-check.sh reports the two conditions separately:

  <tool> update available      a newer version exists at the update source.
  <tool> update not in effect  a newer copy is installed on this host, but
                               PATH still resolves an older one.

PATH skew is measured, never inferred. Every executable copy of a watched
command on PATH is asked for its own version and those answers are compared,
so one lookup cannot hide the skew, and a directory name is never read as a
version because a version manager's "latest" directory can hold an older
build. A copy that will not report a version is a check failure, not a pass.

The watched tools live in local, gitignored config/watched-tools.json, so
adding a tool is a config edit rather than a code change, and the file is
never propagated to another home. Update sources cover both shapes: a local
clone's commit distance from its remote branch, and a command's own version
and update announcement, including a tool like no-mistakes that prints its
version on one command and announces a new release on another.

The check prints one line when something needs attention and prints nothing
otherwise, so it rides the existing watcher state-check contract with its
trust binding instead of introducing a schedule of its own, and
state/.tool-updates keeps the same pending update from being reported on
every poll.

The check only reports. It never installs, updates, reorders PATH, touches a
version manager, or fetches into a watched repository; every git probe is
read-only.

Tests cover the skew case as a regression, and it was verified by mutation:
removing the skew report, or stopping after the first PATH hit as a single
lookup would, each make that test fail.

* no-mistakes(review): fix tool update check probe reporting, budget, and shim write

* no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets

* no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe

* no-mistakes(review): guard git probes at the budget, record uncut findings

* no-mistakes(document): fix stale watched-tool report-record wording in docs and header

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

The behavior shard's watch-triage suite failed on the new worktree-write wedge
tests. Those five tests are the only ones in the file that do not use its
standard waits. They give a fixed 3 second liveness budget to the one poll that
now spawns the bounded worktree walk, and 4 seconds to an escalating watcher
where every other test in the file gives 10. On a loaded runner that poll
outlives the fixed budget, so the round is reaped before the deferral it asserts
on is recorded, and the test reports a lost deferral instead of the deferral
under test. Wait for a completed poll cycle through the file's own
wait_poll_cycle, which is what its header documents this hazard for, and use the
file's standard 100 tick exit budget.

Verified against a load that reproduces the failure: 11 of 12 runs failed
before, 8 of 8 pass after. Verified by mutation too, so the waits still prove
the behavior: removing the write deferral, and keeping a finished deferral chain
across an idle-timer repair, each still fail their test.
* fix: treat yolo as merge authority only, not ask-user finding authority

Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation.

* no-mistakes(document): Clarify yolo documentation ownership and merge posture
… work over (kunchenguid#2767)

* feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost

Step one of the spoken interface: the laptop captures and plays audio, this
desktop holds the model session, and no AWS credential leaves the desktop.

Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte
of reply audio, 6 runs each, all answered, on a question that forces a records
read:

  relay path   1.229 1.379 1.428 1.447 1.481 1.516  median 1.438
  direct       1.147 1.179 1.203 1.237 1.244 1.317  median 1.220

The relay costs about 0.22s of the median. The direct figure reproduces the
earlier survey, which is what makes it a usable control. Excluded: the
captain's own ssh round trip, microphone capture, and speaker output. This
desktop has no microphone and no speaker, so every run used audio files.

Three pieces:

  bin/fm-voice-relay.py    holds the conversation on this host
  bin/fm_voice_records.py  what a spoken answer may read, and the handover
  bin/fm-voice-client.py   the laptop end; audio devices UNVERIFIED
  bin/fm_voice_frame.py    the wire format both machines share

Real work is handed to the existing bin/fm-inbox.sh rather than a second
queueing surface, and the agent says it is handing over rather than answering
as firstmate.

Read scope: Done history and free-form note bodies are never assembled at any
scope, so the wide default cannot reach the places commercial detail
accumulates. config/voice-read-scope narrows it to counts only, and
config/voice-read-deny excludes a named item in one line. The boundary is an
executable test that widening the reader fails.

Push to talk is the default because it is cheaper and the choice is still open;
--listen open-mic is the single flip.

Two traps worth knowing: a clip with no trailing silence is never answered, and
the end of a reply is contentEnd with stopReason END_TURN, not completionEnd.
A second user turn in one session is treated as barge-in unconditionally, and
an interrupted turn that calls a tool is lost, so the session reconnects per
turn and gives up conversational memory. That is the concrete thing step three
has to solve.

* no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing

* no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir

* no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials

* no-mistakes(review): read account config from home, harden deny and turn failures

* no-mistakes(review): close status verb set, fix inbox help, pair data override

* no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion

* no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline

* no-mistakes(review): survive reader failures, release devices, fix claims

A failure while handling a model event, or while sending a tool result,
left the reader task dead with ended and turn_done clear, and close()
re-raised the stored failure on every await. One dropped stream became a
relay that could never build another session. The reader now reports the
session over in a finally whatever killed it, and close() absorbs the
task the same way it already absorbed its sends.

The laptop client releases what it already started when a later startup
step refuses, SystemExit from the handshake wait included, and names a
device refusal instead of leaking a raw PortAudio error. Whether it
releases correctly against a real device is still unverified here.

The records docstring claimed every reading was filtered to open ids.
Only the pull request count and list are; the worker count and the state
histogram cover every live runtime record, finished ids included,
because a meta file still on disk still needs tearing down.

The finished-work deny half of the suite asserted things that held with
the deny list absent. It is replaced by a deny on an open title, which
removes the row and says so while the count stays honest.

* no-mistakes(review): name reader failures, split file and device refusals

A failure inside the model reader released the waiting turn and told
nobody. The session was not marked spent, no notice reached the client,
and the client waits for a reply end or a notice, so the captain got
their whole timeout of silence and then a record saying the turn went
unanswered with nothing about why. Both ends of the relay now name a
failed turn through one function, once per turn, and --self-test carries
the cause in relay_error the way the client's own record does.

Two things that are not failures stay that way. A stream that simply
ends is the end of a session, which serve still reads on its own terms.
A stream that goes away because close() asked it to is an ordinary
renew, and announcing it would have put a failure notice in front of the
captain on every turn.

On the laptop end, the refusal that became a device error covered the
file-backed playback and capture too, so a mistyped --in-file was
reported as an audio device failure and the advice named the flag that
had just failed. The file ends now report the path and the flag that
chose it and stay an OSError; the device ends keep the device advice and
name the flag for that end. The device paths remain unrun here, so only
the file halves are covered by a test.

* no-mistakes(test): survive model session end, order client turn frames

* no-mistakes(document): sync voice relay docs with reviewed relay behavior

* no-mistakes(document): re-measure relay latency and correct its cause

* no-mistakes(document): correct measurement date and name the unmeasured SSH hop

* no-mistakes(document): describe the unpublished control measurement, fix list formatting

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
…unchenguid#2763)

* fix: keep Relay public loops open until retire

Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close.

* no-mistakes(review): Propagate public follow-up registration removal failures

* no-mistakes(review): Persist retire receipts and align parent resolution

* no-mistakes(review): Make rechain resumable after partial obligation creation

* no-mistakes(review): Repair follow-up state, briefs, and expiry escalation

* no-mistakes(review): Serialize follow-up delivery stamps with retirement

* no-mistakes(review): Serialize rechain claims and protect registration terminal states

* no-mistakes(review): Avoid reporting retired delivery loops as open

* no-mistakes(document): Refresh public-loop documentation and verification evidence

* no-mistakes: apply CI fixes

* no-mistakes(review): Preserve delivered follow-up bindings during registration replay

* no-mistakes(review): Harden public follow-up retirement and rechain races

* no-mistakes(review): Fail closed on unresolved secondmate retirement

* no-mistakes(review): Bind secondmate cleanup to its recorded canonical home

* no-mistakes(review): Fix rechain command output and expiry validation

* no-mistakes(review): Validate brief keys and warn on remote promotion

* no-mistakes(document): Document retained public follow-up loops

* no-mistakes(lint): Remove unused bounded-wait loop variable
…ath (kunchenguid#2779)

* feat(bin): merge GitLab merge requests through the guarded PR merge path

bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher,
but bin/fm-pr-merge.sh refused every non-github provider, so a merge request
had to be merged by hand and got none of the recording, guards, or audit
trail a pull request gets.

The merge path now dispatches on the parsed provider. A GitHub URL keeps its
exact previous behavior. A GitLab URL is addressed through glab by the project
URL rebuilt from the parsed host and path, so a merge request on any instance
resolves and no host is hardcoded, and no merge-method flag is added because
the project's own merge method is what should apply.

A GitLab merge happens only after one live read of the merge request confirms
it is open, detailed_merge_status is mergeable, has_conflicts is false,
blocking_discussions_resolved is true, and the head pipeline succeeded at the
exact current head. Every failing condition is reported, not just the first.
The verified head is bound to the merge with glab's --sha, so a push landing
between the read and the merge fails the merge instead of landing commits
nothing verified. Recorded metadata is never the authority for any of this: a
rebase moves the head and leaves a recorded value stale, so a recorded head
that disagrees with the live one is reported rather than trusted, and the
recorded value is read before the recording step because that step drops a
GitLab head it cannot resolve.

* no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent

* no-mistakes(test): state authorised GitHub narrowing of bundled -R guard

This branch NARROWS GitHub behaviour. The narrowing was authorised
deliberately rather than slipping in by accident, and it applies to both
providers, GitHub and GitLab alike, because a script that guards one provider
and not the other is a trap for the next reader.

What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a
bundled short-option cluster that includes R, for example "-dR other/repo".
The forge CLIs expand such a cluster one character at a time, so it carries
"--repo other/repo", and that later value wins over the repository the URL
named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR
other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR
other/repo" and exited 0 with pr= recorded and the merge poll armed. It now
exits 1 with "extra merge arguments must not override the repository", records
nothing, and invokes no forge merge command. Every other GitHub invocation is
byte-identical to the base commit.

Closing that hole honours the existing rule rather than departing from it. The
file header already forbids --repo and -R because the repository must come
only from the URL, so a bundled cluster carrying a repository override was
never legitimate behaviour to preserve: it was that guard being evaded.
Redirecting a merge to a repository the URL does not name is exactly what the
guard exists to prevent.

The refusal is already pinned on both paths by the existing case
test_bundled_repo_override_args_refuse_before_recording in
tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR
https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in
the task meta, no armed merge poll, and no forge merge command invoked, with a
control case proving a cluster that carries no repository override still
reaches the forge. No duplicate assertion was added. Both assertions were
confirmed to have teeth by narrowing the guard back to a bare -R and watching
each path fail.

This commit carries no file change: the guard and its coverage landed in
614853d, and this message exists so the pull request description states the
narrowing.

* no-mistakes(document): fix README pointer for GitLab watch and merge doc

* no-mistakes: apply CI fixes
kunchenguid#2788)

* no-mistakes: apply CI fixes

* fix(bin): drop a private record citation and narrow the review rule

Three corrections to the spoken interface that landed in kunchenguid#2767, plus one
fix carried over from that branch after its pull request had already been
merged.

The confidentiality fix. The module docstring of bin/fm-voice-relay.py
cited a private, gitignored fleet record by exact path and section number.
That widens what this public repository points at, and it cannot resolve
for any reader here, because the path has never been in the repository.
Both traps it pointed at are already described in full in the list
immediately below it, and docs/voice-relay.md carries the same two for
operators with no citation at all, so the pointer is removed and no claim
is weakened by losing it. Two comments that referred to "the survey" as
though it were something a reader could open are reworded the same way.
Neither exposed a path, so that half is comprehensibility rather than
confidentiality.

The review rule. .greptile/rules.md is kept, because its conditions are
right and deleting it would leave the next reviewer to re-litigate a
decision already argued out. What was wrong with it is narrower than its
existence: it read as settled repository policy, when whether VISION.md
itself should be reconciled is an open question belonging to the captain.
One sentence now says so, and says that the conditions listed below it are
what the interpretation depends on. That narrows the claim rather than
widening it.

The carried-over fix. The first commit on this branch is 7f98e79 from
fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes
the window where a transport failure was recorded and then erased, so a
run could be emitted as answered false with relay_error null. That matters
more than it looks: relay_error is the field that keeps an infrastructure
failure from being averaged into a latency figure, so the failure mode is
a dead connection wearing the costume of a slow reply. It landed fifteen
minutes after kunchenguid#2767 merged and so never reached the default branch.

* no-mistakes(review): name a reason on every unanswered-turn close path

* no-mistakes(review): guard the downlink body and pin frames to their turn

* no-mistakes(review): attribute reply audio to its own turn and tell endings apart

* no-mistakes(review): tell a cut-short reply from an unanswered turn

* no-mistakes(review): discard reply audio arriving after the output closes

* no-mistakes(review): count discarded reply audio on the speaker path too

* no-mistakes(review): keep a reason off a turn already answered in full

* no-mistakes(review): say a reset cut a reply short, not that none arrived

* no-mistakes(review): read one turn's audio count once, and hush a tidy exit

* no-mistakes(document): fix stale session-end relay_error claim in voice-relay guide
…repo

Moves the previously untracked state/fm-tg-* scripts into bin/, using the
standard FM_ROOT/FM_HOME/STATE resolution pattern so a configured clone just
works instead of the scripts only running against one hardcoded machine path.

Structural fix: registers the two Stop hooks (fm-tg-guard.sh, fm-tg-hook.sh)
in this repository's own tracked .claude/settings.json, project-scoped,
instead of the user's global settings. That was the root cause of crewmates
draining and answering the captain's Telegram inbox themselves.
fm-tg-isfirstmate.sh stays as defense in depth for a crewmate working on
firstmate's own repo.

Extracts the poller/waiter's duplicated, single-quote-fragile inline Python
into one shared bin/fm-tg-fetch.py, so both fetch paths ack on arrival and
cannot drift apart again. Carries forward every live fix made to the
untracked scripts during this task: arrival-time acknowledgement, the
poller/waiter race fix, the reply/surface race fix in fm-tg-archive.py
(retire on 60s-old-and-unsurfaced, not just surfaced), and the upload fixes
(sendDocument for >1MB images, a payload-scaled upload timeout with an
explicit timeout message instead of 'unparseable reply').

Adds bootstrap integration mirroring the X-mode pattern: a presence-gated
TELEGRAM: report line, a generated state/tg-watch.check.sh poll shim, and a
generated config/tg-mode.env 30s watcher cadence file. Absent or partial
config stays completely silent and inert.

Adds docs/telegram.md (setup, architecture, every guarantee and the defect
that produced it) and tests/fm-telegram.test.sh (hermetic, no live Telegram
traffic - a fake curl stands in for the network so fm-tg-send.sh and
fm-tg-poll.sh still run for real). Drops fm-tg-digest.sh: it was never wired
into the watcher or anything else, so it shipped no live behavior.

Also fixes tests/fm-bootstrap.test.sh to isolate FM_TG_ENV_OVERRIDE, since
Telegram config lives at a fixed $HOME path (deliberately, one bot per
machine) rather than inside each test's scratch FM_HOME.
Amendment 5 (live failures after amendment 4):
- fm-tg-archive.py: the 60s archive window was an invented number and still
  let a ~50s-old message survive a reply and re-surface (third recurrence of
  the duplicate-reply bug). The correct rule is that a reply covers
  everything already in the inbox when it was sent; only a message under 10s
  old (genuinely unread) stays pending.
- fm-tg-send.sh: sendMessage parts now retry up to 3 times with a 2s backoff
  on a transient failure, and never retry a genuine Telegram rejection
  (ok:false), since retrying a bad request cannot help.

Also fixes a real bug found while validating this branch: fm-tg-hook.sh
created state/tg-inbox and state/tg-processed unconditionally, before ever
checking whether Telegram was configured at all - a state mutation on every
Stop hook firing even with zero Telegram config, violating the 'absent
config means completely inert' guarantee. The config-absent test only
passed before because it ran through the real bin/fm-tg-isfirstmate.sh from
inside a live crewmate session, whose ancestry check short-circuits the hook
before the mkdir is ever reached - masking the bug in exactly the dev loop
used to write it. Fixed the hook to check config before creating anything,
and reworked the affected tests to use the identity-stubbed scratch bin
(already used elsewhere in this suite) so the config-absent path is actually
exercised rather than skipped past by ambient ancestry.

Verified the fix catches the regression: reverting fm-tg-hook.sh's check
reproduces the test failure; restoring it passes again.
…ions

A no-mistakes review of this branch flagged two ask-user findings; the
captain decided both directly (both were his own design calls, not the
reviewer's or the crewmate's):

1. tg-no-chat-filter (security): inbound Telegram updates were never
   filtered by TG_CHAT_ID, so anyone who found the bot's public username
   could message it and be recorded/surfaced/acknowledged as the captain,
   then force a blocking turn-end demanding a reply that lands in the real
   captain's chat. bin/fm-tg-fetch.py now drops any update whose chat id
   does not match TG_CHAT_ID before it is ever recorded, with a visible
   (never silent) stderr line naming both chat ids so a legitimately
   wrong-chat message from the captain himself stays diagnosable.

2. tg-archive-unsurfaced-loss (data loss): the arrival-time retirement
   window in fm-tg-archive.py applied to ANY non-ack send, not just an
   actual reply to what was pending, so a proactive/unrelated message from
   firstmate could retire an old message that had never actually been shown
   to the model - a direct 'no message is lost' violation the captain had
   not seen yet. Also, a missing/malformed 'ts' made a record look
   infinitely old (float(None or 0) == 0.0), retiring it on the very first
   send, sight unseen.

   Fixed per the captain's design: a record that has been surfaced at least
   once (bin/fm-tg-drain.py now stamps 'surfaced_ts' on first surfacing)
   always retires on any real reply, however much later - keyed on
   first-surfaced time, not arrival time. A record that has never been
   surfaced retires only within a short, conservative 10s window of
   arrival - the genuine same-turn race the window exists for - and stays
   pending outside that window no matter what unrelated reply goes out. A
   missing/malformed ts is treated as brand new, never as infinitely old.

   This flips the polarity of the two tests that encoded the earlier
   (now-superseded) 'anything older than 10s retires' rule; rewrote them
   along with the rest of the affected suite to match the corrected
   semantics, and added dedicated boundary, surfaced-always-retires,
   missing-ts, and chat-id-filter tests.
… changed

test_unsurfaced_retirement_is_recorded used a 120s-old never-surfaced record
and expected it to retire with a notice. That was the semantics before the
captain's ask-user decision (archive-window fix): a never-surfaced record
over 10s old now stays pending instead of retiring, so this test needed a
fresh (<10s) record to still exercise the same-turn-race retirement path it
is actually testing.
…s decision

The auto-retry from amendment 5 can rarely deliver a duplicate message: an
empty sendMessage reply from a curl that itself exited 0 is ambiguous
(Telegram may have already accepted and answered the send, just too slowly
for --max-time to see the reply), and retrying that case can resend an
already-delivered message. This was flagged as an ask-user finding by
no-mistakes since it is a genuine tradeoff of the captain's own explicit
retry request, with no clean fix available (Telegram's Bot API has no
idempotency key for sendMessage).

The captain's decision: accept the tradeoff as-is. A rare duplicate costs
him nothing to notice and ignore; a dropped message costs him the answer,
and he has already stated that preference in fm-tg-guard.sh's own wording.
Retry stays, unconditionally, for both the definite-non-send and the
ambiguous case.

Per his request, bin/fm-tg-send.sh now distinguishes the two cases in its
retry log line (a nonzero curl exit is a definite non-send; an empty reply
from a curl that itself exited 0 is the ambiguous, may-have-already-sent
case) so the rare case is diagnosable after the fact, and docs/telegram.md
names the tradeoff plainly instead of leaving it implicit. This is
observability only - the actual retry behavior is unchanged for both
cases, per the decision to keep retry unconditional.
The no-mistakes pipeline surfaced these 5 findings, all auto-fix, and began
applying them, but the fix round appeared to stall (no live process, no
filesystem activity for 5+ minutes despite a complete-looking log summary)
and had to be aborted before it committed - a recurring pattern on larger
fix rounds this session. Implementing them directly rather than losing them
a second time:

1. telegram_setup() had no secondmate-home gate. The Telegram config is
   deliberately machine-wide (~/.config, unlike X mode's FM_HOME-scoped
   .env), so a live secondmate home would arm its own competing poller
   against the same bot token - and since getUpdates offsets are
   server-side, whichever home polls first silently steals the update.
   Gated on the primary, matching startup_memory_budget_setup's precedent.

2. fm-tg-fetch.py let a write_record() failure "continue" to the next
   update, which then advanced the offset past the failed one - permanently
   losing it, since Telegram never re-sends an already-acknowledged offset.
   Changed to "break": the whole batch from the failure point on is left
   unacknowledged and retried next poll. A batch that surfaced nothing due
   to a write failure is now reported as a refusal instead of reading as a
   quiet channel.

3. fm_dir_is_child_worktree (bin/fm-primary-scope-lib.sh, shared beyond
   Telegram) compared --git-dir and --git-common-dir as raw strings, but
   git returns them in different path formats from a subdirectory of a
   plain checkout (absolute vs relative) - misclassifying any such
   subdirectory as a crew task worktree. Added fm__git_dir_abs to
   canonicalize both sides via cd+pwd -P before comparing, applied to both
   fm_dir_is_child_worktree and fm_primary_scope_matches.

4. send_part() decided "rejected, not transient" via a raw substring match
   on the reply body, missing any whitespace-formatted variant of the same
   JSON. check() now exits 2 for a parsed-but-not-ok rejection (1 stays for
   unparseable/empty), and send_part branches on that exit code instead.

5. Reworded the missing/malformed-ts comment (in three places: the
   archive.py docstring, its inline comment, and docs/telegram.md) from
   "treated as brand new" to "unknown, never retired by the window" -
   the old wording was self-contradictory, since a genuinely brand-new
   record IS eligible for the window and an unknown-ts one never is.

Re-verified: 36/36 telegram tests pass, shellcheck clean, and the
fm-turnend-guard and fm-cursor-primary suites (both depend on the shared
fm-primary-scope-lib.sh predicate) pass with no regressions.
…QX4MG1J

The pipeline stalled again mid fix-round, so these six findings from that
run's review gate are applied by hand instead of losing another round.

- wait-path-never-marks-surfaced: bin/fm-tg-wait.sh's own fetch success path
  printed the fetched text and exited without ever running the drain, so a
  message delivered through that path (not the drain-first branch at the top
  of the loop) was left at surfaced=0 with .tg-last-surfaced unset. A reply
  composed more than 10s later then found the record still pending under
  fm-tg-archive.py's never-surfaced window and it re-surfaced forever. Run
  the drain silently first, matching the bookkeeping the drain-first path
  already gave every other delivery.

- curl-timeout-mislabelled-definite-non-send: bin/fm-tg-send.sh classified
  every nonzero curl exit as a definite non-send. A timeout-shaped exit
  (curl's own --max-time exit 28, or 124/143 from this script's outer
  fm_run_timed bound) means Telegram may have already accepted and answered
  the send just too slowly to see the reply - the same ambiguous,
  may-have-already-sent case as an empty reply from a curl that exited 0,
  not a definite non-send. Both still retry the same way; only the logged
  reason changes. docs/telegram.md's retry-tradeoff section is corrected to
  match.

- drain-crashes-on-non-dict-record: bin/fm-tg-drain.py appended
  rec.get("ts") outside the try/except wrapping json.load, so a record that
  parses as JSON but is not a dict (e.g. an interrupted write that left a
  bare number or array) crashed the whole drain instead of being skipped,
  losing every pending message in the inbox, not just the corrupt one.

- archive-main-loop-bypasses-load-helper: bin/fm-tg-archive.py's retirement
  loop called raw json.load() instead of the file's own non-dict-tolerant
  load() helper that prune() already relies on. The same non-dict record
  crashed retirement and pruning, invisibly, since the script's only caller
  discards stderr and runs it with "|| true".

- test-suites-not-isolated-from-machine-wide-tg-config: ~20 test suites
  besides tests/fm-bootstrap.test.sh call fm-bootstrap.sh or
  fm-session-start.sh without isolating Telegram's fixed-$HOME config path,
  so a machine with real Telegram captain-comms configured (such as the one
  this repo's own fleet runs on) leaks an unexpected TELEGRAM: line and a
  live cadence-relevant side effect into their output. Default
  FM_TG_ENV_OVERRIDE to a nonexistent path in tests/lib.sh so every sourcing
  suite is isolated without hand-editing each one; a suite that actually
  exercises Telegram config still overrides it explicitly afterward. The one
  suite that does not source tests/lib.sh (the opt-in credentialed Claude
  Stop auto-arm live E2E) gets the same override directly.

- contradictory-cadence-lines: when Telegram captain-comms was active and X
  mode was not, the session-start supervision block printed "X mode:
  inactive; use the default watcher cadence" immediately above "Telegram
  captain-comms: active; source ... 30s cadence is inherited" - two adjacent
  lines giving opposite cadence instructions for the same watcher launch.
  The X mode line now reports its own inactive status without a cadence
  claim when Telegram is carrying that instruction instead.

Added regression tests for all six: the waiter's own fetch path marking a
record surfaced, a curl timeout (exit 28) logged as ambiguous rather than a
definite non-send, a non-dict record surviving both the drain and the
archive pass without losing a good record alongside it, and the
supervision-instructions renderer no longer emitting the contradictory pair
when Telegram is on and X mode is off.

Verified locally: bash tests/fm-telegram.test.sh (40/40), bash
tests/fm-supervision-instructions.test.sh (10/10), shellcheck bin/*.sh
bin/backends/*.sh tests/*.sh (clean), plus spot checks of
fm-bootstrap.test.sh, fm-session-start.test.sh, fm-fleet-sync.test.sh,
fm-x-mode.test.sh, fm-secondmate-sync.test.sh, fm-secondmate-liveness.test.sh,
and fm-shared-captain-inheritance.test.sh against the tests/lib.sh default
change. fm-bootstrap.test.sh and fm-session-start.test.sh each have one
pre-existing unrelated failure, confirmed present against the unmodified
tests/lib.sh from HEAD as well, so neither is a regression from this change.
- tg-cadence-missing-opencode-pi (auto-fix): the OpenCode plugin and Pi
  extension's arm-child spawn only sourced config/x-mode.env before
  starting bin/fm-watch-arm.sh, missing the config/tg-mode.env line the
  Claude and Cursor arm paths already had. A Telegram-configured OpenCode
  or Pi primary polled at the 300s default instead of tg-mode.env's 30s.
  Fixed both spawn lines to match. Testing this exposed a second, more
  severe instance of the same parity gap: the OpenCode plugin's shouldArm()
  only treats x-mode.env presence as a reason to arm the watcher on an idle
  home with no fleet work; without the matching tg-mode.env check, a
  Telegram-configured idle OpenCode primary never armed the watcher at all,
  making the cadence fix moot. Added tg-mode.env there too. The Pi
  extension has no equivalent idle-arm gate (its arm is tool-driven), so it
  needed only the cadence-file fix.

- tg-poll-path-no-surfaced-stamp (ask-user, decided - same bug class as
  wait-path-never-marks-surfaced from the previous fix round, not a new
  judgment call): bin/fm-tg-poll.sh had the identical omission
  bin/fm-tg-wait.sh's own fetch path had before its last fix - it prints
  fm-tg-fetch.py's summary directly and never runs the drain, so a message
  the watcher poll wins the getUpdates race for is recorded but never
  marked surfaced. Applied the identical fix: run fm-tg-drain.py silently
  after a genuinely usable poll (rc==0).

- tg-poll-error-cleared-on-hard-failure (auto-fix): the standing refusal
  record at state/.tg-poll-error was cleared on ANY non-3 exit from
  fm-tg-fetch.py, conflating a genuine success (rc==0) with an unexpected
  failure (python3 vanishing from PATH -> 127, the watcher killing the
  check mid-fetch -> 137/143). An unexpected failure now reports and
  remembers itself the same dedup'd way a channel refusal does, and only a
  confirmed rc==0 poll clears the record.

- tg-test-unsets-isolation-default (auto-fix): three tests in
  tests/fm-telegram.test.sh reset state with a bare
  `unset FM_HOME FM_TG_ENV_OVERRIDE`, discarding tests/lib.sh's Telegram
  isolation default for every test defined after them in the file, not
  just the one resetting - reopening exactly the machine-config leak the
  previous fix round's tests/lib.sh default was meant to close. Each site
  now restores the captured default instead of dropping it. Added a final
  invariant test, run last in the file's own invocation list, that catches
  a regression in any earlier test's reset rather than just these three.

tg-stop-hooks-claude-only (ask-user) is NOT decided here - it is a genuine
scope question (build Stop-hook/extension-hook coverage for five other
harnesses, a multi-day expansion, versus documenting today's Claude-only
scope as a known limitation) rather than a defect against already-accepted
behavior, so it is escalated per the ask-user-authority framework:
data/fm-telegram-ship-t8/needs-decision-round6-stop-hooks-scope.md.

Added regression tests for all four decided fixes: cadence-file sourcing
and idle-arm parity for both the OpenCode plugin and the Pi extension, the
poll path's own surfaced/last-surfaced stamping (plus updating an existing
test whose "surfaced": 1 assertion is now correctly "surfaced": 2 once the
poll's own fix already stamps it once), the poll's unexpected-exit record
preservation, and the isolation-default survival invariant.

Verified locally: bash tests/fm-telegram.test.sh (43/43), bash
tests/fm-pi-watch-extension.test.sh (33/33, covers both the Pi extension
and the OpenCode plugin), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh
(clean), node --check on both edited plugin/extension files, python3
-m py_compile on the touched Python modules.
The captain's decision on tg-stop-hooks-claude-only: decline building
Stop-hook/extension equivalents for five other harnesses (a multi-day
expansion, not this task's scope), but the current unqualified behavior is
worse than unsupported. On a non-Claude primary the captain got a "..."
acknowledging his message and then nothing, forever - a false promise, since
the ack tells him it landed and is being worked when nothing will ever
surface it. That is precisely the "ack then silence" failure this whole
feature exists to end, manufactured by omission instead of a bug.

Two changes, both from the captain's own instructions:

1. Do not ack when nothing can surface the message. bin/fm-tg-fetch.py adds
   harness_can_surface(), which asks bin/fm-harness.sh whether the running
   primary is Claude before ever sending the arrival "...". Only Claude
   registers the Stop hooks (fm-tg-guard.sh/fm-tg-hook.sh) that actually
   drain the inbox and enforce a reply, so acking anywhere else was the false
   promise. The check fails toward "no" (an undetectable harness, a timeout,
   or a missing fm-harness.sh all withhold the ack rather than risk sending
   a false one) and is computed at most once per fetch, not per message.
   bin/fm-tg-drain.py's fallback ack (for a failed arrival-time send) gets
   the identical gate - it is duplicated rather than shared, since a
   hyphenated bin/fm-tg-*.py filename is a script, not an importable module,
   but both copies carry the same reasoning and a cross-reference to keep
   them in sync. This turned out to matter here: bin/fm-tg-poll.sh's own
   recent fix (running fm-tg-drain.py to stamp "surfaced" on its own fetch
   path) meant drain.py's fallback ack was already reachable from the
   harness-agnostic poll path too, not just from Claude's own Stop hook as
   originally assumed - without this gate, drain.py silently un-withheld the
   ack fetch.py had just correctly withheld.

   The message itself is still recorded either way - nothing is lost, only
   the ack (and, by construction, any real surfacing) is withheld.

2. Say so plainly at session start. bin/fm-bootstrap.sh's telegram_setup()
   prints "TELEGRAM: inbound is Claude-only on this setup - ..." once,
   right after the normal arm confirmation, whenever Telegram is configured
   and the detected primary harness is not Claude. Documented in the
   bootstrap-diagnostics skill, same shape as every other TELEGRAM: line.

Documented the split in docs/telegram.md: a new "Inbound is Claude-only"
section states plainly that outbound works on every harness while inbound
does not, with the defect/fix narrative and a note that this was caught
before ship rather than lived through. Cross-referenced from "What it does"
and "No message is lost" so a reader of either does not assume more
coverage than exists.

Added regression tests: the poll withholding the ack (and the recorded
message and a later Claude poll on the same home still acking normally)
on a simulated non-Claude harness, and bootstrap printing (or correctly
omitting) the new diagnostic line depending on the detected harness.
CURSOR_AGENT=1 is used to force a deterministic non-Claude verdict, since
bin/fm-harness.sh checks it before both CLAUDECODE and process ancestry -
unsetting markers alone is not sufficient in an environment where a real
`claude` process is a genuine, close ancestor of the test shell.
tests/fm-telegram.test.sh's scratch-bin helper now also copies
bin/fm-harness.sh and its bin/fm-cursor-lib.sh dependency, which the
existing arrival/guard/reply pipeline test needed once the ack gate reached
into that shared helper too.

Verified locally: bash tests/fm-telegram.test.sh (45/45), bash
tests/fm-pi-watch-extension.test.sh (33/33), bash
tests/fm-supervision-instructions.test.sh (10/10), shellcheck bin/*.sh
bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on the touched
Python modules.
A no-mistakes review of the prior commit (run 01M0P7RF9VYHMKN5ATTY9C1BS8)
found a real, high-severity regression in that commit's own fix
(tg-poll-drain-marks-unsurfaced-messages-surfaced), one auto-fix bug in the
same family that had gone unnoticed in an earlier commit's own fix, one
test-portability bug, and two documentation-accuracy findings. Decided and
fixed all of them myself - all are either straight bugs against already-
accepted behavior or documentation corrections that touch no code, not new
scope calls.

THE REGRESSION. bin/fm-tg-poll.sh and bin/fm-tg-wait.sh both used to run
bin/fm-tg-drain.py afterward with its own stdout discarded, purely to get
its "surfaced" bookkeeping as a side effect. But fm-tg-drain.py surfaces
(and marks) EVERY pending inbox record, not just the one this fetch just
wrote - so a message that arrived earlier and was never actually shown to
the model (e.g. under a non-Claude harness, or simply not yet drained) got
marked surfaced=1 as collateral damage. fm-tg-archive.py treats surfaced>=1
as "retire unconditionally on any real reply, however much later", so the
very next unrelated outbound send silently swept an unseen message into
tg-processed - no "retired unsurfaced" notice, no trace, exactly the "no
message is lost" guarantee this whole feature exists to keep, reopened by
the fix meant to close a different gap.

THE FIX. bin/fm-tg-fetch.py now stamps "surfaced" itself, precisely, for
exactly the record(s) whose content is actually in this call's own real
(non-discarded) output: a poll's "N message(s)... : <preview>" summary only
ever previews message 0 of a batch, so only that record is marked; a wait's
"CAPTAIN: <text>" loop prints every new message in full, so all of them are
marked. bin/fm-tg-poll.sh and bin/fm-tg-wait.sh no longer call
bin/fm-tg-drain.py at all - fetch.py's own mark_surfaced() replaces that
call entirely, inside the fetch's own already-budgeted work rather than as
an unbudgeted extra step after it (which also resolves the review's
tg-poll-drain-outside-check-budget finding as a side effect: it can no
longer overrun FM_CHECK_TIMEOUT because the step it flagged is gone).

OTHER FIXES:
- tg-harness-gate-duplicated-despite-shared-module: harness_can_surface()
  moved from two hand-synced copies (bin/fm-tg-fetch.py and
  bin/fm-tg-drain.py) into bin/fm_tg_records.py, the shared importable
  module both scripts already use for record writes - one owner instead of
  a drift risk, matching this module's existing role for everything else a
  captain message touches.
- tg-tests-depend-on-ambient-claude-harness: three assertions in
  tests/fm-telegram.test.sh required the arrival ack to fire without ever
  setting CLAUDECODE, relying on the host shell resolving to claude by
  ambient marker or process ancestry - true in this dev environment (a real
  claude process is a close ancestor) but false on a CI runner with neither.
  Each now sets CLAUDECODE=1 explicitly. Applied the same fix proactively to
  a fourth site introduced in the prior commit's own new test, not yet
  reviewed.
- tg-doc-claims-no-drain-on-non-claude: resolved as a side effect of the
  regression fix above - fm-tg-drain.py is genuinely never invoked from the
  harness-agnostic poll path any more, so the doc's existing claim is
  accurate again and needed no wording change.
- tg-doc-overstates-argv-token-unavoidability (ask-user, decided - a
  documentation-accuracy correction that changes no code, not a scope
  question): the accepted-tradeoff note claimed no alternative to putting
  the bot token in curl's argv exists at all. `curl -K -` reading the URL
  from a config file on stdin is a real alternative that keeps the token
  off argv; the doc now says so and states the actual reason it is not used
  today - reworking three call sites' request construction to safely quote
  a token, chat id, and free-form captain text into a -K file is its own
  bug surface for a local-only hardening step nobody has asked for. The
  underlying tradeoff itself is unchanged; only the doc's factual claim
  about alternatives is corrected.

Added regression tests for the actual bug: an older, genuinely never-shown
pending message must not get marked surfaced as collateral damage from a
poll fetching an unrelated new one, and must survive an unrelated reply
afterward; a multi-message poll batch marks surfaced only the message its
own preview actually showed, not the rest of the batch.

Verified locally: bash tests/fm-telegram.test.sh (47/47), bash
tests/fm-pi-watch-extension.test.sh (33/33), shellcheck bin/*.sh
bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched
Python module.
… armed

Three consecutive no-mistakes review attempts stalled mid-review on this
branch's now-large cumulative diff (no findings table, no gate, 20+ minutes
of zero activity and no live process each time - a daemon instability, not
a sign of anything wrong with the code). Each stall's partial log still
captured real, legible findings before dying; this commit applies the one
real bug the third attempt's log had already fully diagnosed before it cut
off, rather than losing it to a fourth stall.

bin/fm-supervision-lib.sh's fm_supervision_status() decides whether a
firstmate home needs its watcher kept armed even with zero in-flight tasks.
It already recognized state/x-watch.check.sh (X mode's relay poll shim) as
that reason, but never state/tg-watch.check.sh - the identically-armed
Telegram poll shim (docs/telegram.md). So a Claude primary with Telegram
configured and no other in-flight work never force-armed the watcher at
all: bin/fm-claude-stop-autoarm.sh's need_supervision() gate exits before
ever reaching the Telegram cadence-sourcing lines this session's earlier
commits fixed, making that 30s cadence moot - nothing polled Telegram on an
otherwise-idle Claude primary, on Claude itself, the one harness this whole
feature is supposed to fully support end to end.

This is the same class of gap as the OpenCode plugin's shouldArm() fix
earlier this session (tg-cadence-missing-opencode-pi), just in the shared
predicate bin/fm-guard.sh, bin/fm-turnend-guard.sh,
bin/fm-turnend-guard-cursor.sh, and bin/fm-subagent-pretool-check.sh all
consume too - fixing it here closes the gap for all of them at once rather
than only the one entrypoint the review's partial log named.

Added a regression test mirroring the existing X-mode-poll-need coverage:
a Telegram poll shim alone, with zero in-flight tasks, must still keep the
auto-arm cycle active.

Verified locally: bash tests/fm-claude-stop-autoarm.test.sh (30/30), bash
tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-cursor-primary.test.sh
(27/27), bash tests/fm-procevent.test.sh (42/42), bash
tests/fm-session-lock-ancestry.test.sh (7/7), bash tests/fm-telegram.test.sh
(47/47), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean).
A fourth no-mistakes review attempt also stalled before producing a formal
findings table (the daemon instability from the prior three attempts
persists), but this time completed a full narrative review first,
enumerating five real findings in prose before it died. Decided and fixed
four of them directly (straight bugs against already-accepted behavior, or
naming corrections); documented the fifth as a known, accepted limitation
rather than building new concurrency-safety machinery for it.

THE MAIN FIX. bin/fm-tg-poll.sh's own fetch success path marked the message
it fetched "surfaced" - correct in spirit (only the record actually shown,
not every pending one), but wrong in substance: the poll's own summary
line previews new_texts[0] truncated to 70 characters, a watcher wake
notifying that a message arrived, not a rendering of the message itself.
bin/fm-tg-archive.py treats surfaced>=1 as license to retire the record
unconditionally on the next reply, however much later - correct for a
genuine full-text surfacing (bin/fm-tg-wait.sh's "CAPTAIN: <text>" branch,
unaffected by this fix), wrong for a coarse preview the model may never
have read past. bin/fm-tg-fetch.py's poll branch no longer marks anything
surfaced at all; on Claude, the one harness this ever mattered for, the
record still gets a real, full surfacing every turn end via
bin/fm-tg-guard.sh/bin/fm-tg-hook.sh's own drain regardless of whether the
watcher poll ever ran.

OTHER FIXES:
- The captain-impersonation drop notice (a chat_id mismatch) was never
  actually visible through bin/fm-tg-poll.sh on an otherwise-successful
  poll: fetch.py writes it to stderr and keeps going (rc stays 0), but the
  poll captured stderr to a temp file and deleted it unconditionally before
  the success branch ever ran, so the notice - which docs/telegram.md
  explicitly promises is "visible (never silent)" - was captured and
  discarded, unseen. The diag file is now surfaced on a successful poll
  before it is removed.
- bin/fm-guard.sh's and bin/fm-turnend-guard.sh's "watcher down" banners
  hardcoded "X-mode relay polling" as the reason supervision was needed with
  zero in-flight tasks and zero event sources, the only case that reason was
  ever written for - now that a Telegram poll shim alone can also be that
  reason (this session's earlier bin/fm-supervision-lib.sh fix), an idle
  Telegram-only home got a banner naming the wrong cause. Both banners (and
  a third, identical occurrence in fm-turnend-guard.sh's exhausted-budget
  path that the review's partial log did not reach but shares the exact same
  bug) now name whichever poll shim is actually present.
- mark_surfaced()'s `records` parameter shadowed the module-level `import
  fm_tg_records as records` - dormant (the function never needed the module
  inside its own body) but confusing. Renamed to `surfaced_records`.
- Documented, as a known accepted limitation (not fixed): bin/fm-tg-fetch.py
  and bin/fm-tg-drain.py each do an unsynchronized read-modify-write on the
  same inbox record file, so a genuinely concurrent poll and Stop-hook drain
  can interleave and drop each other's field (surfaced, or an attachment's
  media path). Closing this needs real cross-process locking across three
  scripts - a new concurrency-safety mechanism this feature does not
  otherwise have - for a race whose worst outcomes are one extra surfacing
  or one orphaned attachment, not a lost message or reply.

Updated three existing tests whose assertions assumed the poll's own
truncated preview counted as a full surfacing (it no longer does), and added
regression tests for the drop-notice visibility fix and the corrected
neither-message-marked multi-message-batch behavior.

Verified locally: bash tests/fm-telegram.test.sh (48/48), bash
tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-guard-stale-banner.test.sh
(27/27), bash tests/fm-pi-watch-extension.test.sh (33/33), bash
tests/fm-claude-stop-autoarm.test.sh (30/30), shellcheck bin/*.sh
bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched
Python module.
A fifth no-mistakes review attempt stalled again before a formal gate (the
same daemon instability), but again completed a full narrative first,
catching five more findings - including a real miss in the previous
commit's own fix and a bug in that same commit's fix that never actually
worked in production. Decided and fixed four directly; documented one as a
known, accepted tradeoff rather than risk a change to already-carefully-
tuned retry logic.

THE MISS. The previous commit removed the poll branch's own truncated
70-char preview from being marked "surfaced" - correct - but missed the
identical defect on the far more common path: bin/fm-tg-wait.sh's
"CAPTAIN: <text>" branch truncated at 300 characters and still marked the
WHOLE record surfaced=1. A message longer than 300 characters got the same
treatment: the model could reply from a partial read, and
bin/fm-tg-archive.py would then retire the record as though it had been
shown in full. Unlike the poll branch, this branch genuinely is meant to be
a full surfacing (it is what a Claude Stop hook's own long poll actually
shows the model), so the fix here is not to stop marking it - it is to stop
truncating short of what a real Telegram message can be. Raised the limit
to 4000, matching bin/fm-tg-drain.py's own established ceiling for the
identical job; Telegram's own message cap is 4096, so this is a truncation
limit in name only.

THE BUG IN THE PREVIOUS FIX. bin/fm-tg-poll.sh's fix for the
captain-impersonation drop notice wrote it to stderr - but bin/fm-watch.sh's
real check runner captures only a check's stdout as the wake text and
discards stderr outright, so the fix never actually worked in production.
It passed its own test only because that test captured both streams
together with `2>&1`, masking which one the notice landed on. Now written
to stdout (matching how bin/fm-tg-poll.sh's own record_refusal already
correctly does this), and the test captures stdout alone to catch a
regression back onto the wrong stream.

OTHER FIXES:
- The three watcher-down banners that name a poll-only reason supervision
  is needed (bin/fm-guard.sh, and two sites in bin/fm-turnend-guard.sh) each
  hand-rolled an identical "which poll shim is present" check, introduced
  piecemeal across the last two commits. Consolidated into one shared
  fm_supervision_poll_need_desc() in bin/fm-supervision-lib.sh.
- A .tg-poll-diag.* temp file could leak if the watcher SIGKILLed a slow
  check's whole process group before bin/fm-tg-poll.sh's own cleanup ran.
  SIGKILL cannot be trapped by anything, so this closes only the half that
  can be closed: an EXIT trap now owns the cleanup for every exit path this
  script itself controls, replacing three separate manual rm calls.

Documented, as a known accepted limitation rather than fixed: the waiter's
backoff treats a 409 (the poller/waiter race's own by-design, expected
condition, not a real failure) identically to a genuine unusable reply,
so a busy session can climb its backoff toward the 60s cap during entirely
normal operation. Exempting 409 specifically needs the waiter to also
capture and inspect fm-tg-fetch.py's stderr reason, which it currently does
not; that is a real, scoped fix, just not one to make to carefully-tuned
retry logic - already reshaped by several captain-observed production
incidents this session - without the live traffic needed to verify it does
not reopen the CPU-spin failure the backoff exists to prevent.

Added a regression test for the waiter's 300-character truncation fix, and
corrected the drop-notice test to capture stdout alone rather than both
streams together, so it can no longer pass while the real fix is broken.

Verified locally: bash tests/fm-telegram.test.sh (49/49), bash
tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-guard-stale-banner.test.sh
(27/27), bash tests/fm-claude-stop-autoarm.test.sh (30/30), shellcheck
bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every
touched Python module.
A sixth no-mistakes review attempt stalled again before a formal gate (the
same daemon instability across every attempt on this branch), but again
completed a full narrative first, catching three findings. Fixed two
directly; escalated the third as a genuine architecture question rather
than deciding it myself.

THE NEAR-MISS. The previous commit raised PRINT_MAX from 300 to 4000 to fix
a truncated-preview-marked-as-fully-shown defect, with a comment claiming
4000 "comfortably exceeds Telegram's own 4096-character message cap." It
does not - 4000 < 4096. A 4001-4096-character message (one Telegram
genuinely delivers) was still truncated and still marked surfaced=1, the
identical defect narrowed but not closed. Raised to 4096, Telegram's own
documented cap, in both bin/fm-tg-fetch.py's PRINT_MAX and
bin/fm-tg-drain.py's own separate, less-severely-affected 4000 (same class
of near-miss, same fix). The regression test that exercised only a
500-character message (which 4000 already covered, so it could not have
caught this) now uses a full 4096-character message.

THE 429 DEDUP BUG. state/.tg-poll-error dedups on the literal refusal text
so one standing failure is reported once, not every 30-second cycle -
except Telegram's own 429 description is "Too Many Requests: retry after
N," where N counts down every poll, so the literal line never repeats. A
sustained rate limit therefore woke firstmate on every single cycle, the
exact failure this record exists to prevent, for the one refusal reason
most likely to actually be sustained. record_refusal now dedups on the line
with only that one known-volatile "retry after N" tail normalized out; the
leading error code is left untouched, so a genuinely different refusal (a
different code) still reports as new.

ESCALATED, not decided: harness_can_surface() resolves the harness of the
calling process's own tree, which for the watcher's poll path is the
long-lived watcher process armed once at bootstrap - not necessarily the
harness currently running as the home's primary right now. Nothing ties a
watcher restart to a harness switch on the same FM_HOME (confirmed:
bin/fm-bootstrap.sh only ever suggests a manual restart on a cadence-config
change), so a home switched from Claude to Codex without an intervening
watcher restart could keep sending acks nothing will ever surface - the
exact false promise this session already closed for the common case, just
reopened at an edge the review is right to have found. This is a genuine
architecture question (accept a narrow documented edge case, or build a
watcher-restart-on-harness-mismatch mechanism - new durable state and a new
bootstrap-time check, not a bug fix) rather than a defect against
already-accepted behavior, so it is escalated per the ask-user-authority
framework:
data/fm-telegram-ship-t8/needs-decision-round9-harness-identity-source.md.

Added a regression test for the 429 countdown-dedup fix (a sustained 429
across two polls with different countdowns must report once; a genuinely
different error code afterward must still report).

Verified locally: bash tests/fm-telegram.test.sh (50/50), shellcheck
bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on
every touched Python module.
…ning

The captain's decision on the harness_can_surface() identity-source
escalation: decline building a watcher-restart-on-harness-mismatch
mechanism (new durable state to cover a rare manual reconfiguration is not
worth the complexity), but make the documentation actionable rather than
merely honest - a reader hitting this needs the fix in front of them, not
just a warning that it can happen.

docs/telegram.md's "Inbound is Claude-only" section now states the edge
case plainly and gives the one command that closes it: restart the watcher
(`bin/fm-watch-arm.sh --restart`) after switching a home's primary harness.
Restarting re-arms the watcher under the new primary's own identity, so
harness_can_surface() stops reflecting a stale one immediately.
@Alchemy86
Alchemy86 merged commit b4a8516 into main Aug 23, 2026
10 of 12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.