Telegram captain-comms, shipped into the repo - #1
Merged
Merged
Conversation
…nsion (kunchenguid#205) * fix(bin): use set -u-safe empty-array expansion in pr-merge and spawn Expanding "${arr[@]}" on an empty array under set -u fails on bash < 4.4 (notably macOS bash 3.2). Quote the portable "${arr[@]+"${arr[@]}"}" idiom in fm-pr-merge and fm-spawn batch dispatch so empty arrays expand to nothing. Co-authored-by: Cursor <cursoragent@cursor.com> * test(brief): harden fm-brief regression coverage for parse and scaffolds Tighten bash -n checking, pin literal backtick rendering in the no-mistakes DOD wording assertion, and keep a scout/secondmate scaffold smoke test so the Co-authored-by: Cursor <cursoragent@cursor.com> kunchenguid#166 apostrophe regression cannot return unnoticed. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#693) * fix: make watcher supervision continuous * no-mistakes(review): Bound watcher retries and log attached signals * no-mistakes(review): Add bounded successor-recovery wake fallbacks * no-mistakes(review): Prevent overlapping successor-arm retries * no-mistakes(review): Resume supervision after late arm closes * no-mistakes(review): Bind OpenCode recovery to attempted arm * no-mistakes(test): Synchronize peer beacon regression fixture * no-mistakes(test): Synchronize Pi and OpenCode late-close lifecycle fixtures * no-mistakes(document): Captain: document watcher successor protocol behavior * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* fix: always fetch PR head for review diffs Prefer a freshly fetched refs/pull/<n>/head over a reachable recorded pr_head= so reviewers never hold a merge over a "missing" fix that already landed on the remote PR. Recorded SHA is offline fallback only; local branch is last resort with a warning. Store the tip under refs/fm-review/ so a later base-branch fetch cannot clobber the compare tip via FETCH_HEAD. * no-mistakes(test): Isolate session-start nudge tests from gate state * no-mistakes(document): Correct review-diff documentation
…and skills (kunchenguid#736) * docs: resolve five contract contradictions * no-mistakes(test): align owner-pointer assertions with reworded docs; skip absent shellcheck
* docs(harness): reverify grok exit command * no-mistakes(test): Correct Grok exit resume attribution
* fix(watcher): bound stale wakes for exited paused crew * no-mistakes(review): Gate pause suppression on confirmed agent death * no-mistakes(test): Fixed stale pause cadence * no-mistakes(document): Document dead-agent hold cadence
…id#744) * fix(supervision): distinguish ordinary wakes from repair * no-mistakes(review): Make passive guard follow-ups recovery-only * no-mistakes(document): Clarify recovery-only turn-end guard documentation
* fix(x-mode): dedupe pending mention wakes * no-mistakes(review): fix x-poll claim error deduplication * no-mistakes(review): separate claim diagnostics from relay recovery * no-mistakes(document): Document X-mode once-only mention wakes
…enguid#747) * feat(wake): enrich drained signal context * no-mistakes(review): Bound wake enrichment reads * no-mistakes(document): Document wake-drain annotations * docs(wake): explain at-least-once drain boundary * no-mistakes(review): Prevent symlink races in wake annotations * no-mistakes(review): Exercise wake symlink race regression * test: document intentional AFK marker subprocesses * fix(wake): isolate annotation marker state
* feat(herdr): add optional presentation spaces * no-mistakes(review): Harden Herdr projection creation and spawn serialization * no-mistakes(review): Captain, disarm Herdr cleanup before launch submission * no-mistakes(test): Correct stale Orca metadata failure fixture * no-mistakes(document): Document Herdr presentation projection accurately
…nchenguid#775) * fix(send): treat opencode busy-queued composer state as submitted When fm-send sends a message to a BUSY opencode crewmate on the tmux backend, opencode accepts the Enter and queues the message for the next turn, but leaves the typed text visible in the composer row. The submit-verification loop sees a pending composer, exhausts retries, and reports a false "Enter swallowed" failure while the message is actually delivered. Fix: after Enter retries are exhausted and the composer still shows pending, check fm_pane_is_busy. If the pane is busy (agent mid-turn, footer shows "esc interrupt"), the harness queued the message, so return "empty" (accepted). On an idle pane, keep returning "pending" (genuine swallow detection preserved). Regression tests cover four scenarios: - busy pane + pending composer -> empty (message queued) - idle pane + pending composer -> pending (genuine swallow) - busy pane + composer clears on first Enter -> empty - idle pane + composer clears on first Enter -> empty (existing path) * docs: document busy-queued Enter exception across backend docs and skills Add explanatory comments and backend documentation for the busy-queued Enter fix (opencode 1.18.4 accepts Enter mid-turn but keeps typed text in composer until the turn ends): - bin/fm-tmux-lib.sh: document the busy-aware fallback in the file header and above fm_tmux_submit_enter_core - .agents/skills/afk/SKILL.md: daemon-facing policy note - .agents/skills/harness-adapters/SKILL.md: harness-specific fact - docs/tmux-backend.md: submit-acknowledgement section with the busy-queue exception - docs/herdr-backend.md: record the known gap - docs/architecture.md: cross-reference in the daemon section * test(tmux): fix SC2181 and make busy-submit test executable
…unchenguid#765) * fix(spawn): require two stable reads before accepting worktree path The treehouse-get worktree-detection loop in fm-spawn.sh accepted the first pane_current_path read that differed from the project path, but on some tmux/WSL setups a brand-new window transiently reports a stale-but-real path before the pane actually settles into the worktree. Since that stale path is itself a real, distinct git checkout, it also passes validate_spawn_worktree's isolation check, so the loop silently recorded the wrong worktree in state/<id>.meta (and, for claude harness spawns, installed the turn-end hook there too). Require two consecutive polls to agree on the same non-project path before accepting it, using the existing inter-poll sleep as the confirmation gap so an already-settled pane isn't slowed down by an extra cycle. * fix(tests): drop unused CASE_DIR read in worktree-settle test ShellCheck SC2034: CASE_DIR is split out of the case record but never referenced; discard it with _ instead. --------- Co-authored-by: Freudator86 <tim@allesknut.de>
…anges (kunchenguid#752) * fix(watcher): stabilize Linux process identity * no-mistakes(document): document FM_PROC_ROOT_OVERRIDE and Linux starttime identity rationale
* fix(supervision): verb-aware captain relevance, AFK wedge, head-bound state Stop free-text tokens like "merged" from promoting nonterminal working: lines to captain-relevant, so AFK no longer permanently suppresses idle recovery. Defend wedge aging independently for nonterminal progress verbs, bind no-mistakes current-state attribution to code identity (not branch alone), and mark setup-complete as nonterminal in the ship brief scaffold. * no-mistakes(review): Enforce nonterminal suppression and head-bound run attribution * no-mistakes(document): Document current-code-bound run attribution * no-mistakes(test): Wait for stable Herdr shell readiness * no-mistakes(test): Make Herdr and watcher readiness tests deterministic * no-mistakes(test): Make tmux capture and watcher lifecycle deterministic * no-mistakes(document): Document corrected supervision contracts
* fix: allow safe teardown during watcher recovery * no-mistakes(review): Distinguish unsafe-teardown deny guidance via policy reason code * no-mistakes(document): Sync continuity-gate docs to allow teardown recovery * test: mark dynamic teardown fixture literal * no-mistakes(document): docs: add teardown to continuity gate allow list
…nguid#790) * feat(herdr): order presentation worker spaces * fix(herdr): preserve focus during projected cleanup * no-mistakes(review): Serialize Herdr cleanup and protect active seeded tabs * no-mistakes(review): Serialize Herdr aborts with guarded focus regressions * no-mistakes(review): Fall back flat when Herdr serialization is unavailable * no-mistakes(test): Stabilize watcher startup and AFK handoff tests * no-mistakes(document): Correct Herdr ordering and focus documentation
…#809) * Send literal config reread after inherited config push When declared inherited config changes under an already-running secondmate, build a per-home instruction from validated destination post-write bytes and deliver it on the routed secondmate path. Unchanged config sends nothing; ABSENT represents removal; captain-shared is never inlined. Covers mid-session config-push and the locked bootstrap convergence path without hardening spawn against deliberate runtime choice. * no-mistakes(review): Fix config reread framing, partial propagation, and respawn order * no-mistakes(review): Send config rereads via durable single-line pointers * no-mistakes(review): Make failed config rereads retryable * no-mistakes(review): Make config reread retries generation-safe * no-mistakes(review): Make config rereads durable and ordered * no-mistakes(review): Drain retries, bound history, preserve detect-only read-only mode * no-mistakes(review): Retain write retries and quarantine stale respawn generations * no-mistakes(review): Preserve exact config reread retries and delivery order * no-mistakes(review): Preserve exact retry bytes and bounded quarantine pruning * no-mistakes(document): Consolidated config-reread documentation
* feat(watch): follow GitLab merge requests to merge The merge watch only understood GitHub pull requests, so a task whose deliverable is a GitLab merge request was never followed to merge. Generalize the stored poll identity from owner/repository to a provider-tagged provider/url/host/path/number record. GitLab runs mostly on self-hosted instances and its projects nest under groups at no fixed depth, so the host and the full project path are data in the record rather than constants, and every consumer rebuilds the URL from those parts and refuses any record that does not reconstruct it exactly. The GitLab state is read with plain glab, matching the GitHub path's use of plain gh, so an upstream checkout needs no extra tooling. Two things about glab were established by running it rather than assumed, because a wrong invocation here fails silently into a permanent "not merged": - glab has no field selector, and its JSON would need a JSON processor that firstmate does not require, so the state is read from glab's own field output. Only an exact "merged" wakes firstmate, so a changed format stays silent instead of reporting a merge. - glab cannot take a merge request URL the way gh can, because that form resolves through the current git repository and the watcher has none. It is addressed by project URL and merge request number instead. An absent glab produces no wake rather than a false merge, and arming refuses with a clear message since that is the one point where a missing CLI can still be reported. A GitLab task records no pr_head, which both consumers already treat as optional. The merge path still addresses GitHub only and refuses a merge request URL rather than sending it to the wrong forge. The record version moves to v2, and the existing non-executing migration rebuilds an already-armed watch from its recorded URL, so no watch is lost by upgrading. docs/gitlab-merge-watch.md records the evidence, taken against the public fixture project https://gitlab.com/KarotKris/gitlab-merge-watch-fixture. * no-mistakes(review): Reject github.com host in GitLab MR URL/sidecar validation * no-mistakes(document): Note GitLab MR URLs are explicitly refused, not just malformed ones, in fm-pr-merge.sh docs
…uid#821) * feat(herdr): correct all-home child presentation topology Inherit the presentation opt-in to secondmate homes, label new projected spaces with the approved corner format, insert each child under its owning parent under one session-scoped lock, and keep flat non-destructive fallback. * no-mistakes(review): Exclude secondmates from Herdr presentation projection * no-mistakes(review): Harden shared Herdr locks and ambiguous child ordering * no-mistakes(review): Use adjacency-only Herdr child ownership * no-mistakes(review): Reject foreign legacy projections safely * no-mistakes(review): Validate Herdr session sockets before projection * no-mistakes(test): Fix Herdr teardown fixture session socket metadata * fix(herdr): canonicalize presentation lock socket paths Always resolve the session socket parent directory so symlink parents such as /tmp -> /private/tmp cannot split the shared cross-home lock identity. Refuse relative socket paths. Clarify lock-unavailable warnings. * no-mistakes(test): Fix Bash-compatible GitLab merge request URL parsing * no-mistakes(document): Document all-home Herdr child topology * no-mistakes(lint): Quote fallback provenance string for ShellCheck
* fix(no-mistakes): drop full-suite local Test override Local no-mistakes Test is intent-targeted; CI Behavior keeps the broad tests/*.test.sh suite. Keep commands.lint on bin/fm-lint.sh and add a focused contract test so the override cannot silently return. * no-mistakes(lint): Make CI contract assertion ShellCheck-clean
* feat(test): add canonical timed suite runner and honest CI timeout Introduce bin/fm-test-run.sh as the single serial owner for selecting one script, a family, a conservative changed-file set, or the explicit complete suite, with per-script timing markers and a JSON artifact. Wire CI Behavior through the runner, raise the hang-tripwire timeout to 25 minutes, and document entry points without restoring a full-suite local no-mistakes Test command. * no-mistakes(review): Captain: fix changed selection and empty summaries * no-mistakes(review): Captain: fail closed on unmapped changed sources * no-mistakes(document): Document canonical timed test entry points
* fix: disclose main-home orphan and unstructured inventory gaps Main Bearings could report an empty fleet while structured in-flight rows lacked meta or current backlog rows were free-form. Emit main_inventory from the fleet snapshot, map it into Bearings omitted surfaces and a Charted Next gate, and keep meta as the only live Underway source. * no-mistakes(document): Document Bearings inventory-integrity projection * no-mistakes: apply CI fixes
* feat: add concurrent test isolation proof for Phase 2 Prove an audited portable candidate set passes under concurrent workers with private mode-0700 temp roots, without enabling production CI sharding or fm-test-run --jobs. * no-mistakes(review): Pin isolation proof to audited candidate manifest
* feat(secondmate): parent-owned guards for missed status reports Marked parent-to-secondmate requests now create a durable pending-reply expectation with a privacy-safe correlation id before delivery. Transport success never resolves it; only a correlated parent status or document pointer does. After a completed turn with no report, the parent sends one recovery repost and escalates once if that turn is also missed, without scraping the secondmate conversation or looping. * no-mistakes(review): Deduplicate wrong-home pending-reply sightings * no-mistakes(review): Harden pending-reply recovery and escalation guards * no-mistakes(review): Bound pending-reply backend polling * no-mistakes(review): Cache pending-reply status scans * no-mistakes(review): Protect undelivered pending-reply records from scans * no-mistakes(review): Close pending-reply delivery durability gaps * no-mistakes(review): Separate pending-reply transport outcomes * no-mistakes(review): Escalate stalled pending-reply deliveries once * no-mistakes(review): Resolve attempted deliveries from correlated reports * no-mistakes(review): Resolve late reports after delivery escalation * no-mistakes(document): Document pending-reply grace and ownership * no-mistakes(lint): Silence intentional pending-reply test fixture lint warnings
* feat: add required pinned Herdr CI lane Install exact Herdr 0.7.4 and Treehouse 2.0.1 with official assets and SHA-256 pins, run the real-herdr-gated family serially through fm-test-run with hard-fail on herdr-not-found, and keep portable Behavior free of claimed Herdr coverage. * no-mistakes(document): Consolidate real-Herdr CI documentation ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…guid#841) * feat: shard portable CI tests after isolation proof Balance the Phase 2 proven-isolated set into two LPT portable parallel lanes from Phase 1 timing evidence, keep stateful work in a required portable serial lane, exclude real Herdr to its dedicated required lane, and prove complete inventory coverage with a deterministic guard. Add bounded local --jobs only for the proven set, per-lane timing plus aggregate artifacts, and reduce the interim portable hang tripwire now that the serial remainder owns the long wall-clock path. * no-mistakes(review): Captain, fix CI contracts and completion-order worker scheduling * no-mistakes(review): Captain, preserve stderr gate-skip detection in parallel tests * no-mistakes(document): Document portable sharding and timing aggregation * no-mistakes: apply CI fixes
) * feat: fence primary-session delegation outside the fleet A firstmate primary that delegates through Claude Code's built-in delegation tools creates work with no state/<id>.meta. Because fm-supervision-lib.sh counts *.meta and fm-turnend-guard.sh exits silently at zero, such work does not merely go unsupervised: it makes the whole guard stack structurally inert, and it dies with the primary session. On 2026-07-22 that cost two workers mid-flight and left supervision down for 73 minutes unnoticed. Layer 1, the primary fix: a permissions.deny list in .claude/settings.json removes the 18 delegation, scheduling, worktree, and task-tracking tools from the model's schema, so they are never offered. This is removal rather than interception, so there is no call to intercept and no fail-open path. The list is flat and in one file so its width stays reviewable; the captain owns that width. Layer 2, bin/fm-subagent-pretool-check.sh: a deny list is fail-open against tools that do not exist yet, and permissions.allow is a pre-approval list rather than an availability list, so there is no fail-closed allowlist to use instead. This backstop classifies the tool NAME by shape rather than against a fixed list, so a delegation tool that ships before the deny list is updated is still refused. It excludes mcp__* names and observe-or-stop operations, scopes itself to a genuine primary home via the shared fm_primary_scope_matches predicate so a crewmate's task worktree is unaffected, and offers one deliberate FM_ALLOW_SUBAGENT=1 escape hatch that must be set at launch. Verified live against Claude Code 2.1.217, including a deny-key A/B with a nonsense-name control, layer 2 denying an un-denied Workflow call, the same call allowed in a linked worktree, and the escape hatch. Corrects a prior finding: both Task and Agent work as deny keys, so both are pinned. Codex 0.144.1 verified to expose no delegation tool; grok, opencode, and pi are inspected and documented as not wired because those binaries are absent from this host and the repo requires live validation before trusting a harness hook. Evidence in docs/subagent-guard.md. * no-mistakes(review): Ship scoped Claude delegation guard * no-mistakes(test): Ship Claude delegation deny list * no-mistakes(document): Clarify PreToolUse guard ownership * no-mistakes(lint): Keep Claude deny list local
Reproduction: portable-parallel-2 completed successfully without tasks-axi while fm-decision-hold-lifecycle emitted a gate skip in 30 ms. The pre-shard lane installed tasks-axi and exercised the test fully. Installing tasks-axi is the smallest counterfactual and makes the representative shard execute the test with gate_skip=false in about 20 seconds. Both parallel jobs receive symmetric setup, while the exact 91-test inventory and coverage guard remain unchanged.
* feat: make dispatch profiles quota aware * no-mistakes(review): Fix quota window and Grok product scoping * no-mistakes(document): Document implicit quota-aware dispatch accurately
…lared pause (kunchenguid#2748) * fix(bin): give a captain hold the same bounded pause cadence as a declared pause Two supervisors read a finished task's last status line and disagreed about which declarations mean an idle endpoint is expected. bin/fm-inactive-reconcile.sh suppresses its inactive-outcome scan only on `captain-held`, while the away-mode daemon's wedge path gated deferral on `paused` alone. Both read the LAST line, so the two verbs are mutually exclusive and no finished task waiting on a person could satisfy both at once. Marking 11 such tasks `captain-held:` silenced the 900s outcome scan and immediately produced five possible-wedge escalations in one batch, because the 240s wedge detector no longer saw a pause verb. fm-classify-lib.sh's status_is_paused_or_captain_held already owns the combined question, and bin/fm-watch.sh's ordinary-crew wedge path already asked it. This extends that same answer to the paths still asking the narrower one: - bin/fm-supervise-daemon.sh, all six sites, which form one subsystem and have to move together. classify_stale returns the pause action, reconcile_pause_tracking and migrate_watcher_pause_markers record and migrate the marker, and housekeeping defers the wedge and then re-surfaces the recheck. Changing only the stale-persistence gate would defer the escalation while reconcile_pause_tracking recorded nothing, so the wedge marker would persist and the sweep would `continue` past it forever: quiet, but never re-surfacing. - bin/fm-watch.sh's secondmate stale gate, whose downstream owner pause_state_class already treats both declarations identically. - bin/fm-push-transition-lib.sh's absorb, where either declaration already names the human the transition would report and the wait is already durably recorded. Quieting alone would be half a fix, so the bounded re-surface had to reach a hold too. A hold has no current-state mapping, unlike `paused`, so authoritative crew state reports it as unknown and pause_state_class received `none`. An ordinary crew recovers pause classification from that state through confirmed agent death, which proves no live decision gate is being silenced. A secondmate's endpoint liveness is deliberately never read there, because an idle mate is healthy by design, so that confirmation is unavailable by construction and cannot be required: without recovering the classification for a mate, every caller silenced a held mate outright and its hold would rot invisibly. That promotion is bounded by the declared-wait guard at the top of the function, so it can only reclassify a task that already declared a wait and shows no positive working evidence. Two narrow `status_is_paused` calls are deliberately left alone. bin/fm-crew-state.sh's map_log_state is a current-state reporting contract, not a wedge path; reporting a hold as `paused` would erase the distinction status_key_closing_verb and fm-captain-hold.sh depend on, where a `captain-held` close is a verified durable transfer and a `resolved` close claims outright settlement. fm-classify-lib.sh's call inside status_is_captain_relevant needs no change because that function's own case list already returns non-relevant for `captain-held`. bin/fm-inactive-reconcile.sh keeps its `captain-held` suppression as it is. Its guard exists because a finished task's crew state still reports done from a higher-priority source than the log, and a declared pause needs no such guard: the scan only reports done or failed, and nothing else reaches its record path. Widening it would change a separate subsystem's reporting contract, which this defect does not require. Coverage extends the existing colocated patterns for these predicates and asserts both halves. tests/fm-daemon.test.sh covers the classification, the wedge marker converting to pause tracking with no escalation, the bounded re-surface with its window reset, and the boundary case where an answered hold stops claiming the cadence. tests/fm-watch-triage.test.sh covers a held secondmate re-surfacing on the same bounded cadence without being labeled a wedge. tests/fm-supervision-events.test.sh covers the absorbed push transition. Every one of these fails on the pre-fix code except the answered-hold boundary case, which is there to pin that the quieting was not widened too far. The `paused:` workaround appended to those 11 tasks is live supervision state and is untouched here. It can be retired once this lands. * no-mistakes(review): name the captain in a held task's bounded recheck * no-mistakes(document): extend declared-wait supervision docs to captain-held holds
…guid#2758) * fix(lint): name the installer when ShellCheck or actionlint is missing A missing actionlint exited 127 like a bare command-not-found. Fail with exit 1 and point at the pinned installer, matching the missing-ShellCheck path, without weakening the version pin. * test: isolate kimi and muse detection from inherited Cursor markers Harness detection checks CURSOR_AGENT before ancestry, so these markerless-adapter cases failed when the suite itself ran under Cursor. Clear the verified markers the same way the secondmate harness tests already do. * no-mistakes(document): Document Muse Cursor marker cleanup
…lled but inert (kunchenguid#2684) * feat(checks): report tool updates that are available or installed but inert Firstmate had no way to notice that tooling this home depends on needs an update, and no way at all to notice the worse case: an update that installed correctly and then did nothing. That second case is why this exists. A tool that self-installs into ~/.local/bin while a version manager keeps its own older copy earlier on PATH looks completely up to date to anything that asks only "is a newer version published". On 2026-08-20 a Herdr update landed at 0.8.2 while an older 0.8.0 copy stayed earlier on PATH, so every Herdr command failed on a protocol mismatch and firstmate could not read its own fleet. bin/fm-tool-update-check.sh reports the two conditions separately: <tool> update available a newer version exists at the update source. <tool> update not in effect a newer copy is installed on this host, but PATH still resolves an older one. PATH skew is measured, never inferred. Every executable copy of a watched command on PATH is asked for its own version and those answers are compared, so one lookup cannot hide the skew, and a directory name is never read as a version because a version manager's "latest" directory can hold an older build. A copy that will not report a version is a check failure, not a pass. The watched tools live in local, gitignored config/watched-tools.json, so adding a tool is a config edit rather than a code change, and the file is never propagated to another home. Update sources cover both shapes: a local clone's commit distance from its remote branch, and a command's own version and update announcement, including a tool like no-mistakes that prints its version on one command and announces a new release on another. The check prints one line when something needs attention and prints nothing otherwise, so it rides the existing watcher state-check contract with its trust binding instead of introducing a schedule of its own, and state/.tool-updates keeps the same pending update from being reported on every poll. The check only reports. It never installs, updates, reorders PATH, touches a version manager, or fetches into a watched repository; every git probe is read-only. Tests cover the skew case as a regression, and it was verified by mutation: removing the skew report, or stopping after the first PATH hit as a single lookup would, each make that test fail. * no-mistakes(review): fix tool update check probe reporting, budget, and shim write * no-mistakes(review): keep sweeps alive on broken patterns and oversized budgets * no-mistakes(review): roll back failed arm, widen budget clamp, bound repo probe * no-mistakes(review): guard git probes at the budget, record uncut findings * no-mistakes(document): fix stale watched-tool report-record wording in docs and header * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes The behavior shard's watch-triage suite failed on the new worktree-write wedge tests. Those five tests are the only ones in the file that do not use its standard waits. They give a fixed 3 second liveness budget to the one poll that now spawns the bounded worktree walk, and 4 seconds to an escalating watcher where every other test in the file gives 10. On a loaded runner that poll outlives the fixed budget, so the round is reaped before the deferral it asserts on is recorded, and the test reports a lost deferral instead of the deferral under test. Wait for a completed poll cycle through the file's own wait_poll_cycle, which is what its header documents this hazard for, and use the file's standard 100 tick exit budget. Verified against a load that reproduces the failure: 11 of 12 runs failed before, 8 of 8 pass after. Verified by mutation too, so the waits still prove the behavior: removing the write deferral, and keeping a finished deferral chain across an idle-timer repair, each still fail their test.
* fix: treat yolo as merge authority only, not ask-user finding authority Yolo on/off was documented as also deciding no-mistakes ask-user findings, which hid firstmate's duty to judge unambiguous-toward-design findings itself. Keep every safety boundary; this is a contract clarification, not a relaxation. * no-mistakes(document): Clarify yolo documentation ownership and merge posture
… work over (kunchenguid#2767) * feat(voice): spoken round trip on Nova Sonic 2 with a measured relay cost Step one of the spoken interface: the laptop captures and plays audio, this desktop holds the model session, and no AWS credential leaves the desktop. Measured, amazon.nova-2-sonic-v1:0 in eu-north-1, end of speech to first byte of reply audio, 6 runs each, all answered, on a question that forces a records read: relay path 1.229 1.379 1.428 1.447 1.481 1.516 median 1.438 direct 1.147 1.179 1.203 1.237 1.244 1.317 median 1.220 The relay costs about 0.22s of the median. The direct figure reproduces the earlier survey, which is what makes it a usable control. Excluded: the captain's own ssh round trip, microphone capture, and speaker output. This desktop has no microphone and no speaker, so every run used audio files. Three pieces: bin/fm-voice-relay.py holds the conversation on this host bin/fm_voice_records.py what a spoken answer may read, and the handover bin/fm-voice-client.py the laptop end; audio devices UNVERIFIED bin/fm_voice_frame.py the wire format both machines share Real work is handed to the existing bin/fm-inbox.sh rather than a second queueing surface, and the agent says it is handing over rather than answering as firstmate. Read scope: Done history and free-form note bodies are never assembled at any scope, so the wide default cannot reach the places commercial detail accumulates. config/voice-read-scope narrows it to counts only, and config/voice-read-deny excludes a named item in one line. The boundary is an executable test that widening the reader fails. Push to talk is the default because it is cheaper and the choice is still open; --listen open-mic is the single flip. Two traps worth knowing: a clip with no trailing silence is never answered, and the end of a reply is contentEnd with stopReason END_TURN, not completionEnd. A second user turn in one session is treated as barge-in unconditionally, and an interrupted turn that calls a tool is lost, so the session reconnects per turn and gives up conversational memory. That is the concrete thing step three has to solve. * no-mistakes(review): fix voice relay credential reuse, frame validation and record parsing * no-mistakes(review): test uplink header guard, bound unknown expiry, align state dir * no-mistakes(review): decide deny per item, guard turn failures, bound ambient credentials * no-mistakes(review): read account config from home, harden deny and turn failures * no-mistakes(review): close status verb set, fix inbox help, pair data override * no-mistakes(review): keep profile-free relay alive, unblock loop, fix dead assertion * no-mistakes(review): hide finished pull requests, refuse open mic, keep suite offline * no-mistakes(review): survive reader failures, release devices, fix claims A failure while handling a model event, or while sending a tool result, left the reader task dead with ended and turn_done clear, and close() re-raised the stored failure on every await. One dropped stream became a relay that could never build another session. The reader now reports the session over in a finally whatever killed it, and close() absorbs the task the same way it already absorbed its sends. The laptop client releases what it already started when a later startup step refuses, SystemExit from the handshake wait included, and names a device refusal instead of leaking a raw PortAudio error. Whether it releases correctly against a real device is still unverified here. The records docstring claimed every reading was filtered to open ids. Only the pull request count and list are; the worker count and the state histogram cover every live runtime record, finished ids included, because a meta file still on disk still needs tearing down. The finished-work deny half of the suite asserted things that held with the deny list absent. It is replaced by a deny on an open title, which removes the row and says so while the count stays honest. * no-mistakes(review): name reader failures, split file and device refusals A failure inside the model reader released the waiting turn and told nobody. The session was not marked spent, no notice reached the client, and the client waits for a reply end or a notice, so the captain got their whole timeout of silence and then a record saying the turn went unanswered with nothing about why. Both ends of the relay now name a failed turn through one function, once per turn, and --self-test carries the cause in relay_error the way the client's own record does. Two things that are not failures stay that way. A stream that simply ends is the end of a session, which serve still reads on its own terms. A stream that goes away because close() asked it to is an ordinary renew, and announcing it would have put a failure notice in front of the captain on every turn. On the laptop end, the refusal that became a device error covered the file-backed playback and capture too, so a mistyped --in-file was reported as an audio device failure and the advice named the flag that had just failed. The file ends now report the path and the flag that chose it and stay an OSError; the device ends keep the device advice and name the flag for that end. The device paths remain unrun here, so only the file halves are covered by a test. * no-mistakes(test): survive model session end, order client turn frames * no-mistakes(document): sync voice relay docs with reviewed relay behavior * no-mistakes(document): re-measure relay latency and correct its cause * no-mistakes(document): correct measurement date and name the unmeasured SSH hop * no-mistakes(document): describe the unpublished control measurement, fix list formatting * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
…unchenguid#2763) * fix: keep Relay public loops open until retire Delivering a promised-final reply was deleting the only record that tied a public thread to later work, so a follow-on ship silently owed no closing reply. Retain the registration after delivery, rechain follow-on work onto the same thread, and make retire --reason the only close. * no-mistakes(review): Propagate public follow-up registration removal failures * no-mistakes(review): Persist retire receipts and align parent resolution * no-mistakes(review): Make rechain resumable after partial obligation creation * no-mistakes(review): Repair follow-up state, briefs, and expiry escalation * no-mistakes(review): Serialize follow-up delivery stamps with retirement * no-mistakes(review): Serialize rechain claims and protect registration terminal states * no-mistakes(review): Avoid reporting retired delivery loops as open * no-mistakes(document): Refresh public-loop documentation and verification evidence * no-mistakes: apply CI fixes * no-mistakes(review): Preserve delivered follow-up bindings during registration replay * no-mistakes(review): Harden public follow-up retirement and rechain races * no-mistakes(review): Fail closed on unresolved secondmate retirement * no-mistakes(review): Bind secondmate cleanup to its recorded canonical home * no-mistakes(review): Fix rechain command output and expiry validation * no-mistakes(review): Validate brief keys and warn on remote promotion * no-mistakes(document): Document retained public follow-up loops * no-mistakes(lint): Remove unused bounded-wait loop variable
…ath (kunchenguid#2779) * feat(bin): merge GitLab merge requests through the guarded PR merge path bin/fm-pr-lib.sh already parses a GitLab merge request URL for the watcher, but bin/fm-pr-merge.sh refused every non-github provider, so a merge request had to be merged by hand and got none of the recording, guards, or audit trail a pull request gets. The merge path now dispatches on the parsed provider. A GitHub URL keeps its exact previous behavior. A GitLab URL is addressed through glab by the project URL rebuilt from the parsed host and path, so a merge request on any instance resolves and no host is hardcoded, and no merge-method flag is added because the project's own merge method is what should apply. A GitLab merge happens only after one live read of the merge request confirms it is open, detailed_merge_status is mergeable, has_conflicts is false, blocking_discussions_resolved is true, and the head pipeline succeeded at the exact current head. Every failing condition is reported, not just the first. The verified head is bound to the merge with glab's --sha, so a push landing between the read and the merge fails the merge instead of landing commits nothing verified. Recorded metadata is never the authority for any of this: a rebase moves the head and leaves a recorded value stale, so a recorded head that disagrees with the live one is reported rather than trusted, and the recorded value is read before the recording step because that step drops a GitLab head it cannot resolve. * no-mistakes(review): reject bundled -R clusters and make tool-absence cases host-independent * no-mistakes(test): state authorised GitHub narrowing of bundled -R guard This branch NARROWS GitHub behaviour. The narrowing was authorised deliberately rather than slipping in by accident, and it applies to both providers, GitHub and GitLab alike, because a script that guards one provider and not the other is a trap for the next reader. What bin/fm-pr-merge.sh now refuses is extra merge arguments containing a bundled short-option cluster that includes R, for example "-dR other/repo". The forge CLIs expand such a cluster one character at a time, so it carries "--repo other/repo", and that later value wins over the repository the URL named. Before this change, "fm-pr-merge.sh <task> <github-url> -- -dR other/repo" reached "gh-axi pr merge 12 --repo example/repo --squash -dR other/repo" and exited 0 with pr= recorded and the merge poll armed. It now exits 1 with "extra merge arguments must not override the repository", records nothing, and invokes no forge merge command. Every other GitHub invocation is byte-identical to the base commit. Closing that hole honours the existing rule rather than departing from it. The file header already forbids --repo and -R because the repository must come only from the URL, so a bundled cluster carrying a repository override was never legitimate behaviour to preserve: it was that guard being evaded. Redirecting a merge to a repository the URL does not name is exactly what the guard exists to prevent. The refusal is already pinned on both paths by the existing case test_bundled_repo_override_args_refuse_before_recording in tests/fm-pr-merge.test.sh. On GitHub ("-dR wrong/repo") and on GitLab ("-yR https://other.example/g/p") it asserts exit 1, the refusal wording, no pr= in the task meta, no armed merge poll, and no forge merge command invoked, with a control case proving a cluster that carries no repository override still reaches the forge. No duplicate assertion was added. Both assertions were confirmed to have teeth by narrowing the guard back to a bare -R and watching each path fail. This commit carries no file change: the guard and its coverage landed in 614853d, and this message exists so the pull request description states the narrowing. * no-mistakes(document): fix README pointer for GitLab watch and merge doc * no-mistakes: apply CI fixes
kunchenguid#2788) * no-mistakes: apply CI fixes * fix(bin): drop a private record citation and narrow the review rule Three corrections to the spoken interface that landed in kunchenguid#2767, plus one fix carried over from that branch after its pull request had already been merged. The confidentiality fix. The module docstring of bin/fm-voice-relay.py cited a private, gitignored fleet record by exact path and section number. That widens what this public repository points at, and it cannot resolve for any reader here, because the path has never been in the repository. Both traps it pointed at are already described in full in the list immediately below it, and docs/voice-relay.md carries the same two for operators with no citation at all, so the pointer is removed and no claim is weakened by losing it. Two comments that referred to "the survey" as though it were something a reader could open are reworded the same way. Neither exposed a path, so that half is comprehensibility rather than confidentiality. The review rule. .greptile/rules.md is kept, because its conditions are right and deleting it would leave the next reviewer to re-litigate a decision already argued out. What was wrong with it is narrower than its existence: it read as settled repository policy, when whether VISION.md itself should be reconciled is an open question belonging to the captain. One sentence now says so, and says that the conditions listed below it are what the interpretation depends on. That narrows the claim rather than widening it. The carried-over fix. The first commit on this branch is 7f98e79 from fm/voice-relay-build-v4, taken verbatim rather than rewritten. It closes the window where a transport failure was recorded and then erased, so a run could be emitted as answered false with relay_error null. That matters more than it looks: relay_error is the field that keeps an infrastructure failure from being averaged into a latency figure, so the failure mode is a dead connection wearing the costume of a slow reply. It landed fifteen minutes after kunchenguid#2767 merged and so never reached the default branch. * no-mistakes(review): name a reason on every unanswered-turn close path * no-mistakes(review): guard the downlink body and pin frames to their turn * no-mistakes(review): attribute reply audio to its own turn and tell endings apart * no-mistakes(review): tell a cut-short reply from an unanswered turn * no-mistakes(review): discard reply audio arriving after the output closes * no-mistakes(review): count discarded reply audio on the speaker path too * no-mistakes(review): keep a reason off a turn already answered in full * no-mistakes(review): say a reset cut a reply short, not that none arrived * no-mistakes(review): read one turn's audio count once, and hush a tidy exit * no-mistakes(document): fix stale session-end relay_error claim in voice-relay guide
…repo Moves the previously untracked state/fm-tg-* scripts into bin/, using the standard FM_ROOT/FM_HOME/STATE resolution pattern so a configured clone just works instead of the scripts only running against one hardcoded machine path. Structural fix: registers the two Stop hooks (fm-tg-guard.sh, fm-tg-hook.sh) in this repository's own tracked .claude/settings.json, project-scoped, instead of the user's global settings. That was the root cause of crewmates draining and answering the captain's Telegram inbox themselves. fm-tg-isfirstmate.sh stays as defense in depth for a crewmate working on firstmate's own repo. Extracts the poller/waiter's duplicated, single-quote-fragile inline Python into one shared bin/fm-tg-fetch.py, so both fetch paths ack on arrival and cannot drift apart again. Carries forward every live fix made to the untracked scripts during this task: arrival-time acknowledgement, the poller/waiter race fix, the reply/surface race fix in fm-tg-archive.py (retire on 60s-old-and-unsurfaced, not just surfaced), and the upload fixes (sendDocument for >1MB images, a payload-scaled upload timeout with an explicit timeout message instead of 'unparseable reply'). Adds bootstrap integration mirroring the X-mode pattern: a presence-gated TELEGRAM: report line, a generated state/tg-watch.check.sh poll shim, and a generated config/tg-mode.env 30s watcher cadence file. Absent or partial config stays completely silent and inert. Adds docs/telegram.md (setup, architecture, every guarantee and the defect that produced it) and tests/fm-telegram.test.sh (hermetic, no live Telegram traffic - a fake curl stands in for the network so fm-tg-send.sh and fm-tg-poll.sh still run for real). Drops fm-tg-digest.sh: it was never wired into the watcher or anything else, so it shipped no live behavior. Also fixes tests/fm-bootstrap.test.sh to isolate FM_TG_ENV_OVERRIDE, since Telegram config lives at a fixed $HOME path (deliberately, one bot per machine) rather than inside each test's scratch FM_HOME.
… Cursor stand-down
Amendment 5 (live failures after amendment 4): - fm-tg-archive.py: the 60s archive window was an invented number and still let a ~50s-old message survive a reply and re-surface (third recurrence of the duplicate-reply bug). The correct rule is that a reply covers everything already in the inbox when it was sent; only a message under 10s old (genuinely unread) stays pending. - fm-tg-send.sh: sendMessage parts now retry up to 3 times with a 2s backoff on a transient failure, and never retry a genuine Telegram rejection (ok:false), since retrying a bad request cannot help. Also fixes a real bug found while validating this branch: fm-tg-hook.sh created state/tg-inbox and state/tg-processed unconditionally, before ever checking whether Telegram was configured at all - a state mutation on every Stop hook firing even with zero Telegram config, violating the 'absent config means completely inert' guarantee. The config-absent test only passed before because it ran through the real bin/fm-tg-isfirstmate.sh from inside a live crewmate session, whose ancestry check short-circuits the hook before the mkdir is ever reached - masking the bug in exactly the dev loop used to write it. Fixed the hook to check config before creating anything, and reworked the affected tests to use the identity-stubbed scratch bin (already used elsewhere in this suite) so the config-absent path is actually exercised rather than skipped past by ambient ancestry. Verified the fix catches the regression: reverting fm-tg-hook.sh's check reproduces the test failure; restoring it passes again.
…ions A no-mistakes review of this branch flagged two ask-user findings; the captain decided both directly (both were his own design calls, not the reviewer's or the crewmate's): 1. tg-no-chat-filter (security): inbound Telegram updates were never filtered by TG_CHAT_ID, so anyone who found the bot's public username could message it and be recorded/surfaced/acknowledged as the captain, then force a blocking turn-end demanding a reply that lands in the real captain's chat. bin/fm-tg-fetch.py now drops any update whose chat id does not match TG_CHAT_ID before it is ever recorded, with a visible (never silent) stderr line naming both chat ids so a legitimately wrong-chat message from the captain himself stays diagnosable. 2. tg-archive-unsurfaced-loss (data loss): the arrival-time retirement window in fm-tg-archive.py applied to ANY non-ack send, not just an actual reply to what was pending, so a proactive/unrelated message from firstmate could retire an old message that had never actually been shown to the model - a direct 'no message is lost' violation the captain had not seen yet. Also, a missing/malformed 'ts' made a record look infinitely old (float(None or 0) == 0.0), retiring it on the very first send, sight unseen. Fixed per the captain's design: a record that has been surfaced at least once (bin/fm-tg-drain.py now stamps 'surfaced_ts' on first surfacing) always retires on any real reply, however much later - keyed on first-surfaced time, not arrival time. A record that has never been surfaced retires only within a short, conservative 10s window of arrival - the genuine same-turn race the window exists for - and stays pending outside that window no matter what unrelated reply goes out. A missing/malformed ts is treated as brand new, never as infinitely old. This flips the polarity of the two tests that encoded the earlier (now-superseded) 'anything older than 10s retires' rule; rewrote them along with the rest of the affected suite to match the corrected semantics, and added dedicated boundary, surfaced-always-retires, missing-ts, and chat-id-filter tests.
… changed test_unsurfaced_retirement_is_recorded used a 120s-old never-surfaced record and expected it to retire with a notice. That was the semantics before the captain's ask-user decision (archive-window fix): a never-surfaced record over 10s old now stays pending instead of retiring, so this test needed a fresh (<10s) record to still exercise the same-turn-race retirement path it is actually testing.
…s decision The auto-retry from amendment 5 can rarely deliver a duplicate message: an empty sendMessage reply from a curl that itself exited 0 is ambiguous (Telegram may have already accepted and answered the send, just too slowly for --max-time to see the reply), and retrying that case can resend an already-delivered message. This was flagged as an ask-user finding by no-mistakes since it is a genuine tradeoff of the captain's own explicit retry request, with no clean fix available (Telegram's Bot API has no idempotency key for sendMessage). The captain's decision: accept the tradeoff as-is. A rare duplicate costs him nothing to notice and ignore; a dropped message costs him the answer, and he has already stated that preference in fm-tg-guard.sh's own wording. Retry stays, unconditionally, for both the definite-non-send and the ambiguous case. Per his request, bin/fm-tg-send.sh now distinguishes the two cases in its retry log line (a nonzero curl exit is a definite non-send; an empty reply from a curl that itself exited 0 is the ambiguous, may-have-already-sent case) so the rare case is diagnosable after the fact, and docs/telegram.md names the tradeoff plainly instead of leaving it implicit. This is observability only - the actual retry behavior is unchanged for both cases, per the decision to keep retry unconditional.
The no-mistakes pipeline surfaced these 5 findings, all auto-fix, and began applying them, but the fix round appeared to stall (no live process, no filesystem activity for 5+ minutes despite a complete-looking log summary) and had to be aborted before it committed - a recurring pattern on larger fix rounds this session. Implementing them directly rather than losing them a second time: 1. telegram_setup() had no secondmate-home gate. The Telegram config is deliberately machine-wide (~/.config, unlike X mode's FM_HOME-scoped .env), so a live secondmate home would arm its own competing poller against the same bot token - and since getUpdates offsets are server-side, whichever home polls first silently steals the update. Gated on the primary, matching startup_memory_budget_setup's precedent. 2. fm-tg-fetch.py let a write_record() failure "continue" to the next update, which then advanced the offset past the failed one - permanently losing it, since Telegram never re-sends an already-acknowledged offset. Changed to "break": the whole batch from the failure point on is left unacknowledged and retried next poll. A batch that surfaced nothing due to a write failure is now reported as a refusal instead of reading as a quiet channel. 3. fm_dir_is_child_worktree (bin/fm-primary-scope-lib.sh, shared beyond Telegram) compared --git-dir and --git-common-dir as raw strings, but git returns them in different path formats from a subdirectory of a plain checkout (absolute vs relative) - misclassifying any such subdirectory as a crew task worktree. Added fm__git_dir_abs to canonicalize both sides via cd+pwd -P before comparing, applied to both fm_dir_is_child_worktree and fm_primary_scope_matches. 4. send_part() decided "rejected, not transient" via a raw substring match on the reply body, missing any whitespace-formatted variant of the same JSON. check() now exits 2 for a parsed-but-not-ok rejection (1 stays for unparseable/empty), and send_part branches on that exit code instead. 5. Reworded the missing/malformed-ts comment (in three places: the archive.py docstring, its inline comment, and docs/telegram.md) from "treated as brand new" to "unknown, never retired by the window" - the old wording was self-contradictory, since a genuinely brand-new record IS eligible for the window and an unknown-ts one never is. Re-verified: 36/36 telegram tests pass, shellcheck clean, and the fm-turnend-guard and fm-cursor-primary suites (both depend on the shared fm-primary-scope-lib.sh predicate) pass with no regressions.
…QX4MG1J
The pipeline stalled again mid fix-round, so these six findings from that
run's review gate are applied by hand instead of losing another round.
- wait-path-never-marks-surfaced: bin/fm-tg-wait.sh's own fetch success path
printed the fetched text and exited without ever running the drain, so a
message delivered through that path (not the drain-first branch at the top
of the loop) was left at surfaced=0 with .tg-last-surfaced unset. A reply
composed more than 10s later then found the record still pending under
fm-tg-archive.py's never-surfaced window and it re-surfaced forever. Run
the drain silently first, matching the bookkeeping the drain-first path
already gave every other delivery.
- curl-timeout-mislabelled-definite-non-send: bin/fm-tg-send.sh classified
every nonzero curl exit as a definite non-send. A timeout-shaped exit
(curl's own --max-time exit 28, or 124/143 from this script's outer
fm_run_timed bound) means Telegram may have already accepted and answered
the send just too slowly to see the reply - the same ambiguous,
may-have-already-sent case as an empty reply from a curl that exited 0,
not a definite non-send. Both still retry the same way; only the logged
reason changes. docs/telegram.md's retry-tradeoff section is corrected to
match.
- drain-crashes-on-non-dict-record: bin/fm-tg-drain.py appended
rec.get("ts") outside the try/except wrapping json.load, so a record that
parses as JSON but is not a dict (e.g. an interrupted write that left a
bare number or array) crashed the whole drain instead of being skipped,
losing every pending message in the inbox, not just the corrupt one.
- archive-main-loop-bypasses-load-helper: bin/fm-tg-archive.py's retirement
loop called raw json.load() instead of the file's own non-dict-tolerant
load() helper that prune() already relies on. The same non-dict record
crashed retirement and pruning, invisibly, since the script's only caller
discards stderr and runs it with "|| true".
- test-suites-not-isolated-from-machine-wide-tg-config: ~20 test suites
besides tests/fm-bootstrap.test.sh call fm-bootstrap.sh or
fm-session-start.sh without isolating Telegram's fixed-$HOME config path,
so a machine with real Telegram captain-comms configured (such as the one
this repo's own fleet runs on) leaks an unexpected TELEGRAM: line and a
live cadence-relevant side effect into their output. Default
FM_TG_ENV_OVERRIDE to a nonexistent path in tests/lib.sh so every sourcing
suite is isolated without hand-editing each one; a suite that actually
exercises Telegram config still overrides it explicitly afterward. The one
suite that does not source tests/lib.sh (the opt-in credentialed Claude
Stop auto-arm live E2E) gets the same override directly.
- contradictory-cadence-lines: when Telegram captain-comms was active and X
mode was not, the session-start supervision block printed "X mode:
inactive; use the default watcher cadence" immediately above "Telegram
captain-comms: active; source ... 30s cadence is inherited" - two adjacent
lines giving opposite cadence instructions for the same watcher launch.
The X mode line now reports its own inactive status without a cadence
claim when Telegram is carrying that instruction instead.
Added regression tests for all six: the waiter's own fetch path marking a
record surfaced, a curl timeout (exit 28) logged as ambiguous rather than a
definite non-send, a non-dict record surviving both the drain and the
archive pass without losing a good record alongside it, and the
supervision-instructions renderer no longer emitting the contradictory pair
when Telegram is on and X mode is off.
Verified locally: bash tests/fm-telegram.test.sh (40/40), bash
tests/fm-supervision-instructions.test.sh (10/10), shellcheck bin/*.sh
bin/backends/*.sh tests/*.sh (clean), plus spot checks of
fm-bootstrap.test.sh, fm-session-start.test.sh, fm-fleet-sync.test.sh,
fm-x-mode.test.sh, fm-secondmate-sync.test.sh, fm-secondmate-liveness.test.sh,
and fm-shared-captain-inheritance.test.sh against the tests/lib.sh default
change. fm-bootstrap.test.sh and fm-session-start.test.sh each have one
pre-existing unrelated failure, confirmed present against the unmodified
tests/lib.sh from HEAD as well, so neither is a regression from this change.
- tg-cadence-missing-opencode-pi (auto-fix): the OpenCode plugin and Pi extension's arm-child spawn only sourced config/x-mode.env before starting bin/fm-watch-arm.sh, missing the config/tg-mode.env line the Claude and Cursor arm paths already had. A Telegram-configured OpenCode or Pi primary polled at the 300s default instead of tg-mode.env's 30s. Fixed both spawn lines to match. Testing this exposed a second, more severe instance of the same parity gap: the OpenCode plugin's shouldArm() only treats x-mode.env presence as a reason to arm the watcher on an idle home with no fleet work; without the matching tg-mode.env check, a Telegram-configured idle OpenCode primary never armed the watcher at all, making the cadence fix moot. Added tg-mode.env there too. The Pi extension has no equivalent idle-arm gate (its arm is tool-driven), so it needed only the cadence-file fix. - tg-poll-path-no-surfaced-stamp (ask-user, decided - same bug class as wait-path-never-marks-surfaced from the previous fix round, not a new judgment call): bin/fm-tg-poll.sh had the identical omission bin/fm-tg-wait.sh's own fetch path had before its last fix - it prints fm-tg-fetch.py's summary directly and never runs the drain, so a message the watcher poll wins the getUpdates race for is recorded but never marked surfaced. Applied the identical fix: run fm-tg-drain.py silently after a genuinely usable poll (rc==0). - tg-poll-error-cleared-on-hard-failure (auto-fix): the standing refusal record at state/.tg-poll-error was cleared on ANY non-3 exit from fm-tg-fetch.py, conflating a genuine success (rc==0) with an unexpected failure (python3 vanishing from PATH -> 127, the watcher killing the check mid-fetch -> 137/143). An unexpected failure now reports and remembers itself the same dedup'd way a channel refusal does, and only a confirmed rc==0 poll clears the record. - tg-test-unsets-isolation-default (auto-fix): three tests in tests/fm-telegram.test.sh reset state with a bare `unset FM_HOME FM_TG_ENV_OVERRIDE`, discarding tests/lib.sh's Telegram isolation default for every test defined after them in the file, not just the one resetting - reopening exactly the machine-config leak the previous fix round's tests/lib.sh default was meant to close. Each site now restores the captured default instead of dropping it. Added a final invariant test, run last in the file's own invocation list, that catches a regression in any earlier test's reset rather than just these three. tg-stop-hooks-claude-only (ask-user) is NOT decided here - it is a genuine scope question (build Stop-hook/extension-hook coverage for five other harnesses, a multi-day expansion, versus documenting today's Claude-only scope as a known limitation) rather than a defect against already-accepted behavior, so it is escalated per the ask-user-authority framework: data/fm-telegram-ship-t8/needs-decision-round6-stop-hooks-scope.md. Added regression tests for all four decided fixes: cadence-file sourcing and idle-arm parity for both the OpenCode plugin and the Pi extension, the poll path's own surfaced/last-surfaced stamping (plus updating an existing test whose "surfaced": 1 assertion is now correctly "surfaced": 2 once the poll's own fix already stamps it once), the poll's unexpected-exit record preservation, and the isolation-default survival invariant. Verified locally: bash tests/fm-telegram.test.sh (43/43), bash tests/fm-pi-watch-extension.test.sh (33/33, covers both the Pi extension and the OpenCode plugin), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), node --check on both edited plugin/extension files, python3 -m py_compile on the touched Python modules.
The captain's decision on tg-stop-hooks-claude-only: decline building Stop-hook/extension equivalents for five other harnesses (a multi-day expansion, not this task's scope), but the current unqualified behavior is worse than unsupported. On a non-Claude primary the captain got a "..." acknowledging his message and then nothing, forever - a false promise, since the ack tells him it landed and is being worked when nothing will ever surface it. That is precisely the "ack then silence" failure this whole feature exists to end, manufactured by omission instead of a bug. Two changes, both from the captain's own instructions: 1. Do not ack when nothing can surface the message. bin/fm-tg-fetch.py adds harness_can_surface(), which asks bin/fm-harness.sh whether the running primary is Claude before ever sending the arrival "...". Only Claude registers the Stop hooks (fm-tg-guard.sh/fm-tg-hook.sh) that actually drain the inbox and enforce a reply, so acking anywhere else was the false promise. The check fails toward "no" (an undetectable harness, a timeout, or a missing fm-harness.sh all withhold the ack rather than risk sending a false one) and is computed at most once per fetch, not per message. bin/fm-tg-drain.py's fallback ack (for a failed arrival-time send) gets the identical gate - it is duplicated rather than shared, since a hyphenated bin/fm-tg-*.py filename is a script, not an importable module, but both copies carry the same reasoning and a cross-reference to keep them in sync. This turned out to matter here: bin/fm-tg-poll.sh's own recent fix (running fm-tg-drain.py to stamp "surfaced" on its own fetch path) meant drain.py's fallback ack was already reachable from the harness-agnostic poll path too, not just from Claude's own Stop hook as originally assumed - without this gate, drain.py silently un-withheld the ack fetch.py had just correctly withheld. The message itself is still recorded either way - nothing is lost, only the ack (and, by construction, any real surfacing) is withheld. 2. Say so plainly at session start. bin/fm-bootstrap.sh's telegram_setup() prints "TELEGRAM: inbound is Claude-only on this setup - ..." once, right after the normal arm confirmation, whenever Telegram is configured and the detected primary harness is not Claude. Documented in the bootstrap-diagnostics skill, same shape as every other TELEGRAM: line. Documented the split in docs/telegram.md: a new "Inbound is Claude-only" section states plainly that outbound works on every harness while inbound does not, with the defect/fix narrative and a note that this was caught before ship rather than lived through. Cross-referenced from "What it does" and "No message is lost" so a reader of either does not assume more coverage than exists. Added regression tests: the poll withholding the ack (and the recorded message and a later Claude poll on the same home still acking normally) on a simulated non-Claude harness, and bootstrap printing (or correctly omitting) the new diagnostic line depending on the detected harness. CURSOR_AGENT=1 is used to force a deterministic non-Claude verdict, since bin/fm-harness.sh checks it before both CLAUDECODE and process ancestry - unsetting markers alone is not sufficient in an environment where a real `claude` process is a genuine, close ancestor of the test shell. tests/fm-telegram.test.sh's scratch-bin helper now also copies bin/fm-harness.sh and its bin/fm-cursor-lib.sh dependency, which the existing arrival/guard/reply pipeline test needed once the ack gate reached into that shared helper too. Verified locally: bash tests/fm-telegram.test.sh (45/45), bash tests/fm-pi-watch-extension.test.sh (33/33), bash tests/fm-supervision-instructions.test.sh (10/10), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on the touched Python modules.
A no-mistakes review of the prior commit (run 01M0P7RF9VYHMKN5ATTY9C1BS8) found a real, high-severity regression in that commit's own fix (tg-poll-drain-marks-unsurfaced-messages-surfaced), one auto-fix bug in the same family that had gone unnoticed in an earlier commit's own fix, one test-portability bug, and two documentation-accuracy findings. Decided and fixed all of them myself - all are either straight bugs against already- accepted behavior or documentation corrections that touch no code, not new scope calls. THE REGRESSION. bin/fm-tg-poll.sh and bin/fm-tg-wait.sh both used to run bin/fm-tg-drain.py afterward with its own stdout discarded, purely to get its "surfaced" bookkeeping as a side effect. But fm-tg-drain.py surfaces (and marks) EVERY pending inbox record, not just the one this fetch just wrote - so a message that arrived earlier and was never actually shown to the model (e.g. under a non-Claude harness, or simply not yet drained) got marked surfaced=1 as collateral damage. fm-tg-archive.py treats surfaced>=1 as "retire unconditionally on any real reply, however much later", so the very next unrelated outbound send silently swept an unseen message into tg-processed - no "retired unsurfaced" notice, no trace, exactly the "no message is lost" guarantee this whole feature exists to keep, reopened by the fix meant to close a different gap. THE FIX. bin/fm-tg-fetch.py now stamps "surfaced" itself, precisely, for exactly the record(s) whose content is actually in this call's own real (non-discarded) output: a poll's "N message(s)... : <preview>" summary only ever previews message 0 of a batch, so only that record is marked; a wait's "CAPTAIN: <text>" loop prints every new message in full, so all of them are marked. bin/fm-tg-poll.sh and bin/fm-tg-wait.sh no longer call bin/fm-tg-drain.py at all - fetch.py's own mark_surfaced() replaces that call entirely, inside the fetch's own already-budgeted work rather than as an unbudgeted extra step after it (which also resolves the review's tg-poll-drain-outside-check-budget finding as a side effect: it can no longer overrun FM_CHECK_TIMEOUT because the step it flagged is gone). OTHER FIXES: - tg-harness-gate-duplicated-despite-shared-module: harness_can_surface() moved from two hand-synced copies (bin/fm-tg-fetch.py and bin/fm-tg-drain.py) into bin/fm_tg_records.py, the shared importable module both scripts already use for record writes - one owner instead of a drift risk, matching this module's existing role for everything else a captain message touches. - tg-tests-depend-on-ambient-claude-harness: three assertions in tests/fm-telegram.test.sh required the arrival ack to fire without ever setting CLAUDECODE, relying on the host shell resolving to claude by ambient marker or process ancestry - true in this dev environment (a real claude process is a close ancestor) but false on a CI runner with neither. Each now sets CLAUDECODE=1 explicitly. Applied the same fix proactively to a fourth site introduced in the prior commit's own new test, not yet reviewed. - tg-doc-claims-no-drain-on-non-claude: resolved as a side effect of the regression fix above - fm-tg-drain.py is genuinely never invoked from the harness-agnostic poll path any more, so the doc's existing claim is accurate again and needed no wording change. - tg-doc-overstates-argv-token-unavoidability (ask-user, decided - a documentation-accuracy correction that changes no code, not a scope question): the accepted-tradeoff note claimed no alternative to putting the bot token in curl's argv exists at all. `curl -K -` reading the URL from a config file on stdin is a real alternative that keeps the token off argv; the doc now says so and states the actual reason it is not used today - reworking three call sites' request construction to safely quote a token, chat id, and free-form captain text into a -K file is its own bug surface for a local-only hardening step nobody has asked for. The underlying tradeoff itself is unchanged; only the doc's factual claim about alternatives is corrected. Added regression tests for the actual bug: an older, genuinely never-shown pending message must not get marked surfaced as collateral damage from a poll fetching an unrelated new one, and must survive an unrelated reply afterward; a multi-message poll batch marks surfaced only the message its own preview actually showed, not the rest of the batch. Verified locally: bash tests/fm-telegram.test.sh (47/47), bash tests/fm-pi-watch-extension.test.sh (33/33), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched Python module.
… armed Three consecutive no-mistakes review attempts stalled mid-review on this branch's now-large cumulative diff (no findings table, no gate, 20+ minutes of zero activity and no live process each time - a daemon instability, not a sign of anything wrong with the code). Each stall's partial log still captured real, legible findings before dying; this commit applies the one real bug the third attempt's log had already fully diagnosed before it cut off, rather than losing it to a fourth stall. bin/fm-supervision-lib.sh's fm_supervision_status() decides whether a firstmate home needs its watcher kept armed even with zero in-flight tasks. It already recognized state/x-watch.check.sh (X mode's relay poll shim) as that reason, but never state/tg-watch.check.sh - the identically-armed Telegram poll shim (docs/telegram.md). So a Claude primary with Telegram configured and no other in-flight work never force-armed the watcher at all: bin/fm-claude-stop-autoarm.sh's need_supervision() gate exits before ever reaching the Telegram cadence-sourcing lines this session's earlier commits fixed, making that 30s cadence moot - nothing polled Telegram on an otherwise-idle Claude primary, on Claude itself, the one harness this whole feature is supposed to fully support end to end. This is the same class of gap as the OpenCode plugin's shouldArm() fix earlier this session (tg-cadence-missing-opencode-pi), just in the shared predicate bin/fm-guard.sh, bin/fm-turnend-guard.sh, bin/fm-turnend-guard-cursor.sh, and bin/fm-subagent-pretool-check.sh all consume too - fixing it here closes the gap for all of them at once rather than only the one entrypoint the review's partial log named. Added a regression test mirroring the existing X-mode-poll-need coverage: a Telegram poll shim alone, with zero in-flight tasks, must still keep the auto-arm cycle active. Verified locally: bash tests/fm-claude-stop-autoarm.test.sh (30/30), bash tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-cursor-primary.test.sh (27/27), bash tests/fm-procevent.test.sh (42/42), bash tests/fm-session-lock-ancestry.test.sh (7/7), bash tests/fm-telegram.test.sh (47/47), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean).
A fourth no-mistakes review attempt also stalled before producing a formal findings table (the daemon instability from the prior three attempts persists), but this time completed a full narrative review first, enumerating five real findings in prose before it died. Decided and fixed four of them directly (straight bugs against already-accepted behavior, or naming corrections); documented the fifth as a known, accepted limitation rather than building new concurrency-safety machinery for it. THE MAIN FIX. bin/fm-tg-poll.sh's own fetch success path marked the message it fetched "surfaced" - correct in spirit (only the record actually shown, not every pending one), but wrong in substance: the poll's own summary line previews new_texts[0] truncated to 70 characters, a watcher wake notifying that a message arrived, not a rendering of the message itself. bin/fm-tg-archive.py treats surfaced>=1 as license to retire the record unconditionally on the next reply, however much later - correct for a genuine full-text surfacing (bin/fm-tg-wait.sh's "CAPTAIN: <text>" branch, unaffected by this fix), wrong for a coarse preview the model may never have read past. bin/fm-tg-fetch.py's poll branch no longer marks anything surfaced at all; on Claude, the one harness this ever mattered for, the record still gets a real, full surfacing every turn end via bin/fm-tg-guard.sh/bin/fm-tg-hook.sh's own drain regardless of whether the watcher poll ever ran. OTHER FIXES: - The captain-impersonation drop notice (a chat_id mismatch) was never actually visible through bin/fm-tg-poll.sh on an otherwise-successful poll: fetch.py writes it to stderr and keeps going (rc stays 0), but the poll captured stderr to a temp file and deleted it unconditionally before the success branch ever ran, so the notice - which docs/telegram.md explicitly promises is "visible (never silent)" - was captured and discarded, unseen. The diag file is now surfaced on a successful poll before it is removed. - bin/fm-guard.sh's and bin/fm-turnend-guard.sh's "watcher down" banners hardcoded "X-mode relay polling" as the reason supervision was needed with zero in-flight tasks and zero event sources, the only case that reason was ever written for - now that a Telegram poll shim alone can also be that reason (this session's earlier bin/fm-supervision-lib.sh fix), an idle Telegram-only home got a banner naming the wrong cause. Both banners (and a third, identical occurrence in fm-turnend-guard.sh's exhausted-budget path that the review's partial log did not reach but shares the exact same bug) now name whichever poll shim is actually present. - mark_surfaced()'s `records` parameter shadowed the module-level `import fm_tg_records as records` - dormant (the function never needed the module inside its own body) but confusing. Renamed to `surfaced_records`. - Documented, as a known accepted limitation (not fixed): bin/fm-tg-fetch.py and bin/fm-tg-drain.py each do an unsynchronized read-modify-write on the same inbox record file, so a genuinely concurrent poll and Stop-hook drain can interleave and drop each other's field (surfaced, or an attachment's media path). Closing this needs real cross-process locking across three scripts - a new concurrency-safety mechanism this feature does not otherwise have - for a race whose worst outcomes are one extra surfacing or one orphaned attachment, not a lost message or reply. Updated three existing tests whose assertions assumed the poll's own truncated preview counted as a full surfacing (it no longer does), and added regression tests for the drop-notice visibility fix and the corrected neither-message-marked multi-message-batch behavior. Verified locally: bash tests/fm-telegram.test.sh (48/48), bash tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-guard-stale-banner.test.sh (27/27), bash tests/fm-pi-watch-extension.test.sh (33/33), bash tests/fm-claude-stop-autoarm.test.sh (30/30), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched Python module.
A fifth no-mistakes review attempt stalled again before a formal gate (the same daemon instability), but again completed a full narrative first, catching five more findings - including a real miss in the previous commit's own fix and a bug in that same commit's fix that never actually worked in production. Decided and fixed four directly; documented one as a known, accepted tradeoff rather than risk a change to already-carefully- tuned retry logic. THE MISS. The previous commit removed the poll branch's own truncated 70-char preview from being marked "surfaced" - correct - but missed the identical defect on the far more common path: bin/fm-tg-wait.sh's "CAPTAIN: <text>" branch truncated at 300 characters and still marked the WHOLE record surfaced=1. A message longer than 300 characters got the same treatment: the model could reply from a partial read, and bin/fm-tg-archive.py would then retire the record as though it had been shown in full. Unlike the poll branch, this branch genuinely is meant to be a full surfacing (it is what a Claude Stop hook's own long poll actually shows the model), so the fix here is not to stop marking it - it is to stop truncating short of what a real Telegram message can be. Raised the limit to 4000, matching bin/fm-tg-drain.py's own established ceiling for the identical job; Telegram's own message cap is 4096, so this is a truncation limit in name only. THE BUG IN THE PREVIOUS FIX. bin/fm-tg-poll.sh's fix for the captain-impersonation drop notice wrote it to stderr - but bin/fm-watch.sh's real check runner captures only a check's stdout as the wake text and discards stderr outright, so the fix never actually worked in production. It passed its own test only because that test captured both streams together with `2>&1`, masking which one the notice landed on. Now written to stdout (matching how bin/fm-tg-poll.sh's own record_refusal already correctly does this), and the test captures stdout alone to catch a regression back onto the wrong stream. OTHER FIXES: - The three watcher-down banners that name a poll-only reason supervision is needed (bin/fm-guard.sh, and two sites in bin/fm-turnend-guard.sh) each hand-rolled an identical "which poll shim is present" check, introduced piecemeal across the last two commits. Consolidated into one shared fm_supervision_poll_need_desc() in bin/fm-supervision-lib.sh. - A .tg-poll-diag.* temp file could leak if the watcher SIGKILLed a slow check's whole process group before bin/fm-tg-poll.sh's own cleanup ran. SIGKILL cannot be trapped by anything, so this closes only the half that can be closed: an EXIT trap now owns the cleanup for every exit path this script itself controls, replacing three separate manual rm calls. Documented, as a known accepted limitation rather than fixed: the waiter's backoff treats a 409 (the poller/waiter race's own by-design, expected condition, not a real failure) identically to a genuine unusable reply, so a busy session can climb its backoff toward the 60s cap during entirely normal operation. Exempting 409 specifically needs the waiter to also capture and inspect fm-tg-fetch.py's stderr reason, which it currently does not; that is a real, scoped fix, just not one to make to carefully-tuned retry logic - already reshaped by several captain-observed production incidents this session - without the live traffic needed to verify it does not reopen the CPU-spin failure the backoff exists to prevent. Added a regression test for the waiter's 300-character truncation fix, and corrected the drop-notice test to capture stdout alone rather than both streams together, so it can no longer pass while the real fix is broken. Verified locally: bash tests/fm-telegram.test.sh (49/49), bash tests/fm-turnend-guard.test.sh (67/67), bash tests/fm-guard-stale-banner.test.sh (27/27), bash tests/fm-claude-stop-autoarm.test.sh (30/30), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched Python module.
A sixth no-mistakes review attempt stalled again before a formal gate (the same daemon instability across every attempt on this branch), but again completed a full narrative first, catching three findings. Fixed two directly; escalated the third as a genuine architecture question rather than deciding it myself. THE NEAR-MISS. The previous commit raised PRINT_MAX from 300 to 4000 to fix a truncated-preview-marked-as-fully-shown defect, with a comment claiming 4000 "comfortably exceeds Telegram's own 4096-character message cap." It does not - 4000 < 4096. A 4001-4096-character message (one Telegram genuinely delivers) was still truncated and still marked surfaced=1, the identical defect narrowed but not closed. Raised to 4096, Telegram's own documented cap, in both bin/fm-tg-fetch.py's PRINT_MAX and bin/fm-tg-drain.py's own separate, less-severely-affected 4000 (same class of near-miss, same fix). The regression test that exercised only a 500-character message (which 4000 already covered, so it could not have caught this) now uses a full 4096-character message. THE 429 DEDUP BUG. state/.tg-poll-error dedups on the literal refusal text so one standing failure is reported once, not every 30-second cycle - except Telegram's own 429 description is "Too Many Requests: retry after N," where N counts down every poll, so the literal line never repeats. A sustained rate limit therefore woke firstmate on every single cycle, the exact failure this record exists to prevent, for the one refusal reason most likely to actually be sustained. record_refusal now dedups on the line with only that one known-volatile "retry after N" tail normalized out; the leading error code is left untouched, so a genuinely different refusal (a different code) still reports as new. ESCALATED, not decided: harness_can_surface() resolves the harness of the calling process's own tree, which for the watcher's poll path is the long-lived watcher process armed once at bootstrap - not necessarily the harness currently running as the home's primary right now. Nothing ties a watcher restart to a harness switch on the same FM_HOME (confirmed: bin/fm-bootstrap.sh only ever suggests a manual restart on a cadence-config change), so a home switched from Claude to Codex without an intervening watcher restart could keep sending acks nothing will ever surface - the exact false promise this session already closed for the common case, just reopened at an edge the review is right to have found. This is a genuine architecture question (accept a narrow documented edge case, or build a watcher-restart-on-harness-mismatch mechanism - new durable state and a new bootstrap-time check, not a bug fix) rather than a defect against already-accepted behavior, so it is escalated per the ask-user-authority framework: data/fm-telegram-ship-t8/needs-decision-round9-harness-identity-source.md. Added a regression test for the 429 countdown-dedup fix (a sustained 429 across two polls with different countdowns must report once; a genuinely different error code afterward must still report). Verified locally: bash tests/fm-telegram.test.sh (50/50), shellcheck bin/*.sh bin/backends/*.sh tests/*.sh (clean), python3 -m py_compile on every touched Python module.
…ning The captain's decision on the harness_can_surface() identity-source escalation: decline building a watcher-restart-on-harness-mismatch mechanism (new durable state to cover a rare manual reconfiguration is not worth the complexity), but make the documentation actionable rather than merely honest - a reader hitting this needs the fix in front of them, not just a warning that it can happen. docs/telegram.md's "Inbound is Claude-only" section now states the edge case plainly and gives the one command that closes it: restart the watcher (`bin/fm-watch-arm.sh --restart`) after switching a home's primary harness. Restarting re-arms the watcher under the new primary's own identity, so harness_can_surface() stops reflecting a stale one immediately.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves the Telegram captain-comms machinery out of the gitignored state/ directory and into tracked bin/, so a configured clone works without hand-holding.
Includes the four guarantees measured against real failures tonight: only firstmate contacts the captain (crew sessions are refused), no message is retired unanswered, no message is answered twice, and the '...' acknowledgement does not count as a reply.
Also carries an upstream catch-up (16 commits) picked up during the work.
Validation note, stated plainly: the no-mistakes pipeline could not complete. Six consecutive review attempts stalled before emitting a findings table, each time after 20+ minutes of no filesystem activity and no live process. Every finding that any partial run did surface was fixed by hand rather than lost. Directly verified instead: fm-telegram 50/50, fm-pi-watch-extension 33/33, fm-supervision-instructions 10/10, shellcheck clean.