feat(bin): reconcile 18 upstream changes as a recorded merge - #37
Merged
Merged
Conversation
…as a proven empty composer (kunchenguid#4455) * fix(composer): accept Grok title overhang * no-mistakes(review): summary: named Grok overhang constant, doc caveat, restored tmux typed-title coverage
…ailure (kunchenguid#4474) * fix(bin): recover Claude auto-arm after timeout * no-mistakes(document): Add host-timeout signal coverage to autoarm test-coverage list
* fix(spawn): establish Claude task channel authority * no-mistakes(document): Document Claude task-worker control-channel trust in harness-adapters reference
…or pending text (kunchenguid#4458) * fix: guard relaunch exit against pending input * no-mistakes(review): Verifying test run in progress * no-mistakes(document): docs(agent-control): document exit's composer-empty fail-safe guard * no-mistakes(ci): fixed 2 tests broken by approved do_exit fail-safe change (empty-only composer gate). herdr-smoke test's sleep-stand-in never renders a real composer -> updated assertion to expect "not proven empty" refusal instead of stale "did not stop" msg. secondmate-restart fake tmux capture-pane returned bare '> ' glyph (never valid empty proof) -> changed to bordered empty box matching fm-control-relaunch fixture. all 4 related suites pass locally now
…unchenguid#4460) * fix: reconcile diverged secondmate updates * no-mistakes(document): Fix stale fm-update.sh/fm-ff-lib.sh purpose lines in docs/scripts.md * no-mistakes(document): docs: reflect secondmate divergence reconcile in README/SKILL.md
…d#4497) * fix(dispatch): support Codex Luna max effort * no-mistakes(review): use portable CODEX_HOME path in codex effort reference
kunchenguid#4498) * feat(calm): render smooth Unicode swell * feat(calm): make sails asymmetric * feat(calm): use quarter sail glyph * no-mistakes(review): docs: sync calm feasibility sprite passage with approved renderer * no-mistakes(document): docs: sync calm wave phase doc comment * no-mistakes(ci): CI の Lint 失敗は tests/fm-calm-pi-extension.test.sh の test_interactive_terminal_e2e 関数で `boat_narrow_sails` が local 宣言に残っていたことによる ShellCheck SC2034 でした。関数内での参照を確認したところ、狭幅端末の検査は boat_narrow_previous / boat_narrow_direction / boat_narrow_reversed に移行済みで、boat_narrow_sails は代入も参照も一切ありませんでした。そのため local 宣言からこの 1 語のみを削除しました(3315 行目)。Calm の描画実装、他のテストアサーション、ドキュメントは変更していません。検証: bin/fm-lint.sh(ローカル変更ファイルモード)exit 0、CI 相当の `shellcheck --norc --external-sources tests/fm-calm-pi-extension.test.sh` exit 0(SC2034 解消)、`bash -n` 構文チェック通過、actionlint 1.7.12 でワークフロー 3 件 valid。
kunchenguid#4491) * fix: supersede scout delivery brief on promotion * fix: preserve ship safety contract after promotion * no-mistakes(document): Document fm-promote.sh now supersedes brief.md on relaunch
…d stop cleanup dropping accents from a held body (kunchenguid#4471) * fix(bin): let captain holds work on hosts with an older JSON::PP Holding a task for the captain, and the cleanup that keeps a captain-held row open, both fail outright on any host whose JSON::PP defaults allow_nonref off - 2.27202 on a Linux desk is one. Both read a task's body back with `decode_json`, but tasks-axi shows a scalar field as a JSON-encoded bare string, and an older library rejects that whole value with "must be object or array". The consequence is fleet-wide on such a host, not one broken command: a worker there cannot formally record a decision for the captain at all. It can only mention the decision in passing in a status line, where it can be missed - which is how a real decision goes unrecorded. The hold reports that the task lost its hold-set stamp; the cleanup cannot return the row to Queued. Both call sites now ask for allow_nonref explicitly rather than inheriting whatever the installed library defaults to. The second one is worth naming: its `/\A"/` guard reads as deliberate, but a leading quote is exactly the bare-string case that fails, so the guard selects for the failing input rather than protecting against it. The regression case forces the older default back off for every perl the commands spawn, then drives both paths - holding a task that carries a body, and tearing down a captain-held row whose deliverable must still be appended. It also probes that the simulation genuinely rejects a bare scalar, so the case cannot pass vacuously on a lenient host. Each half was verified failing on its own unfixed call site with that site's real error message. Suites: fm-captain-hold-lifecycle 51 cases, fm-backlog-atomicity 99 cases, 0 failures. Verification limit: the mechanism is reproduced and tested, but neither fix is verified against a real JSON::PP 2.27202 host, because none is in the loop. This laptop runs 4.06, where the bug does not manifest. `bin/fm-procevent-lavish.sh:471` was checked and left alone - it matches a brace-delimited object before decoding, so allow_nonref never applies. * fix(bin): stop cleanup silently dropping accented characters from a held body Cleanup rewrites a captain-held row's body to append the finished work's deliverable, and the decoder it reads that body with printed decoded characters to a stream with no `:raw` layer. A character at or below U+00FF then came out as one latin-1 byte instead of two UTF-8 ones, so a body reading "café" lost the accent. `fm_backlog_retain` writes that body straight back through `--body-file`, and nothing reported an error - the character was simply gone from a row still waiting on the captain. The decoder now writes bytes, the same `binmode STDOUT, ":raw"` plus `utf8::encode` that the sibling decoder in `bin/fm-captain-hold.sh` already used. Review of the parent commit found this on one of the lines that commit already changed. It predates that change. The test asserts bytes rather than decoded strings, because comparing strings cannot tell latin-1 from UTF-8. It uses two separate rows on purpose: any character above U+00FF makes perl print the whole string as UTF-8, so one body carrying both an accent and an em dash passes even unfixed and proves nothing. Verified failing before the fix on the accented row, passing after. Suites: fm-captain-hold-lifecycle 52 cases, fm-backlog-atomicity 99 cases, 0 failures. * no-mistakes(document): record body-decode regression proofs in captain-hold lifecycle doc * no-mistakes(review): drop whole-file UTF-8 check from retained-body test * no-mistakes(review): correct stale JSON::PP fleet-host claim in lifecycle doc * no-mistakes(review): anchor native-reproduction claims per defect in lifecycle doc
…furniture (kunchenguid#4532) * fix(composer): read codex 0.154's idle starfield and status footer as furniture codex-cli 0.154.0 animates a braille "starfield" around its idle composer: on the row above the bold `›` prompt row, on the `›` row behind the SGR-2 dim `Ask Codex to do anything` placeholder, and on the row below it, then draws a bright status footer (`<model> <effort>[ fast] · <path> · <title>`). The cells are truecolor greys on both sides of the ghost luminance ceiling, so the brighter ones survive ghost stripping, and the rows below the glyph carry no structural edge. The shared classifier selected the bare `›` shape, extended its wrap region over the two rows beneath the glyph, read the survivors and the footer as wrapped typed input, and answered `pending`; the steering doorbell defers on exactly that verdict, so no doorbell ever reached an idle codex 0.154 pane. bin/fm-composer-lib.sh now recognises that furniture by shape, declared once next to the idle placeholders and reached from the two wrap-region boundary points: - a row whose non-whitespace content is entirely braille cells (U+2800..U+28FF, detected byte-exactly under LC_ALL=C) is furniture: it never counts as wrapped typed content and bounds a bare composer's wrap region; braille behind the glyph row's content is stripped before the emptiness decision when nothing else follows the glyph; a row mixing braille with other text stays typed content; - the codex status footer bounds the wrap region exactly as omp's status row does, anchored on the effort token, a spaced middle dot, and a `~` or `/` path cell, so a typed `fix · tests` stays composer input; - `^Ask Codex to do anything$` joins the verified idle-placeholder set; the ghost strip remains what proves that row empty, and the bare-row rule that bright placeholder text is real input is unchanged. Unchanged: the strict blank-row rule, the styled=0 degradation (a plain cmux/orca capture of this screen still reads `unknown`, never `pending`), FM_COMPOSER_GHOST_LUMA_MAX, and every other harness's shape. tests/fm-composer-lib.test.sh carries both live Herdr samples byte-for-byte with the divergence (letters in place of the starfield read `pending`) and the over-stripping negatives; tests/fm-composer-codex-idle-live-e2e.test.sh is the default-on live guard (token-free, skips explicitly without codex or tmux) that launches the installed codex idle and asserts `empty` through both the tmux and the cursorless styled reads, naming codex --version on failure. docs/verification/runtime-backends.md records the dated Herdr evidence: `pending` before, `empty` after, on the captured screen. * no-mistakes(review): drop unreachable codex footer rule and inert placeholder entry --------- Co-authored-by: Todd Billings <todd@usdvcapital.com>
* fix(bin): refuse empty text steers in fm-send A marked secondmate request sent with an empty message delivered only marker and correlation bytes and minted a pending-reply expectation the parent could never see resolved, stalling the fleet with no loud error (kunchenguid#4255). Fail closed on an empty or whitespace-only message on the text path, mirroring the existing --resolve-key refusal. * chore: retain ambient Pi-lens autoformat as its own commit Formatting-only edits produced by ambient Pi-lens autoformat during the msg-loss investigation, kept separate from the behavioural change in c23acba so the fix stays reviewable on its own. AGENTS.md is deliberately excluded: its only autoformat edit stripped the trailing space from the documented FM_OPERATIONAL_PREFIX value, which bin/fm-operational-input.sh:28 defines as "FIRSTMATE_OP: " and line 11 records as permanent compatibility. Documenting that constant without its trailing space makes the doc wrong about the contract, so that one line was restored rather than retained.
…chenguid#4554) On rose-pine-moon the two-color water (cyan crests over blue troughs) read as a pink stripe over aqua, the yellow left sail and mast clashed with the red right sail, and the hull carried a blue interior run. Every water cell is now blue so the swell reads through glyph height alone, and both sail halves, the mast, and the whole hull are one yellow run. Geometry, cadence, animation, direction flip, resize clamping, and the narrow fallback are unchanged. Update the unit and real-TUI color assertions to the new palette and the Calm docs that described the old one.
…chenguid#4270) * fix(watch): stop aging a second mate's active turn from its launch The parent watcher's second-mate wake-loop stall check exempts a mate that is demonstrably inside an active turn, but secondmate_in_active_turn asked busy_turn_over_age first and returned "not in a turn" whenever that said the bound was crossed. busy_turn_over_age ages from state/<task>.turn-ended, falling back to state/<task>.meta. A second mate's turns end in its own home, so the parent never gets a turn-ended mark for it and the fallback ages the mate's last launch. Every mate launched more than BUSY_TURN_MAX_SECS ago was therefore permanently "over age", the busy pane was never consulted, and any turn outstripping FM_SECONDMATE_WAKE_STALL_SECS raised a false wake-loop stall. The gate now bounds the busy exemption by <idle> - how long the queue's drain position has not moved - which is evidence this home actually holds. A busy mate stays exempt while the queue has been frozen for less than BUSY_TURN_MAX_SECS, and a mate stuck busy forever still alarms, so the bound that stops a busy pane from proving liveness forever is kept rather than removed. busy_turn_over_age is untouched; its remaining callers are the ordinary crew busy-pane bound. The regression pins the case that actually broke: a mate whose launch record predates BUSY_TURN_MAX_SECS and which is demonstrably mid-turn must not escalate, while the same mate with its queue frozen past the bound still publishes exactly one notification. The existing coverage only exercised a freshly launched mate, which passes either way. Reaching that alert now costs a pane capture inside the gate, so the three checkpoints in this suite that assert an alert move from a 1s to a 4s bound - the value the neighbouring active-turn cases already use. The bound is a ceiling, not a wait: the checkpoint returns on the first actionable wake. On a loaded machine a 1s bound missed the alert repeatedly; at 4s it did not miss in 20 runs under the same load. * no-mistakes(review): scope the second-mate active-turn regression test's coverage claim * no-mistakes(document): fix stale second-mate active-turn comments in fm-watch
…unchenguid#4278) * feat(bin): add read-only PR blocker and reviewer-discovery commands Two focused, opt-in commands that read GitHub and never write to it. fm-pr-state.sh reports what still blocks one pull request from the author's side: a closed or merged state, draft state, unknown or conflicting mergeability, absent or failing required checks, and a blocking CHANGES_REQUESTED decision explained by each reviewer's latest verdict, marked STALE when it was left at a superseded head. A pull request that only awaits an approval is not reported as blocked, and advisory checks are omitted. Every reading is taken against one exact head; a push that lands mid-read invalidates the whole result rather than mixing two snapshots. fm-pr-reviewers.sh suggests reviewers from the most recent commits to the pull request's exact changed paths, counting each commit once, resolving handles through GitHub's own commit author.login mapping, and excluding the author and Bot accounts. Both stay read-only: no review request, no approval, no merge. Unresolved review-thread state is left unreported because the REST API does not expose it and unattended commands may not use GraphQL. Closes kunchenguid#3731 * no-mistakes(review): accept only PR URLs and stop at terminal state * no-mistakes(review): report unconfirmed required checks; make URL-only guards discriminate * no-mistakes(review): stop attributing readings to unverified heads * no-mistakes(review): narrow readiness contract to checks that have reported * no-mistakes(review): read the pull request once, drop the head guard * no-mistakes(document): scope pr-forge isolation proof to its measured members * no-mistakes(document): record uncovered pr-forge members and their pending proof * docs(isolation-proof): re-prove pr-forge at its full membership tests/fm-pr-state.test.sh and tests/fm-pr-reviewers.test.sh joined the pr-forge family in this branch, and script_allows_concurrency grants four workers by family membership alone, so both ran concurrently on a proof measured before they existed. Re-proved the family at all eight members: two consecutive runs, 0 failures, each begun with the one-minute load average below 6.0 so the result measures isolation rather than contention. A third run taken between them is disclosed rather than recorded, because it started while the previous run's workers were still decaying. The new durations are not comparable with the six-member measurement above them, so they are not presented as evidence about the two new members, and that record's 1.72x four-worker figure is left as a statement about its own run rather than restated as current. * no-mistakes(review): disclose gh error-text coupling at its matching site and tests
…uid#2752) * fix(bin): teach validation-round pauses in briefs * no-mistakes(document): Point classifier comments to authoritative pause examples
…guid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree
* feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, #5769f7 light) and the whole boat takes the Claude orange of the stock spinner (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence
…kunchenguid#4586) * fix(watch): honour a declared wait before wedge-escalating a quiet pane wedge_timer_check escalated on elapsed idle time alone. Nothing asked whether the worker had already said why its pane was quiet, so a lane that declared a bounded external wait climbed the escalation ladder for as long as the wait lasted, and past FM_WEDGE_DEMAND_INSPECT_COUNT every repeat carried demand-deep-inspection - which by its own wording forbids re-absorbing on the run-step or pane state, so the supervisor could not use the evidence that was there either. The generated brief promises that declaring `paused:` buys the long recheck cadence instead of a wedge, but the timer was still reachable while that declaration stood: a crew that declares a wait and then has an active run or busy pane attributed to it is handed to the timer as provably-working. The declaration is what the worker said about its own silence, so it now outranks a liveness verdict that only says something is running. The consult runs in the at-threshold branch that was about to escalate, beside the worktree walk already there, and costs one status-line read. Either status-line record defers to the same FM_PAUSE_RESURFACE_SECS recheck the declared-wait absorber already uses, so the wait is still rechecked and cannot rot invisibly. Which verb declared it decides the wording, because the two block on different people: a `paused:` wait is owed by an external dependency and asks the reader to confirm it still holds, while a `captain-held:` transfer is owed by the captain reading the recheck and asks them to answer or release the hold. A hold is not rechecked at all while the away-posture record exists, as on every other captain-held path, and that absorb arms no throttle so the recheck is owed in full on return. A declared clearing time that has already passed stops counting, and a lane that never declared one keeps the identical escalation schedule, reason, count and demand-deep-inspection wording, so detection and its worst-case time are unchanged. The deferral restarts the idle timer rather than cancelling it, so a lane that stops waiting escalates again within one threshold. A lane quiet because its own validation run is parked at a gate awaiting a human decision is deliberately out of scope: reading that state needs a signal carrying who the wait is on and what clears it, rather than one inferred from a parked verdict that also covers gates awaiting the crewmate itself. Tests pin both directions for each case and were each confirmed to fail with the consult removed. * no-mistakes(document): docs: honour declared waits in stale-escalation docs
* fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library
kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower
…nguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks
* Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check`
…#4656) * fix(bin): re-record PR poll identity after a volume device renumber (Fixes kunchenguid#4260) A volume remount can renumber the state filesystem's st_dev while every inode and byte stays the same; APFS does this across a reboot. A poll registration records its sidecar and check as device:inode, so every poll armed before the remount failed strict validation and the watcher refused all of them as unauthenticated state checks until each was re-armed by hand. There are two device comparisons. fm_pr_private_file_valid compares a live file's device with the state directory's device read in the same invocation: it refuses a file that is not on the state directory's own filesystem and already survives a renumber, so it is unchanged. The registration's recorded identity versus the live identity (from kunchenguid#556, reused by the kunchenguid#932 retirement receipt) binds the registration to the exact files published in its own transaction; its device part is what breaks. When strict capture fails, the watcher now proves the device is the only difference: every other artifact check passes (template bytes, both hashes, private mode, single link, live device, metadata), both recorded identities name one device, and each recorded inode equals its live inode. Only then, under the task's control lock, does it rewrite the two identity lines, repeating the whole proof and comparing the registration's file identity and bytes just before the rename, and then capture strictly again. A swapped, altered, re-moded, relinked, split-device, or foreign-device artifact still fails a proof and is still refused, and a pending retirement receipt blocks the rewrite. Reproduction: on macOS a poll armed on an APFS disk image that was detached and re-attached behind another image moved st_dev 16777239 -> 16777243 with inodes, bytes, mode, and link count unchanged; the real watcher refused it on main and reports its merge with this change. The portable regression test rewrites a real registration's recorded device and drives the watcher. Not changed here: the status presentation cursor keys rows by its own device:inode identity in bin/fm-classify-lib.sh, a different helper that needs its own fix; a retirement receipt left by a reboot between its publication and removal still names the old device and stays refused; custom check trust binds only a content hash and is unaffected. * fix(review): Serialize PR poll publication writers * fix(review): Bound PR poll publication lock scope
…llow-up to kunchenguid#4627) (kunchenguid#4661) A budget that expires partway through an observation no longer records an error or prints the unavailable wake; the URL keeps its prior record and is observed first next poll. forge() flags budget exhaustion at the point it refuses, or when a read is killed at the budget's own deadline, so a genuine forge failure still records the error and wakes. Each distinct URL is now observed once per poll and applied to every owning task.
…kunchenguid#4680) * fix(bin): clear parent pending-replies on local secondmate retirement Local secondmate teardown left resolved parent pending-reply records behind after home removal (seen after papa-hdds / pxmx retirement). Refuse non-forced retirement while any reply for that id is still unresolved, and delete every matching record plus its delivery confirmation after a successful local or remote retirement, matching the remote cleanup path. * no-mistakes(document): Align secondmate retirement docs with pending-reply cleanup * no-mistakes(review): Lokale Pending-replies-Sicherheitsprüfung vor Home-Entfernung * no-mistakes(review): Pending-replies-corr_id auf 16-Hex absichern * no-mistakes(review): Pending-replies Basename und corr_id abgleichen * no-mistakes(document): Clarify forced retirement pending-reply cleanup --------- Co-authored-by: ladwein <ladwein@firstmate.bost8.thelad.loc>
kunchenguid#4677) * fix(bin): accept Orca's composite worktree id at teardown Teardown refused every Orca-backed task because the endpoint validator checked orca_worktree_id with the simple-atom rule meant for tmux-style window names, which rejects any character outside [A-Za-z0-9._@%+-]. Orca returns that id as `<orca id>::<absolute worktree path>`, so the colon and slashes in every real value made validation fail and finished Orca tasks could never be cleaned up. Validate the field as the composite it is: both halves of the first `::` split present, the path half absolute, and no embedded newline, carriage return, or tab. The terminal field keeps the atom check, which is correct for it, and no other backend's validation changes. The existing Orca fixtures recorded ids like `wt-teardown`, a shape Orca never returns, which is why the suite passed a check the real value fails. They now carry the composite form, so the tests exercise the real value. * no-mistakes(document): name Orca's repo id in the composite worktree id * no-mistakes(document): list teardown endpoint safety suite in Orca regression entry points
* feat(bin): add opt-in typed dispatch resolution through typesafe.ai Add bin/fm-dispatch-resolve.sh, which resolves one concrete crewmate or scout profile from a written brief with typesafe.ai's System One model: one Choice question over the rules' `when` texts, then the confidence floor, the rule's `approval` and `floor`, each profile's `provider` and `floor`, one quota-axi snapshot, and the spendPriority argmax all in code. It is off unless TYPESAFE_API_KEY is in the environment or the home's gitignored .env; off means one stderr line, exit 0, and no network call, so firstmate dispatches exactly as before. The key reaches curl on a file descriptor, never argv. Extract fmx_env_get into bin/fm-env-lib.sh as the one .env accessor and the harness-to-provider table into bin/fm-quota-axi-lib.sh so the new tool and bin/fm-quota-choose.sh share one owner each. Bootstrap validates the four new optional dispatch fields. Document the schema, the operator contract, the AGENTS.md intake step, and the live and benchmark evidence. * no-mistakes(review): Harden typed dispatch resolution and quota bounds * no-mistakes(review): Validate dispatch floors and ranking evidence * no-mistakes(review): Tighten dispatch response and floor evidence * no-mistakes(review): Neutralize none matching and resolve defaults locally * no-mistakes(review): Preserve providerless profiles outside typed resolution * no-mistakes(review): Validate response usage and reject duplicate profiles * no-mistakes(review): Escalate unverifiable floors and validate probabilities * no-mistakes(review): Validate probability mass and unknown profile floors * no-mistakes(review): Simplify resolver interface and preserve fallback routing * no-mistakes(review): Fix constants and rank partial quota evidence * no-mistakes(review): Add authoritative provider mapping and enforce explicit providers * no-mistakes(review): Declare provider for documented Pi profile * no-mistakes(review): Validate provider identifiers and support Gemini dispatch * no-mistakes(review): Strictly anchor provider identifiers * no-mistakes(review): Validate selectors and preserve fallback candidate evidence * no-mistakes(review): Gate typed validation and harden resolver evidence * no-mistakes(review): Preserve opt-in routing and harden candidate evidence * no-mistakes(review): Prioritize known exhaustion over quota uncertainty * no-mistakes(review): Isolate API secrets and preserve no-key diagnostics * no-mistakes(review): Fallback safely when dispatch rules are absent * no-mistakes(review): Prioritize quota vetoes and isolate bootstrap secrets * no-mistakes(document): Document typed dispatch safety and fallback behavior
…nguid#6179) * fix(tests): cut the fixed sleeps in supervision-host cycles The serial CI lane keeps brushing its 30-minute cap because fm-supervision-host.test.sh spends ~903s of the job, and per the run-36635306527 case profile the top nine cases are all multi-cycle ones (3-10 park/close/turn cycles each): every close waits out the host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan cycle, and every engine turn waits out the fixed sleep 1 descendant snapshot. That is ~3s of pure sleep per cycle before any real work. The host poll now accepts positive decimal seconds through a new seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a positive decimal defaulting to one second - the smallest seam at each wait's single owner. The suite drives them at 0.2 alongside the existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real poll loops still run. The park-boundary case moves onto the injected test clock instead of a real 3s wait, per-case cleanup polls the host pid rather than sleeping a full second, and the proof-by-absence windows (flood re-escalation, successor re-announce, watcher persistence, recovery staying off main) shrink from 2-3s to 1s, which still spans two watcher polls at the test cadence. Every assertion, process lifecycle, and reaping path is unchanged; production defaults stay at one second. Isolated case timings on a contended host, base vs branch: attended-latch 54.3->34.6s, undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence 47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s, registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s, latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and shellcheck clean. * no-mistakes(review): Wait for scan lock release before duplicate check * no-mistakes(document): Correct supervision snapshot cadence documentation * fix(tests): keep production poll cadence, probe exits at 0.1s The fractional poll cadences multiplied the cost of each loop body: full process-table scans in the engine turn and process refreshes in await_close ran five times more often, which swamped the thin CI runner and nearly doubled every multi-cycle case (serial 5 was cancelled at its 30-minute limit on run 36635306527's successor). Restore the production cadence and notice arm/engine exits with a cheap kill -0 probe at a tenth of a second between the one-second bodies instead: strictly less dead time than baseline with no added CPU. Also hold each injected-clock park bound well past its case's wall-clock checks so a host that ignored the test clock fails instead of silently passing at a real-time boundary, and restore the shortened proof windows (watcher liveness, recovery-off-main absence, first-cycle stream) to their baseline depth. * no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192) * fix: rebalance portable CI from current duration measurements * no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup * no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216) * fix(bin): run no repository hook when core.hooksPath is empty The per-task hook wrapper refused every commit in a repository whose own config sets core.hooksPath to the empty string, because git rev-parse --git-path hooks fails on it. Plain git reads that setting as no hooks, so the wrapper now runs none; every other lookup failure still refuses and shows git's error. Fixes kunchenguid#6171 * no-mistakes(review): Refuse commits when core.hooksPath is a valueless key * no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs * no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213) * fix(bin): let a stale record on a reassigned slot retire records-only When a pool slot's owner claim names another task, the stale record's teardown touches nothing under the slot, so the exclusive-slot record scan no longer refuses it. Full teardowns of a slot this task still claims, or one with no claim, keep the refusal. Fixes kunchenguid#6184 * no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240) * fix(bin): keep the steering doorbell short under deep homes The doorbell printed the task inbox's absolute path twice, so under a deep home it grew to about 290 characters and a Herdr submit reported it never reached the pane on every re-ring. It now names the inbox once by its short <task>.inbox name and points at the full path the worker's brief already gives, so its length no longer depends on the home's depth. Fixes kunchenguid#6120 * no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell * no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
… no turns (kunchenguid#4859) * fix(dod): drive no-mistakes with one foreground call, not a background poll The brief told workers to background the drive call and poll `axi status` because one call "routinely outlives what your harness lets a single command run". That advice contradicts the tool it drives: `no-mistakes axi run --help` documents `--wait` with an 8m default, existing precisely "so an agent harness with a 10-minute tool cap gets a structured return instead of an unbounded hang". Following the old text, a worker could never idle - a backgrounded call returns in milliseconds, so it does not wait at all - and each attempt leaked a live timer that later fired as a paid wake. Tell workers to make one foreground call, let it block, and repeat it when it returns on elapsed wait rather than on a gate or outcome. Also drops the generalisation that told workers on any unestablished harness to assume a command cap and use the same shape, which exported the defect to harnesses with no such cap. * fix(bin): let a waiting worker spend no turns until it is answered A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot kept taking model turns: the brief told it to list its inbox at any natural checkpoint, and six automatic senders nudged secondmates whatever their open decisions. - The ship and scout briefs gain one Waiting section: end the turn after needs-decision or blocked, and hold an external wait inside ONE blocking command bounded by the harness's own command ceiling. The checkpoint clause is deleted. Forbidding the wrong shapes is not enough on its own, so the section also names the blocking foreground `until` loop as the wait a Claude Code worker may use, because that harness can refuse a sleep-then-check command while pointing at backgrounding, which is the one shape a waiting worker must not take. - fm-send --automatic defers (exit 4, nothing written or rung) while the target has an open decision or blocker of its own; every automatic sender passes it and keeps its retry state, and the pending-reply recovery waits the same way. - The two senders that report the result classified it by matching the text of the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh as a supervision warning, and that guard prints its worktree-tangle banner whenever the primary checkout is on a feature branch, which is exactly what a CI pull-request checkout is. The banner lands ahead of the `deferred:` line, so the match fell through and a waiting mate was reported as a failed send, with the banner as the reason. Both senders now classify on fm-send's exit status, which is the contract the deferral is actually stated in, and select the `deferred:` line out of the output rather than assuming it came first. The third root cause, a no-mistakes definition of done that backgrounded the drive call and polled axi status, is fixed by this branch's parent commit "drive no-mistakes with one foreground call, not a background poll"; this commit takes that text as is and adds the regression test. Upstream's spawn abort path no longer calls the lease-return helper at all, so the fork's missing-helper guard and its pin-feature test line are moot here and are not ported. The command ceilings each harness enforces, and the probes behind the named Claude Code wait, are recorded in docs/verification/runtime-backends.md. * no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses * no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery * no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more * no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up * no-mistakes(review): Hold automatic wakes until a mate's own decision closes * no-mistakes(document): Document watcher delivery of deferred remote re-read nudges * no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD * no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture * no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions * no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI * Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI" This reverts commit c719928. * no-mistakes(review): Retry deferred local instruction nudges via the watcher * no-mistakes(review): Document watcher retry for deferred local instruction nudges * no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean * Pin autoarm supervision model in secondmate restart T3b The fresh watcher beat the test writes proves a live watcher only under the autoarm model; on CI hosts with no detected harness the persistent model demands a lock-holding watcher, so the watcher-down banner became the reported reason and the deferral assertion failed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep deferred secondmate nudges retryable under the inheritance lock. A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes. * no-mistakes(document): Document watcher retry of deferred restart re-read nudges * Send secondmate reread and restart nudges immediately again. Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main. * Make the no-turn wait opt-in behind config/wait-no-turns. Homes that do not create the file keep the previous briefs, drive text, and sends. * no-mistakes(document): Document wait-no-turns inbox wording change in configuration * no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting * no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording * no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…note (kunchenguid#6140) * fix(bin): record Gerrit change URLs as close notes Teardown's backlog_done_args hands every ship's recorded pr= URL to fm_backlog_done as --pr, and tasks-axi refuses any --pr that is not a canonical GitHub or Forgejo pull request. A Gerrit change URL therefore left the item In flight after cleanup, and the pending backlog-close record replayed into the same refusal at every session start. fm_backlog_done now rewrites a --pr whose value fm_pr_url_parse reads as a Gerrit change into --note "Gerrit change <url>". The mapping sits at the tasks-axi call rather than in the pending-close record, so records already written with --pr replay to a close unchanged. The captain-held retain path records the URL in its deliverable line and skips the update --pr it cannot make. * no-mistakes(review): Note retained Gerrit change URL when captain answers early * no-mistakes(document): Document Gerrit change URL handling in captain-hold retention
* perf(remote): separate active job sampling from dispatcher cadence * no-mistakes(document): Link remote wait timing to its authoritative contract * no-mistakes(ci): Fixed ci-1 with two narrowly scoped SC2030 annotations documenting intentional subshell-local legacy and active cadence overrides in tests/fm-remote-job.test.sh. Runtime behavior is unchanged. Reproduced the lint failure before the fix; afterward ShellCheck 0.11.0 with source following, Bash syntax validation, the complete remote-job behavior suite, and git diff --check all passed * perf(supervision): reduce park, delta and dispatcher polling * no-mistakes(document): Clarify poll latency contracts and authoritative documentation pointers
…henguid#6221) * fix(bin): load backend sibling libraries under zsh fm_backend_source kept each backend's sibling list in one space-separated string and iterated it unquoted. zsh does not word-split an unquoted expansion, so the readability check saw the whole list as one path and refused every backend with more than one sibling. Hold the list in the function's positional parameters instead, which needs no word splitting in Bash 3.2, Bash 5, or zsh. The existing zsh case in tests/fm-backend.test.sh covers it wherever zsh is installed. * test: run the Calm mod suite on stock Bash 3.2 The suite injected shell values into its generated Node scripts with the ${value@Q} transformation, which needs Bash 4.4. Stock macOS Bash 3.2 reports a bad substitution, so every case failed before it asserted anything. Build each JavaScript string literal with JSON.stringify through a small helper instead, which works on any Bash and is a valid literal for any value. * no-mistakes(review): fix(bin): rename zsh-special path local in fm_backend_source * test: narrow the zsh backend claim to name matching Under zsh the adapters locate their siblings through BASH_SOURCE, so a successful fm_backend_source is not a full load. Assert only what the contract states, and pass js_string values after -- so node never reads a leading-dash value as its own option. --------- Co-authored-by: Nova Agent B <novaagentb@gmail.com>
…tatus scans (kunchenguid#5263) * fix(bin): exclude a remote mate's own parent channel from self-home scans A remote secondmate home's outbound parent channel lives at state/parent-replies.status inside its own state dir, so the watcher's signal scan enumerated it as a task status file and the open-decisions fold classified it as a phantom task named parent-replies: every parent-channel append spun a spurious signal wake and a phantom open decision in the mate's own home. fm-parent-channel-lib.sh gains fm_parent_channel_outbound_status, which resolves the channel into the mate's own state dir for the remote route only, and fm-classify-lib.sh's status_scan_parent_channel_exclude wraps it for the fleet-wide scans. The watcher's scan_signals and heartbeat fail-safe backstop, the whole-file and incremental open-decisions folds, the presentation snapshot, and the unread-surface scan now skip exactly that resolved path. The exclusion is home-shape-aware: a parent-replies.status in a main home or a local mate is an ordinary task log and keeps waking and folding, and every other status file is untouched. * no-mistakes(review): exclude a remote mate's parent channel from the daemon heartbeat scan * no-mistakes(document): Document remote mate parent-channel scan exclusion * ci: retrigger portable serial 4 * no-mistakes(ci): CI check 'Behavior portable serial 7' failed in tests/fm-contributions.test.sh ('reservation poll failed'). CI stderr showed bin/fm-contributions.sh:345 arithmetic 'DEADLINE - 6\n90077104: syntax error in expression': the fixture's fake date returned a torn two-line clock value. Root cause: the fake forge wrapper in wrap_forge advances the shared controllable clock via a non-atomic read-modify-write ('$(cat $FORGE/clock) + 6' with truncate-in-place '> $FORGE/clock') while concurrent background gh calls run and the fake date reads the same file; an interleaved truncate+write publishes a half-written value (CI's torn '6\n90077104', tail of 1790077104) or an emptied-read value ('6'), which either breaks the poll's arithmetic (nonzero exit -> 'reservation poll failed') or defeats the 15-second reservation defer. This is a pre-existing test-fixture race, not caused by the PR's diff (base..target touches no contributions code; the same commit passed this shard in run 35711207830 earlier the same day). Fixed the flaky fixture at its root: clock_bump() now writes each new value to a per-process mktemp file in the same directory and publishes it with mv (atomic rename), so concurrent forge callers and the fake date always read one complete old-or-new clock; fault patterns and deltas are unchanged. Verified: minimal 3-way concurrency repro shows the old wrapper corrupting (12/32/38 outcomes incl. empty-read) while the rename-based wrapper never corrupts (20/20 clean); the full tests/fm-contributions.test.sh passes twice (all 38 assertions ok, incl. the reservation, budget-exhaustion, genuine-failure, shared-once, and latency tests); 10 isolated reservation runs pass; shellcheck rc=0; worktree contains only this one-file change * no-mistakes(document): drop stale file-set copy in daemon catch-all comment
…6307) * fix(bin): name the accepted verdict actors in fm-contributions help and refusal * fix(ci): Updated tests/fm-contributions.test.sh to assert exactly captain, fleet, maintainer, and nobody in command-emitted help and refusal output. Three focused regressions passed; all three extra-actor mutations were rejected. ShellCheck, syntax, and diff checks passed. Production code remains unchanged
…chenguid#6306) * fix(bin): recognise a clone root git names with different path spelling fm-fleet-sync compared git's --show-toplevel with pwd -P as strings, so a clone root that git recorded with different casing (case-insensitive volume) was skipped as not a clone root and never refreshed. Compare filesystem identity instead, which also covers symlink spelling. * fix(document): Remove stale clone-root comparison comment
) * test(calm): pin Pi's regular TUI mode where pane assertions read scrollback Pi 1.0.0 defaults its TUI to a fullscreen alternate-screen mode whose scrollable transcript is application-owned, so rows that leave the viewport never enter terminal scrollback and tmux capture-pane -S can no longer see them. The Pi Calm e2e launches now pass --tui-mode regular wherever the flag exists so the transcript assertions keep reading real scrollback on both the Pi 1.0.0 line and earlier Pi lines, which have no such flag and render regular-only anyway. * no-mistakes(document): Correct Pi TUI documentation and scrollback rationale
…ries (kunchenguid#6331) * fix(bin): encode captain-hold reasons and reject self-inventory in complete hold now stores a reason with parentheses, line breaks, or percent signs through a reversible percent encoding that every reader decodes, instead of refusing it. hold --origin records the origin on the held task, and complete refuses the origin as its own inventory entry and an entry held for a different origin; holds with no recorded origin are accepted and flagged. * fix(review): Decode marked hold reasons consistently across readers * fix(review): Remove unnecessary lifecycle test dispatch * fix(review): Correct hold origin identity and inventory recovery * fix(review): Record origins before placing backend holds * fix(document): Clarify captain-hold validation and reason reader documentation * fix(ci): Fixed both findings: failed backend holds restore the previous origin, and invalid base64/UTF-8 reasons remain verbatim. Added regressions and documented valid-literal ambiguity. Both failures were reproduced before fixes. Verification: 54 lifecycle tests and 9 wrapper tests passed; 7 Beads-specific cases skipped because tasks-axi is markdown-only. Focused lint and diff checks passed. No pipeline or publication actions performed
* fix(bin): take over the watcher cycle a main-only pass-through leaves An attended main-only pass-through leaves a successor watcher cycle running through main's handling turn. The session's next park attached to that cycle instead of owning it, so the successor's arm, orphaned by its host's exit, kept owning the watcher while the new park's arm polled it twice a second until the next close or the park boundary, hours later in a quiet second mate. A second-mate restart hit this every time, since its persist request is a main-only close. The host now records the successor it leaves for main, and the next host's first cycle runs bin/fm-watch-arm.sh --take-over on it: when that arm still owns the healthy watcher, the new arm stops it, reports a reason the cycle delivered first, and otherwise owns a fresh cycle. The stop's own downtime publication is undone over an acknowledged episode when no wake was appended in between, so the handover wakes nobody. * no-mistakes(review): Keep left-arm record until the orphaned arm is gone * no-mistakes(review): Relinquish successor arm only after durably recording it * no-mistakes(review): Relinquish successor only after its record reads back * no-mistakes(document): Clarify successor takeover guarantees and authoritative documentation * no-mistakes(ci): Fixed ci-1: acknowledgement restore now requires the exact taken-over arm/watcher ledger row with signal=TERM, awaited within a short bound. Otherwise takeover proceeds without erasing downtime. Added a self-exit regression confirmed failing before the fix and passing afterward; ordinary takeover tests and the full watcher-arm suite pass. Updated Generation reuse documentation. Syntax, diff checks, and ShellCheck pass with existing SC1091/SC2034 warnings excluded. ci-2 remains unchanged * no-mistakes(document): Clarify watcher take-over recovery and restart limits
…ile-upstream-18 # Conflicts: # .pi/extensions/fm-branch-supervision.ts # bin/fm-backlog-transition-lib.sh # bin/fm-captain-hold.sh # bin/fm-task-inbox-lib.sh # docs/captain-hold-lifecycle.md # tests/fm-calm-pi-extension.test.sh # tests/fm-captain-hold-lifecycle.test.sh # tests/fm-send-inbox-doorbell-live-e2e.test.sh
…restore e2e caveat
…fication parenthetical
… note, add archived-origin test
…nt archived origin rule
…ed-origin verification record
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Yes
That "Yes" answers firstmate's question: start the reconcile for the 18 new upstream changes. Context: running /updatefirstmate reported the fork rub-a-dub-dub/firstmate diverged from upstream kunchenguid/firstmate. Since the last reconcile point (upstream commit a774c44, reconciled in #36), upstream main has 18 new commits, including: reduce remote-job and supervision polling churn (kunchenguid#6255); load backend sibling libraries when sourced under zsh (kunchenguid#6221); exclude a remote mate's own parent channel from self-home status scans (kunchenguid#5263); document accepted contribution verdict actors (kunchenguid#6307); recognize clone roots across path spelling differences (kunchenguid#6306); preserve Pi calm transcript captures with Pi 1.0 (kunchenguid#6338); preserve hold reasons and reject invalid completion inventories (kunchenguid#6331); reclaim orphaned watcher arms on the next park (kunchenguid#6335). PR #36 was squash-merged, so the fork's history does not record upstream's 220 commits as merged even though their content is in the fork; GitHub shows the fork 238 commits behind. The plan accepted with this Yes: reconcile the 18 and land it as a real merge so upstream history is recorded and later updates only see what is new.
The captain's standing update policy (2026-09-22): "the update process should be to incorporate upstream as well as my changes. Where my changes do/don't make sense, adjust/discard them. That's the standing way this should work." Issue 5269 concern: if any incoming commit introduces an unvetted decision or gating layer into supervision code, name it in the PR description and escalate.
What Changed
a774c448, then upstream head8690c411) so upstream history is recorded rather than squashed, bringing the 18 new changes: reduced remote-job/supervision polling churn (FM_REMOTE_JOB_ACTIVE_POLL_SECONDS, 0.25s result/active sampling), zsh-safe sibling sourcing inbin/fm-backend.sh(positional params instead of an unquoted word-split string), exclusion of a remote mate's own parent channel from self-home status scans via the newfm_parent_channel_outbound_statusresolver, clone-root recognition across path spellings, an explicit verdict-actor allowlist inbin/fm-contributions.sh's refusal message, Pi 1.0 calm transcript capture, and the opt-inconfig/wait-no-turnsflag with its single retry ring (fm_task_inbox_mark_retry/_clear_retry).bin/fm-hold-reason-lib.shstores reasons asfm-hold-v1:<base64>so parentheses and newlines survivetasks-axi, andbin/fm-captain-hold.shwrites aCaptain hold origin:body line thatcomplete/verifyenforce — an entry that is its own origin, or held for a different origin, is refused, while a hold with no recorded origin is accepted on durability alone and named in the output (archived Done rows accept any matching recorded origin).bin/fm-watch-arm.shadds--take-over <arm-pid>— an incoming decision path (upstream fix: reclaim orphaned watcher arms on the next park kunchenguid/firstmate#6335) that stops the named arm's own child watcher by locked identity and conditionally restores the acknowledged recovery token through new handover helpers inbin/fm-wake-lib.sh— andbin/fm-supervision-host.shrecords a pass-through's left-behind successor instate/.supervision-host-leftand reads park elapsed time fromSECONDSinstead of forkingdate. Fork-side adjustments shorten the steering doorbell to the exported$FM_TASK_INBOXplus the inbox name (acknowledge-on-read, drain-all, under the length bound), rebalance the portable test shards with aserial_budget_mspacking check, and addtests/fm-parent-channel-scan-exclusion.test.shandtests/fm-live-lab-up-mate.test.shalongside expanded captain-hold, watch-arm, supervision-host, and task-inbox coverage.Risk Assessment
Testing
Validated the reconcile at the product level rather than only through unit runs: git ancestry evidence shows the two real merges and the fork's unrecorded-upstream count going 238 to 0, targeted suites cover the eight upstream changes the captain's "Yes" named, and hand-driven CLI transcripts demonstrate the steering doorbell (198 characters, depth-independent, drain-all clause restored), the captain-hold archived-origin accept/refuse boundary through the real tasks-axi, and the supervision-host gate change that the standing issue-5269 policy requires the PR to name. tests/fm-captain-hold-lifecycle.test.sh normally gate-skips here for a missing tasks-axi, so I installed it into a throwaway npm prefix (removed afterwards, nothing global touched) and ran its four archived-answer cases, all of which passed. The one gap is local-only: four watcher/supervision/inbox suites fail nondeterministically on this machine at a different assertion nearly every run and fail the same way at the base commit, traced to a measurable macOS ps race under load average 8-18 that makes fm_pid_identity intermittently return nothing; every individual case I isolated reached green on retry, and remote CI owns the regression verdict there. A capture of the doorbell actually landing inside a live worker pane needs a real harness process in that pane, which is the opt-in live guard already recorded as blocked from this worktree by a first-run directory-trust prompt, so the doorbell evidence is the exact emitted line plus a real fm-send run against a real tmux pane instead. All transient artifacts were removed; the worktree is clean.
Evidence: Merge ancestry: 238 unrecorded upstream commits before, 0 after
Source: Merge ancestry: 238 unrecorded upstream commits before, 0 after
$ git rev-list --count 8690c411 ^7a9cbcc8 # upstream commits the fork did NOT record, BEFORE 238 $ git rev-list --count 8690c411 ^c5467750 # upstream commits the fork does NOT record, AFTER 0 $ git log -1 --format="%H %P" 3817f1d6 # real two-parent merge of upstream/main 3817f1d6a42dcfbbea52e926f4e74386cdf7232b ea5ce37f1d0b1593c4578e8dfddb84e20d911d05 8690c4117a298ee870f0f72f571c72b846b4c5fbEvidence: Steering doorbell at the product surface, before vs after
Source: Steering doorbell at the product surface, before vs after
AFTER this reconcile (198 chars): : Firstmate instruction waiting: list "$FM_TASK_INBOX"/*.msg in 'fm-firstmate-reconcile-upstream-18-scout.inbox' steering inbox: read each in order, mv to handled/ on read, leaving none behind, act. Upstream #6240 "keep the steering doorbell short under deep homes": AFTER : shallow=198 deep=198 -> identical=yes, length grows with home depth=no BEFORE: shallow=483 deep=621 -> identical=no, length grows with home depth=yesEvidence: Captain-hold archived-origin binding, real CLI end to end
Source: Captain-hold archived-origin binding, real CLI end to end
$ captain complete sample-archived-other-review sample-archived-origin-call fm-captain-hold: captain-held task sample-archived-origin-call was held for origin sample-archived-origin-review, not sample-archived-other-review; hold a task for sample-archived-other-review or list the right one [exit 1] $ teardown sample-archived-other-review REFUSED: scout task sample-archived-other-review has not passed the captain-call completion gate. [exit 1] $ captain complete sample-archived-origin-review --none complete: sample-archived-origin-review captain-call inventory reviewed (sample-archived-origin-call) [exit 0] $ teardown sample-archived-origin-review teardown sample-archived-origin-review complete (...) [exit 0]Evidence: Supervision-host gate change (kunchenguid#6154) - issue-5269 escalation item for the PR
Source: Supervision-host gate change (#6154) - issue-5269 escalation item for the PR
home config/ primary BEFORE merge AFTER merge (no supervision files) claude host ON host ON (no supervision files) codex host off host off supervision-host (empty) claude host ON host ON supervision-host (empty) codex host ON host ON supervision-host containing off claude host off host ON supervision-host containing off codex host off host ON supervision-host-off (new file) claude host ON host off supervision-host-off (new file) codex host off host off => the old spellingconfig/supervision-hostholding "off" now reads as opt-in; the new opt-out is the separate, fleet-inheritableconfig/supervision-host-off. No home under ~/git carries either file, so nothing flips on this merge.Evidence: Behaviour coverage for the eight named upstream changes
Source: Behaviour coverage for the eight named upstream changes
Evidence: Host flakiness: base-commit comparison and measured root cause
Source: Host flakiness: base-commit comparison and measured root cause
Evidence: Captain-hold archived-answer cases run against real tasks-axi
Source: Captain-hold archived-answer cases run against real tasks-axi
Evidence: Evidence index
Source: Evidence index
Pipeline
Updates from git push no-mistakes
... (13 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
🔧 Fix: enforce origin binding on archived-answer path; fix doc order
3 warnings still open:
bin/fm-task-inbox-lib.sh:300- The fork's acknowledge-on-read doorbell wording blows through the 200-character bound this merge inherited from upstream fix(bin): keep the steering doorbell short under deep homes kunchenguid/firstmate#6240, and the guard meant to enforce it can no longer detect that. Measured: with the test's inbox name 't1.inbox' the emitted line is 195 characters, so the fixed prefix is 187 and the line is 193 + len(task-id). It therefore stays within the bound only for a task id of 7 characters or fewer. Upstream's wording at this same line (166 chars for 't1.inbox', fixed prefix 158) left room for a 36-character id. The repo's own task ids run far past 7: data/ and docs/ carry e.g. fm-claude-mods-calm-sailboat-s1 (31), which yields a 224-character doorbell. tests/fm-send-inbox.test.sh:164 asserts[ "${#typed}" -le 200 ], but its fixture task ist1(2 chars), so the assertion passes while every realistic home exceeds the bound - the guard is vacuous for the ids it is supposed to protect. The bound is not arbitrary: the function's own rationale at lines 283-287 states that a long line "wraps past what a harness composer read can prove, and a Herdr submit then reports it did not reach the pane on every re-ring", and both proof readers do whole-line comparisons (fm_task_inbox_composer_holds at line 360, fm_backend_herdr_composer_payload_shown in bin/backends/herdr.sh) whose failure makes fm_task_inbox_ring return 1 (permanent skip in thependingbranch) or 2 (send-failed). I did not find a length at which 224 characters provably breaks those readers - the Herdr capture is the full viewport and the comparison strips whitespace - so I am reporting the violated contract, not a proven delivery failure. The captain's standing policy is "where my changes do/don't make sense, adjust/discard them": this fork wording is the change that spends upstream's headroom. Remedy (author's call, which is why this is ask-user): shorten the line to fit 200 characters at a realistic id length while keeping the receipt-not-completion meaning, or raise the bound deliberately and say why - and in either case re-point the guard at a realistic-length task id so the assertion constrains the shipped line.bin/fm-control.sh:954- The acknowledge-on-read consistency pass changed four surfaces (bin/fm-task-inbox-lib.sh:300, bin/fm-brief.sh:362, tests/fm-send-inbox-doorbell-live-e2e.test.sh:5-10, docs/verification/runtime-backends.md:804) and missed a fifth worker-facing one. record_note appends to the relaunched worker's brief: "First, check your instruction inbox: list $STATE/$ID.inbox/*.msg, act on / each message in numeric order, then mv each handled file into / $STATE/$ID.inbox/handled/. A steer sent before the relaunch survives there." That is act-then-ack. A grep for ordering phrasings across bin/, tests/, docs/, .agents/ and skills/ returns this as the only remaining surface with the opposite order; bin/fm-brief.sh:100 and bin/fm-dod-lib.sh:116 are order-neutral. Concrete sequence: the note is appended to $DATA/$ID/brief.md, the same file that already carries the INBOX_SECTION the fix round made ack-first, so after a relaunch one file hands the worker two contradictory orders for the same mv, and the relaunch path is precisely the one where the note itself says a pre-relaunch steer is already sitting unhandled. A worker that follows the note withholds the mv until the steered work finishes; bin/fm-task-inbox-lib.sh:22-27 states the consequence directly - it "would keep drawing re-rings and, once the ladder is spent against an idle pane, escalate as stuck while its wait was legitimate all along." bin/fm-control.sh is not in this change's diff, so the contradiction predates the merge; it is in scope only because this run's authorized goal was to apply the corrected decision consistently. Remedy: restate the note in the shipped order (list, read each in numeric order, mv into handled/ as receipt rather than completion, then act). Flagged ask-user because it is worker-facing instruction text in a file this change has not otherwise touched.bin/fm-captain-hold.sh:977- The newest fix round added a refusal and a new output label on the archived-answer path with no test that can observe either. The change makes verify_entry_durable refuse when an archived Done row'sCaptain hold origin:differs from the completing origin, and reportrecordedinstead ofunrecordedwhen one is present. Neither existing archived test distinguishes the new behavior from the old: in test_pruned_answered_call_satisfies_completion_gate (tests/fm-captain-hold-lifecycle.test.sh:1238) the archived origin equals the origin being completed, so the comparison passes whether or not the origin is recovered at all - if archived_row_hold_origin returned nothing, the branch would silently fall back tounrecorded, enforce nothing, and the test would still pass; test_reused_archived_call_id_refuses_a_bare_second_close never reaches the loop because body_has_resolution_record already fails on the bare row. So the whole claimed fix could be a no-op and CI would be green. I traced the parsing by hand on a synthetic archive row and it does recover the origin, but that is my trace, not a guard. Missing regression, following the fork's own fixture shape: hold C --origin A, complete A C, answer C, prune --keep 0 --state done, then assertcomplete B Cfor an unrelated scout B refuses with "captain-held task C was held for origin A, not B" and that teardown of B stays refused - and separately assert that the accepted same-origin case no longer prints "[no recorded origin on: C; not checked against B]" from command_complete's output at line 1885. That test fails against the pre-fix code and passes against this one.🔧 Fix: shorten doorbell under 200 chars, align relaunch note, add archived-origin test
2 issues (1 warning, 1 info) still open:
bin/fm-captain-hold.sh:982- The fix round's archived-origin loop requires EVERY archived origin for a call id to equal the completing origin, which newly refuses a completion whose own answered row is sitting in the archive, and permanently wedges that scout's teardown. The live path at lines 947-953 reads one row and so has one stored origin; the archived path at 977-986 iterates thesort -uset of origins across every- [x] <call> -row and fails on the first mismatch. Reused call ids with different origins are not hypothetical - tests/fm-captain-hold-lifecycle.test.sh:1300 (test_reused_archived_call_id_refuses_a_bare_second_close) proveshold <call> --origin <second>succeeds once the first incarnation has been pruned, with its own assertion "the pruned call id could not be re-held for a second incarnation". Concrete sequence, differing from that fixture only in that the second incarnation is answered rather than bare-closed: (1)hold C --origin A;complete A C;answer C;prune --keep 0 --state done-> archive row 1 carriesCaptain hold origin: Aand a resolution record. (2)hold C --origin B;complete B C;answer C;prune-> archive row 2 carries origin B and a record. (3) Nowcomplete A --nonere-verifies A's persisteddecision_keys=C: resolve_entry finds no live row, archive_has_resolution_record returns 0 with origins {A,B}, the loop reaches B and fails with "captain-held task C was held for origin B, not A". Against commit 07c591d the same step printedC archived-answer unrecordedand succeeded. Because command_verify (line 1927) runs the identical check and is called by scout teardown, A's teardown is refused from then on with onlybin/fm-teardown.sh --forceas an escape - even though row 1 is exactly the evidence the gate asks for. The all-rows rule is correct forbody_has_resolution_record(an earlier answer must not vouch for a later bare close) but does not transfer to origins: the invariant the gate states is "was held for this origin", which row 1 satisfies. Remedy, inside the same block rather than any new machinery: accept when at least one archived origin matches, and refuse only when at least one origin is recorded and none matches (keeping the existing requirement that every archived row carry a resolution record). Flagged ask-user because it changes an accept/refuse boundary on the captain-hold completion gate that the fix round chose deliberately, and because neither the new test nor test_reused_archived_call_id_refuses_a_bare_second_close covers two answered incarnations, so whichever rule is chosen needs paired coverage for that shape.bin/fm-supervision-host.sh:243- Informational, for the PR-description step: the user intent's issue-5269 clause asks that any incoming commit introducing a decision or gating layer into supervision code be named in the PR description. This merge brings exactly one such layer, upstream fix: reclaim orphaned watcher arms on the next park kunchenguid/firstmate#6335: a new durable.supervision-host-leftrecord (line 243) written by detach_successor (line 673), read back and liveness/identity-checked in activate, and consumed asstart_arm ... --take-over "$LEFT_ARM"(line 1051), with the matching new--take-overmode and take_over_cycle decision in bin/fm-watch-arm.sh:487,556. It arrives vetted - tests/fm-watch-arm.test.sh:1117 and :1147 cover both the not-owner attach and the owning take-over, and tests/fm-supervision-host.test.sh gained 251 lines - so this is not a defect, only the item the intent asks to be named. Upstream fix: reduce remote-job and supervision polling churn kunchenguid/firstmate#6255's polling change in the same file is a cadence change (the 0.5s arm-exit probe inside the unchanged POLL cadence), not a new gate.🔧 Fix: accept any matching archived origin, add reused-id coverage
2 warnings still open:
docs/captain-hold-lifecycle.md:113- The contract doc's origin rule no longer matches the archived path the fix rounds built, and the archived path's origin rule is documented nowhere. Line 113 states flatly "An entry whose recorded origin differs from the one being completed is refused." For a live row (bin/fm-captain-hold.sh:953-959) that is exact: one stored origin, refuse on any difference. For an entry whose row has been pruned, bin/fm-captain-hold.sh:978-994 now implements a different rule: archive_has_resolution_record returns the sort -u set ofCaptain hold origin:values across every- [x] <call> -row, the loop accepts on the FIRST match and breaks, and refuses only when at least one origin is recorded and none matches. So an entry whose archive carries origins {A, B} is accepted for A even though an origin recorded for that id does differ from A - the opposite of what line 113 tells a reader. The archive paragraph at line 132 documents the sibling rule for resolution records in detail ("an archived identity counts only when every archived row for that id carries a record") and says nothing about origins, so there is no second place that corrects line 113 either. This matters concretely: the same check runs in command_verify (bin/fm-captain-hold.sh:1935), which scout teardown calls, so an operator debugging a refused teardown reads line 113, expects per-row refusal, and cannot reconcile it withcompletesucceeding for one of two archived origins - and the only documented escape isbin/fm-teardown.sh --force. Remedy is one sentence beside line 115 (and matching wording near line 132): when the live row is gone, the gate accepts the entry if any archived Done row for that id records the completing origin, and refuses only when one or more origins are recorded and none matches, while every archived row must still carry its own resolution record. Classified auto-fix because it only writes down a boundary the user already decided in the previous round; it changes no behavior.bin/fm-captain-hold.sh:976- Simplification: the fix round added a second copy of the self-inventory rule on the archived branch that cannot fire, and the component is not needed by any intent requirement. Reaching line 976 requires resolve_entry to have returned 1 (line 968), which meanstask_show "$entry"failed - $entry has no live row in this home's backlog.origin_idat line 975 is one of exactly two values: iftask_show "$origin"succeeds, task_identity returns that row's ownidfield, and a row found by that id is by definition showable, sotask_show "$entry"would have succeeded and the live branch would have been taken; iftask_show "$origin"reports NOT_FOUND, task_identity returns $origin verbatim, and[ "$entry" = "$origin" ]was already refused at line 939 before resolve_entry ran. Either way[ "$entry" != "$origin_id" ]is always true and refuse_self_inventory is never called from here. The live branch's equivalent check at line 951-952 is the reachable one, because there $id is the resolved identity of a row that does exist. Note that theorigin_id=$(task_identity "$origin")assignment on line 975 IS required - the comparison loop at line 983 reads it - so the narrower form that satisfies the intent is to keep the assignment and drop line 976. Flagged ask-user rather than auto-fix because removing a refusal line on the captain-hold completion gate touches an accept/refuse boundary a prior fix round chose deliberately, even though it is provably inert, and because the two new archived tests cannot distinguish its presence from its absence.🔧 Fix: drop inert archived self-inventory check, document archived origin rule
2 issues (1 warning, 1 info) still open:
bin/fm-task-inbox-lib.sh:300- The pipeline fix rounds rewrote the one agent-facing prompt whose effectiveness rests on a dated live-model verification, and then recorded in the project's own verification file that the verification no longer matches the emitted line. Upstream fix(bin): keep the steering doorbell short under deep homes kunchenguid/firstmate#6240 arrived in merge 3817f1d self-consistent: its doorbell measured 189 characters with the test's then-currentt1fixture, under the 200-character bound tests/fm-send-inbox.test.sh:172 asserts. Fix round a1f7745 ('adopt upstream doorbell line with fork receipt clause') appended a fork clause, pushing it to 239 and breaking that bound; rounds 07c591d and 9a685c7 then shortened it to the current text, which also drops upstream's explicit drain-all clause 'leaving none behind' (tests/fm-task-inbox.test.sh no longer asserts it; the surviving 'read each in order' is the only drain-all signal). In the same rounds the fixer added docs/verification/runtime-backends.md:806 - 'The wording the guard steers against has since changed; a refresh should re-date this entry against the linebin/fm-task-inbox-lib.shcurrently emits.' That makes the merge ship an unverified behavioral assumption where upstream shipped one verified on 2026-08-23 against every installed harness. Nothing closes the gap in CI: tests/fm-send-inbox-doorbell-live-e2e.test.sh gates onfm_live_gate opt-in FM_SEND_INBOX_LIVE_E2E tmux, so no automated run observes whether a real worker still comprehends the line; the shipped tests only assert substrings of the emitted string. Concrete consequence: a worker that no longer acknowledges every record leaves records unmoved, firstmate re-rings and eventually treats the worker as stuck - exactly the failure the 2026-08-23 run was spent to rule out, now backed only by a stale verdict the fixer itself flagged. Reported ask-user because the remedy is a decision, not a mechanical correction: either spend a live run (FM_SEND_INBOX_LIVE_E2E=1 tests/fm-send-inbox-doorbell-live-e2e.test.sh) and re-date the entry, or restore upstream's wording verbatim and revert the fixture tightening at tests/fm-send-inbox.test.sh:151 so the merged pair is self-consistent again. Note that restoring upstream verbatim is worse for real homes - at a realistic 41-character task id upstream's line measures 265 characters - so the wording the fix rounds landed is likely the one to keep and re-verify, but that is the author's call, not the reviewer's.tests/fm-send-inbox.test.sh:172- Informational, noting a tradeoff the change deliberately accepts. The 200-character doorbell bound arrives with this merge (it does not exist at base 7a9cbcc) and fix round 9a685c7 made it bite by replacing thet1fixture with a realistic 41-character task id. Measured against the current code, the emitted line is 198 characters for that fixture - two characters of headroom - and nothing bounds task-id length: validate_slug (bin/fm-captain-hold.sh:283) constrains the character set only, and fm_task_inbox_doorbell_line embeds<task-id>.inboxverbatim, so a 44-character task id emits 201 characters and crosses the threshold the code comment says makes 'a long line wrap past what a harness composer read can prove', with a Herdr submit then reporting it never reached the pane on every re-ring. No action is warranted here: the change is a large net improvement - the fork emitted 459 characters at base for the same id, and upstream's merged line would emit 265 - and the only remedies (capping task-id length, or shortening the prompt further) would extend this change beyond reconciling the 18 upstream commits. Flagging it so the thin margin is a known property of the invariant the fix round just made constraining, rather than a surprise the next doorbell edit discovers.🔧 Fix: restore drain-all clause in doorbell; live refresh blocked
2 warnings still open:
docs/verification/runtime-backends.md:807- The live run the captain asked for in the previous round could not be performed, and the reason is structural rather than transient, so it is returned as the requested escalation instead of being left in the doc. The wording leg of that instruction is satisfied: the emitted line is now 198 characters for the test's 40-character task id (fm-firstmate-reconcile-upstream-18-scout), under the 200-character bound at tests/fm-send-inbox.test.sh:172, and it keeps both "steering inbox" and the restored drain-all clause "leaving none behind", which tests/fm-task-inbox.test.sh:179 now asserts. The live leg is not: tests/fm-send-inbox-doorbell-live-e2e.test.sh:81 opens its isolated tmux session with-c "$ROOT", and every harness window is launched with that same cwd (:139), so the harness is asked to trust whatever checkout the guard runs from. Run from this pipeline, $ROOT is the ephemeral worktree/.no-mistakes/worktrees/00644a303d6a/01M3XSA2Q1W3GBVVDSQH6C68RM, which neither harness has ever trusted, and wait_ready (:118) deliberately refuses to dismiss a screen matching 'trust' because the header (:32-34) treats a trust dialog as a real unready state. That is exactly the failure the fixer recorded: claude 2.1.287 returned ready_rc=1 (/git/firstmate) or trust this path once in each harness, then re-date the entry for Claude and Codex and name the harnesses not re-run. Accepting the merge with the entry as it stands - an explicit, dated, non-evidence record of a blocked refresh - is also a defensible containment call, but it is a call to make rather than to assume.composer stayed visibly pending), codex-cli 0.154.0 proceeded on ready_rc=2, received the steer, and sat on its own trust prompt withacted=no acked=no. Consequence: the merge ships the one agent-facing prompt whose effectiveness rests on live-model comprehension backed only by the 2026-08-23 verdict for wording the code no longer emits, and the note at line 806 says so. Nothing in CI closes the gap - the guard is gated behindfm_live_gate opt-in FM_SEND_INBOX_LIVE_E2E tmux, and the deterministic tests only assert substrings of the generated line. The concrete risk if a model now reads it differently: a worker that moves only the first record leaves the rest unmoved, the watcher re-rings and eventually reports the mate stuck. Reported ask-user because the remedy is the captain's, not this phase's: runFM_SEND_INBOX_LIVE_HARNESSES="claude codex" FM_SEND_INBOX_LIVE_E2E=1 bash tests/fm-send-inbox-doorbell-live-e2e.test.shfrom a checkout both harnesses already trust (bin/fm-supervision-engine-lib.sh:74- The user intent carries a standing instruction: "if any incoming commit introduces an unvetted decision or gating layer into supervision code, name it in the PR description and escalate." Of the 18 incoming commits, one meets that description and is surfaced here so the escalation happens at review time. Upstream fix: inherit supervision host opt-out across secondmates kunchenguid/firstmate#6154 (12e90e1) rewrites the supervision-host home gate: fm_supervision_host_enabled previously readconfig/supervision-hostand opted the home out when its first word wasoff; it now refuses on the presence of a separateconfig/supervision-host-offand treats ANYconfig/supervision-hostfile as opt-in (bin/fm-supervision-engine-lib.sh:74-76). Two consequences worth the captain's explicit vetting. First, the opt-out became fleet-wide and primary-authoritative:supervision-host-offwas added to FM_INHERITABLE_CONFIG (bin/fm-config-inherit-lib.sh:85), and per .agents/skills/secondmate-provisioning/SKILL.md:118 the primary's absence of the flag now REMOVES a secondmate's own copy at convergence - a mate can no longer hold a supervision posture its primary does not share. Second, the old spelling is not migrated: a home whoseconfig/supervision-hostholdsoffis now read as opt-in and would start running the host. I checked reachability for this fleet and the second consequence is not currently live - there is noconfig/supervision-hostin ~/git/firstmate/config and none anywhere under ~/git - so no home flips on this merge; the item needing a decision is the new propagation rule, plus whether a one-line migration note belongs in docs/configuration.md "Supervision host" before a mate is ever opted out. Classified ask-user because the remedy is a product decision about an upstream gating contract (accept it under the standing update policy, or diverge), not a correction to what the merge already does; the two neighbouring supervision changes the PR phase should also name are fix: reduce remote-job and supervision polling churn kunchenguid/firstmate#6255 (polling cadence) and fix: reclaim orphaned watcher arms on the next park kunchenguid/firstmate#6335 (the --take-over arm-reclaim gate in bin/fm-watch-arm.sh:551-604).tests/fm-watch-arm.test.sh- Four watcher/supervision/inbox suites are nondeterministic on this machine and could not be driven to a stable all-green locally: tests/fm-watch-arm.test.sh, tests/fm-wake-queue.test.sh, tests/fm-supervision-host.test.sh and tests/fm-task-inbox.test.sh. Each run stopped at a DIFFERENT assertion (these suites' fail() exits, so a later stop proves the earlier assertions passed that run), and every one of them also fails at the base commit 7a9cbcc in the same environment, so this is not introduced by the reconcile. Isolating one case (fm-supervision-host's attended captain outcome) and running it four times gave fail/pass/fail/pass; the host's own log named the reason, 'pass-through attended: the main session could not be identified', which is bin/fm-supervision-host.sh:1023 -> fm_supervision_host_main_key -> fm_pid_identity, a function byte-identical at base and HEAD. Measured root cause: fm_pid_identity resolves a pid withps -p <pid> -o lstart= -o command=, and on this loaded host (load average 8-18 on 10 cores from other work on the machine) ps intermittently returns nothing for a live, freshly forked process - 4 failures in 60 direct probes. That same probe backs the watcher lock holder, the main-session key, and tests/lib.sh's fixture temp-root owner, which covers every failure observed. The captain-hold archived cases hit the same instability on repeat runs ('task record authorized directory cannot be resolved', 'whether its backlog item is still held for the captain could not be read'), though all four of those cases did reach green. Nothing in the change needs fixing for this; remote CI owns the regression verdict for these suites. Recorded so the local red is not mistaken for a product regression.git rev-list --count 8690c411 ^7a9cbcc8and^c5467750- upstream commits unrecorded by the fork, 238 before the reconcile and 0 aftergit merge-base --is-ancestor <c> c5467750for each of the 18 commits ingit rev-list 8690c411 ^a774c448git log -1 --format='%H %P' 3817f1d6andea5ce37f- both reconcile commits are real two-parent mergesbin/fm-test-run.sh tests/fm-parent-channel-scan-exclusion.test.sh tests/fm-fleet-sync.test.sh tests/fm-backend.test.sh tests/fm-contributions.test.sh tests/fm-remote-job.test.sh tests/fm-calm-pi-extension.test.sh- all pass (calm-pi gate-skips, Pi package not installed)bin/fm-test-run.sh tests/fm-send-inbox.test.sh- passes, covers the doorbell and the 200-character boundbin/fm-test-run.sh tests/fm-task-inbox.test.sh- passes (1 of 5 attempts; see flakiness finding)bin/fm-test-run.sh tests/fm-brief.test.sh- passes, covers the reworded brief inbox sectionbin/fm-test-run.sh --per-script-timeout-secs 900 tests/fm-watch-arm.test.sh tests/fm-wake-queue.test.sh tests/fm-supervision-host.test.sh- run serially and concurrently, plus the same three at base 7a9cbcc8 for comparisonTrimmed copy oftests/fm-captain-hold-lifecycle.test.sh(same helpers, fixtures and assertions) running its four archived-answer cases against real tasks-axi 0.2.6 - each passed: pruned recorded answer satisfies completion, pruned answered call binds to its recorded origin, archived earlier incarnation does not answer a bare close, each reused-id incarnation completes under its own originManual CLI transcript: realbin/fm-captain-hold.sh hold/complete/answer,tasks-axi prune --keep 0 --state done, thencomplete <wrong-origin> <call>(refused by name),bin/fm-teardown.sh <wrong-origin>(refused),complete <right-origin> --noneandbin/fm-teardown.sh <right-origin>(both accepted)Manual doorbell comparison:fm_task_inbox_doorbell_linefrom HEAD and from base 7a9cbcc8 over the same shallow and deep fixture homes with a 40-character task id, measuring emitted length and clause presenceManual pane run: realbin/fm-send.shagainst a real tmux 3.6b pane on an isolated socket, confirming the payload text never reaches the terminal and the dead-pane path reports the durable recordManual supervision-gate truth table:fm_supervision_host_enabled <config-dir> <primary>from HEAD and base over four config shapes x two primaries, plus a filesystem scan forconfig/supervision-hostunder ~/gitfor i in 1..60: sleep 30 & ; fm_pid_identity $!- direct measurement of the host ps race behind the flaky suites (4 failures in 60)docs/fm-test-portable-shards.md:15- Two portable-serial members now carry no measured duration hint and are balanced on PORTABLE_SERIAL_DEFAULT_WEIGHT_MS instead of evidence: tests/fm-parent-channel-scan-exclusion.test.sh, which this merge adds (upstream fix(bin): exclude a remote mate's own parent channel from self-home status scans kunchenguid/firstmate#5263, after the 2026-09-30 hint refresh), and tests/fm-test-lib-plan.test.sh, which the fork owns. Measured here: bin/fm-test-run.sh --check-coverage reports serial=203 against a 201-entry portable_serial_weight_hints table, serial_unhinted=2, and still passes the guard's PORTABLE_SERIAL_MAX_UNHINTED_PERCENT bound. I corrected the doc's coverage claim to the baseline it actually describes and pointed at serial_unhinted= rather than inventing a count, because the remedy the doc itself prescribes ('Refresh the hints whenever a serial member grows materially or the lane gains scripts') needs per-script duration_ms from several green Ubuntu CI runs, which this documentation phase cannot produce and must not fabricate. No action is warranted in this change: the coverage guard keeps the partition complete and disjoint whatever the hints say, so the cost is a slightly less even shard, not lost coverage. Recorded so the next hint refresh picks up both scripts.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.