Conversation
…ng state (#2) * fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413) A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload. * fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary. * docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes * fix(calm): keep Pi's export confirmation visible (kunchenguid#2461) Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared. * feat(bin): render the captain's four-section status board from existing state fm-status-board.sh prints one HTML page (Needs you to continue, In progress, Waiting, Recently completed) rendered entirely from fm-bearings-snapshot.sh's already-correct cross-home classification and fm-fleet-snapshot.sh's untruncated per-item detail, with no second store for agents to keep in sync. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
Tearing down a task deletes state/<id>.meta, the record bin/fm-pr-merge.sh resolves the PR through. A branch can be fully pushed and reachable from a remote (so the existing dirty/unpushed/landed checks find nothing to refuse) while its own PR still sits open, stranding a mergeable PR with no guarded path to land it - this happened twice in one night. The new check runs after the existing dirty/unpushed/landed checks inside validate_worktree_teardown_safety, so their refusal messages take precedence and never contradict this one. It fires only when a pr= is actually recorded (scout, local-only, and secondmate teardowns never record one). A gh lookup error refuses loudly too, naming the PR firstmate could not confirm, rather than tearing down blind. --force already skips the dirty/landed checks and now explicitly skips this one too, since it is the same "captain says discard" escape hatch - the usage comment spells out what that means here (the PR needs merging or closing by hand, outside fm-pr-merge.sh's guarded path). Co-authored-by: joliverMI <joliver@sensibletech.biz>
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413) A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload. * fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary. * docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes * fix(calm): keep Pi's export confirmation visible (kunchenguid#2461) Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared. * fix(procevent): deliver captured results at most once A Lavish-captured answer reached a secondmate five times because forwarding it ran fm-send.sh and only afterward, manually, remembered to call `handled` - so every re-announcement of the still-unacknowledged capture (a restart, a compaction, a re-read wake) repeated the forward. The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves capture is exactly-once, but nothing paired a downstream effect with its acknowledgement atomically. Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`, which checks handled status, runs the command, and marks the generation handled as one call under the source lock, so a repeated delivery attempt against an already-delivered generation runs the command zero times and a failed command stays eligible for retry instead of being dropped. Point the process-event-sources skill and the runner's operating contract at it for exactly this class of forward. Five delivery attempts against one captured generation now run the downstream command once instead of five times - the read-and-discard cost of the other four is eliminated structurally. The historical incident's own token cost was not measured at the time and is not reconstructable after the fact. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
…#4) Adds a fifth board section between "Needs you to continue" and "In progress" for finished work that is deployed and usable, not merely merged. It renders only when firstmate has written a "ready: <where to see it>" note on a landed backlog row, confirming deployment and naming the surface - never inferred from merge/landed state alone, so a merged-but-not-yet-deployed change never appears here. The item shows its pointer text, never a PR link or repo path. Approval buttons (part 2) are not implemented in this PR: investigation found Lavish's artifact bridge (window.lavish.queuePrompt) can queue an exact, single-item approval string, but transmission still needs the existing manual "Send to Agent" tap by Lavish's own design, and the delivery path could not be confirmed end-to-end from this environment. Ship the verified section now and report the button investigation as a plan. Co-authored-by: joliverMI <joliver@sensibletech.biz>
#5) The board named locations the captain could not reach from his phone: a pull request he does not review, a raw data/<id>/report.md path, a "fuller detail lives in the X home record" cross-home line with nothing to tap, and a bearings-truncated in-progress summary with no full text behind it. Pull requests and GitHub links are now never rendered anywhere on the page. A report links only when its backlog note carries an explicit report-url: marker firstmate wrote after actually publishing it (the same convention as the existing ready: marker) - investigated whether Lavish's existing route could serve reports directly and confirmed it cannot: Lavish is a publish-one-artifact-at-a-time surface, not a path-mapped static file server, so there is no passive route to piggyback on. A Ready item's "where to see it" text is flagged honestly instead of shown verbatim when it looks like a filesystem path or a pull-request URL. A cross-home item's bounded detail says plainly that the rest is not reachable from the page, instead of naming an unreachable home record. An in-progress row now prefers its fuller current-state text over the shortened one-line summary. Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488) * feat(stow): persist the open records a session is holding /stow curated memory and captured session knowledge, but never touched record state, while AGENTS.md called it an "unfinished-work sweep" and the receipt declared the session "safe to reset" - wording that implied a record-correctness guarantee stow does not make. A shipped PR with no backlog item, a queued umbrella whose phases had merged, and four decision holds left open after their answers shipped all survived repeated stows. Add a bounded pass that files record state from the same volatile input the rest of stow already uses: the open threads in context, minutes before the reset destroys them. It creates a record for an unfiled thread and corrects one the session knows is wrong, through the owning path, and states its boundary as part of the contract - it never enumerates the backlog, lists holds, or queries a forge, because it cannot be a reconciliation and must not be read as one. Correct the wording in AGENTS.md and the completion receipt so reset-safe means what it actually guarantees: nothing this session knew was lost. * no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi * no-mistakes(document): note /stow open-record persistence in README command catalog * refactor(stow): state open-record persistence as principle, not procedure The first version enumerated triggers, named commands, and prescribed an ordered procedure. That is too rigid for an agent skill: it invites literal execution of a checklist instead of judgment, and every enumerated example is a way for the guidance to go stale. Reduce it to the intent - before a reset, the important open work you are holding in context must end up durably recorded rather than dying with the session, filing what is unfiled and correcting what is stale - and let the agent judge importance, the record, and the owning write path. Keep the scope bound, since it is a decided contract and not a mechanic: this covers the open work the session is holding, never a reconciliation of durable records against repository or forge reality. The wording corrections in AGENTS.md and the completion receipt are unchanged. * fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490) * fix(decisions): close captain holds at answer time Firstmate had two "a decision is open" ledgers with asymmetric closing mechanics. The live status-log ledger closes atomically at answer time, because bin/fm-send.sh --resolve-key makes answering a decision be the act that closes it. The durable backlog hold ledger had no such coupling: answering and recording were two separate acts, and only the first was forced by the workflow. That asymmetry lost four real captain decisions. Their answers were captured durably to disk, keyed character for character by the hold decision keys, acknowledged, and even implemented and shipped, yet the holds stayed open for two days and the captain was asked to re-answer decisions already on his own disk. Give the hold ledger the same answer-time-closure property: - bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's counterpart to --resolve-key. It shares one unrouted close implementation with `decline`, so it carries every existing guard - the captain decision file, the active-hold requirement, retry identity, and the refusal to release still-routed work - and differs only in the resolution mode it records. `decline` keeps its stronger meaning that the answer routes no follow-up work at all. - bin/fm-procevent-lavish.sh wires the channel that actually carried the lost answers. `arm --decisions-origin` binds a deck to the origin whose holds it carries, `answers` reads the structured choices out of a captured poll result, `close-decisions` maps each key to its hold and closes it through the command above, and `autohandle` lets the runner apply that at capture time. Safety is preserved rather than traded away. Only rows tagged `choice` are read, so freeform captain prose cannot forge a decision key. Closure is confined to the one bound origin. The decision text is a pure function of the captured result, so a replayed capture is idempotent. A hold that is absent, already closed, or still blocking routed work is skipped and left for `resolve`, never forced. A deck armed without the binding touches no hold at all. And autohandle deliberately never reports full handling, because recording an answer is transcription while acting on it is firstmate's judgement - so the check wake still reaches the handler. fm-send --resolve-key is untouched. * no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory * refactor(decisions): make keyed-answer closure one general capability The previous pass gave holds answer-time closure but built it as bespoke Lavish wiring: the review adapter carried the source-to-origin binding, mapped keys to hold identities, wrote decision records, decided what to skip, and closed holds itself. That treated a review deck as a special decision source. It is not - it is an ephemeral discussion format that happens to carry answers. Collapse it into ONE general capability with one owner. bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its matching hold": - `answers <origin> --source <provenance>` is the channel-agnostic intake. It reads key/answer/label lines on stdin, maps each key to its hold, and closes it through the same `answer` path, so every guard applies identically whatever channel the answer came from. --source is provenance recorded in the decision, never a behavior switch; there is no per-channel branch and no knowledge of chat, decks, or transports. - `bind`/`unbind`/`binding` own the source-to-origin binding for any channel whose answers arrive detached from their origin. Every channel is now an ordinary caller that only turns what it received into keyed lines: - bin/fm-send.sh (chat) feeds the intake for a key that names an active hold. This also fixes a real gap: once `complete` transfers a decision to its hold it closes the live status copy, so --resolve-key alone could never answer a transferred decision. - bin/fm-procevent.sh feeds it generically. A bound source's captured result goes to `<adapter> answers <result-file>` and whatever that prints is piped into the intake. The runner names no adapter, parses no result, and carries no decision rule, so any future adapter with an `answers` command works with no change here. - bin/fm-procevent-lavish.sh keeps only `answers`, which reports the structured choices a review captured and stops. It maps nothing to a hold and closes nothing; it lost ~160 lines of decision logic. Feeding is independent of handling, so it never acknowledges a result and never suppresses a wake - recording an answer is transcription, acting on it stays firstmate's judgement. The regression that proves closure now drives a FIXTURE adapter that is not the review adapter, so what is proven is that any bound channel reaches the intake rather than that one channel is wired specially. A new regression drives the real fm-send over a stubbed transport for the chat side. Every prior guarantee still holds, and fm-send's status-log behavior is unchanged. * no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression * fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512) A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md. The installer now creates and migrates to a recoverable two-line pointer file. * fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515) * feat(dashboard): add the Admiral's Fleet Dashboard task board Replaces the generated Lavish status page with a purpose-built board: a zero-dependency Python backend (stdlib http.server + sqlite3), a vendored React frontend styled to match Spectra, an agent-facing CLI (bin/fm-dashboard.sh) that is the only way agents touch it, and the fleet-dashboard skill teaching the fleet auditor and other agents how to use it. The board owns its own persistent records rather than mirroring any backlog - its six-status vocabulary and four review tabs don't map onto tasks-axi state, and the fleet auditor needs a separate claim to check against live reality. See docs/dashboard.md for the full architecture, the deploy story, and the drift-risk tradeoff that choice implies. Link posting rejects GitHub/PR links and local-only hosts structurally (standing order 17), not just by agent memory. The page shows a loud, persistent banner when it cannot reach its own API, and the auditor's "never run" state is rendered distinctly from a clean sweep, so an absence of confirmation is never shown as a confirmed-good state. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
Testing and needs-attention were conflated: both put a finished-enough card in front of the Admiral, but only needs-attention actually blocks on him. A card can now be set to needs-attention with a --reason that renders directly on the card, sorts above every other status, and gets its own loud always-visible section on the page. The fleet auditor's sweep now treats a needs-attention card's age as a finding, since that status - unlike testing - is not supposed to sit unanswered. Existing databases migrate in place; the new column is nullable and existing cards, history, and statuses are untouched. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…orce Audit button (#8) The auditor's 15-minute cadence never actually ran: it registered a durable check in its own home, but nothing polled it because that home had no live supervision session. Host cron (already running independent of any chat session) now drives bin/fm-fleet-audit-tick.sh, which heartbeats every tick and sweeps via bin/fm-fleet-audit-sweep.sh once the dashboard's own interval setting says it's due, self-healing a lock abandoned by a crashed sweep. The Force Audit button uses the identical sweep script through a detached subprocess so a scheduled and a forced run can never stack. Also fixes the "unassigned" agent noise on every card (shows the captain instead) and adds a Go to card jump from each discrepancy log entry. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…ard (#9) Every dashboard card's agent and backlog_ref sat empty, so status transitions only happened when someone remembered to make them by hand - worst of all at teardown, the exact moment the worker that would have remembered is being removed. fm-spawn.sh --card <card-id> now records dashboard_card=<id> in the task's own state/<id>.meta and best-effort links the card's ref/agent to the task identity, advancing a not_started card to working. fm-teardown.sh reads that identity back and advances the card to testing once a real (non-forced) landing succeeds, leaving scout, secondmate, and already-complete cards untouched. No card is the normal case and stays a complete no-op. A bad card id or an unreachable dashboard never fails the spawn or the teardown; it warns loudly and best-effort records fm-dashboard.sh audit-log --fleet so a failed link is visible to the fleet auditor rather than silently dropped. All 21 live cards still carry no agent/ref - 20 belong to a project this task has no access to verify task identities for, and the one firstmate-owned card is already complete with no live task to link - so none were backfilled; see the PR description for the per-card list. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…nd sends (#12) * fix(bin): tmux endpoint liveness reports dead targets as alive fm_backend_target_exists's tmux branch trusted display-message's exit status alone. tmux silently falls back to the caller's own active pane (or an unrelated live session/window whose name is an unambiguous PREFIX of the missing target) while still exiting 0, so a dead tmux endpoint reads alive whenever the caller itself runs inside tmux - which firstmate always does - or shares a socket with any other live session/window. Pin the tmux probe to an exact target: session:window pairs use list-panes -t "=$session:=$window" (the same exact-match form fm_backend_tmux_kill already uses), bare %N pane ids / @n window ids / $N session ids pass straight through since list-panes already resolves them exactly, a bare (colon-less) target is checked against exact session and window-name inventory since tmux's `=` pin does not make a bare target exact the way it does the session component of a session:window target, an all-digit window component resolves exactly against the session's own window-index inventory (FM_SUPERVISOR_TARGET _DEFAULT is "firstmate:0", an index not a name), and a dotted window component is tried first as an exact window name and, if that fails, falls back to pane-qualified addressing (window.pane or window-index.pane) - both are valid, competing readings of the same string in real tmux. fm-crew-state.sh's pane_readable duplicated the same probe and now defers to the fixed primitive. A task record naming a remote host is no longer judged by the local tmux server at all in either caller: fm-session-start.sh's digest reports those as an explicit "not checked from here" verdict instead of probing a session that can never resolve locally, and fm-crew-state.sh skips the local probe entirely for a remote_host record so its status-log fallback - previously reached only by accident, through the old probe's own bug - still runs. fm-fleet-snapshot.sh, fm-control.sh, and fm-send.sh were checked and already guard remote_host before reaching this probe. herdr, zellij, orca, and cmux already address their target by id through a structured inventory query, so they do not share this defect shape. Adds a real-tmux regression suite that runs the probe from inside a live tmux pane (the one condition that reproduces the false positive, since the fallback needs a live session to fall back to) against nonexistent, prefix-colliding, dotted, pane-qualified, bare-name, and window-index targets, and a crew-state case proving a remote secondmate's status is still read when the local probe always fails. All fail against the pre-fix code and pass after. Updates fake-tmux test fixtures that modeled only display-message's semantics to model the pinned probe instead. Known gaps, deliberately not fixed here and tracked in one follow-up, fm-tmux-agent-state-session-prefix-match: the recovery-grade sibling fm_backend_tmux_agent_state (bin/backends/tmux.sh) still resolves its session component by unpinned prefix match, same defect class, unfixed; and fm_backend_tmux_kill can silently destroy a DIFFERENT live window (not merely fail to remove the target one) when given a dotted window name - verified live, also unfixed. No task id in this fleet currently contains a dot, so nothing live is exposed by either gap today. The follow-up will converge every tmux target-resolution helper in this repo onto one shared resolver rather than each parsing target strings independently. Delivery note: this branch was built by hand directly on this fork's own current main tip (git apply --3way of the reviewed diff) rather than by rebasing, because this fork's history has diverged from the upstream template repository it was created from and a rebase would pull in commits that exist only on that upstream - this repo's standing rule is that nothing is ever proposed to it. * no-mistakes(review): refuse tmux key sends to targets that resolve inexactly * no-mistakes(review): guard and exact-pin every tmux input-delivering send * no-mistakes(review): resolve tmux probe and send targets through one resolver * no-mistakes(review): refuse ambiguous bare tmux names, correct stale disclosures * no-mistakes(test): add list-panes arm to shared secondmate fake tmux * no-mistakes(document): align endpoint-verdict and supervisor-target docs with exact tmux resolution * no-mistakes: apply CI fixes --------- Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(dashboard): mechanically link handed-off backlog items to their card Cards for work routed to a secondmate had no mechanism to advance: status advances only through bin/fm-spawn.sh --card and bin/fm-teardown.sh, both keyed off local task metadata that bin/fm-backlog-handoff.sh never creates, so a card handed to a secondmate froze at not-started forever. bin/fm-backlog-handoff.sh --card <card-id> gives a single handed-off item the same best-effort link: once it lands in the secondmate's backlog, it sets the card's ref/agent and advances a not_started card to working. The durable identity lives in this home's own state/handoff-cards/<id> record (never in the backlog item, which the handoff does not own) and is retired per pair only once the board genuinely confirms that pair's link - never on a merely-attempted one, so a failed link is retried by the next arrival instead of silently orphaning the card. bin/fm-dashboard.sh's own curl calls are now bounded so an unreachable board cannot stall a held handoff lock. Includes a fix to the shared dashboard-card-link test's own server-start wait: it previously trusted bin/fm-dashboard.sh's 1-second process-liveness sleep as proof the server was ready, which flakes on a busier CI runner (confirmed pre-existing on main, not introduced here). It now polls the real /api/health reachability check before running any test. * no-mistakes(review): bound dashboard curl calls, fix card-link identity, 404, and record races * no-mistakes(review): confirm card links only on read state, keep unlinkable pairs * no-mistakes(review): fail card-link ownership check closed, never link superseded cards * no-mistakes(document): fix stale usage line for multi-item handoff * fix(tests): show captured output on a failed exit-code assertion expect_code only ever reported "expected exit 0, got 1", with no way to see why - the same missing-evidence shape as the CI failure this is meant to help diagnose. Accept the command's own captured output as an optional argument and print it on failure, matching how assert_contains/assert_not_contains already show theirs. Wired into every expect_code call in tests/fm-dashboard-card-link.test.sh. Also fold in the changed-file test-family mapping this branch's rebuild had dropped: bin/fm-backlog-handoff.sh, bin/fm-spawn.sh, and bin/fm-teardown.sh (the three entry points to the mechanical card link) plus bin/fleet-dashboard/* now select the dashboard family, so the suites protecting this change actually run when the files they protect change. * no-mistakes(review): sweep local card records on --resume-pending, reject zero timeout * no-mistakes(review): hold back only still-staged pairs on --resume-pending * no-mistakes(document): document handoff card record write and resume sweep * no-mistakes(review): fail loudly on card record rewrite errors, clear stale warnings ledger * no-mistakes(review): commit superseded marks after record write, gate ledger reset * no-mistakes(review): reject separator-bearing --card ids before they reach the record * no-mistakes(review): drop durable warnings ledger for in-process report suppression * no-mistakes(review): gate unreadable-card warning through reason-keyed report mark * no-mistakes(review): replace bash-4 assoc array with 3.2-portable report set * no-mistakes(document): scope outbox recovery claims, fix card-link report bound * fix(dashboard): stop flooding the fleet audit log for unretirable pairs An unretirable pending pair (the board says a card id does not exist) was writing a durable audit-log --fleet entry every time it was swept. bin/fm-bootstrap.sh's --resume-pending sweep now runs unattended on every session start, so a single permanently-unlinkable pair would grow that entry on a cadence no operator controls - the exact always-red-marker failure the fleet auditor's own design already guards against. The stderr warning stays exactly as built, still bounded in-memory per process; only the fleet-log write is removed. Surfacing a permanently unlinkable pair to the Admiral belongs to the fleet auditor's own sweep, raised once as a finding, not to this handoff path on every boot. * no-mistakes(review): handle unterminated final lines in handoff card state * no-mistakes(document): explain unterminated-line handling in handoff card state * fix(dashboard): never write the fleet audit log from --resume-pending The no-such-card case already stopped writing to the fleet audit log from bin/fm-bootstrap.sh's unattended --resume-pending sweep. This applies the same rule to the general transport-failure case in dashboard_link_card: a ref/agent/status write that fails still warns on stderr everywhere, but only an operator-initiated invocation (no --resume-pending) may write it to audit-log --fleet, and at most once per invocation. * no-mistakes(review): surface pending handoff card links at session start * no-mistakes(document): document spawn/teardown fleet-log writes and card-link resume gating * no-mistakes(document): note superseded pairs as a persisting card-link cause * no-mistakes(ci): drain the denied arm attempt before flipping the lock tests/fm-pi-watch-extension.test.sh's OpenCode session-lock case wrote a foreign pid, fired session.idle, slept a fixed 120ms, then took the lock and fired again. The event hook does not await its arm attempt, and that attempt walks the process tree with one `ps` per level, so on a loaded runner the denied attempt was still in flight at 120ms. ensureArm coalesces onto launchInFlight, so the second event joined the stale read-only answer and never armed - the shard's "expected exit 0, got 1". Wait on the plugin's own coordinator handle instead of the clock: ensureArmed joins an in-flight launch, so once it resolves no attempt is outstanding, and its status is the denial under test. Also pass the node output to expect_code so a future failure here shows why rather than an exit code alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes --------- Co-authored-by: joliverMI <joliver@sensibletech.biz> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#13) * fix(bin): count a trust-registered custom check as supervision-needed fm_supervision_status only looked for state/*.meta, state/x-watch.check.sh, and state/procevent/*.source. A home whose only duty is a check registered through bin/fm-check-register.sh (a state/<id>.check-trust binding paired with its state/<id>.check.sh) matched none of those, so FM_SUP_NEEDED stayed false and the Claude Stop auto-arm never armed a watcher for it. Count a state/<id>.check-trust file paired with its state/<id>.check.sh as supervision-needed, mirroring the existing presence-only style used for process-event sources and the X-mode relay poll. Update the turn-end guard and pull-warning banners to name a registered check specifically instead of falling through to the X-mode wording, and update docs/turnend-guard.md and its cross-references (docs/architecture.md, AGENTS.md, docs/scripts.md, docs/subagent-guard.md, docs/supervision-protocols/grok.md, and bin/fm-claude-stop-autoarm.sh's own header) to point at docs/turnend-guard.md as this predicate's single owner rather than restating its input list. * no-mistakes(document): note check-predicate rationale, generalize stale guard comments * no-mistakes(document): record check-predicate verification evidence entry --------- Co-authored-by: joliverMI <joliver@sensibletech.biz>
…14) * fix(watcher): re-evaluate arm coalescing when the fleet lock changes Two independent causes made the arm-readiness suite fail a different assertion almost every run: - ensureArm() in the OpenCode watch plugin reused a still-resolving earlier caller's beginArm() result unconditionally. Every ordinary session.idle produces two callers, so when the fleet lock was reacquired while an earlier attempt was mid-flight, the later caller inherited that attempt's stale read-only verdict and never armed. Fixed with premise-validated coalescing: a caller shares an in-flight attempt only while the lock file content it captured is still current; otherwise it evaluates fresh. Two callers on an unchanged lock still coalesce into one subprocess walk, so this costs nothing on the ordinary turn a serialized-everything fix would have doubled. - Both adapters spawn their arm child through a login shell, which sources /etc/profile in addition to the account's own profile files; the system-wide half is not relocatable via HOME. main already raised the readiness timeout to absorb that cost (250ms to 2000ms) and its own measurement shows a worst case of ~1740ms under contention against that budget - narrower headroom, not a removed confound. This adds FM_WATCH_ARM_NO_LOGIN_SHELL, a test-only opt-out that spawns the arm child under plain bash -c, removing the confound instead of padding around it. Production keeps the login shell as the unconditional default. Also fixes three unhandled-EPIPE crash sites found while proving this change under load: child.stdin.end() in fm-primary-turnend-guard.js, fm-primary-turnend-guard.ts, and fm-operational-input.js raises an unhandled 'error' event on the stdin stream (not the ChildProcess) when the child exits before the write lands, which was crashing the whole session process. Each site now no-ops that stream error since the child's own close/error handlers already drive resolution. Adds two regression tests (opencode coalescing on an unchanged lock, login-shell default vs. opt-out for both adapters) and one EPIPE regression test, and rewrites the existing OpenCode lock test to force the stale-verdict race deterministically via a gated ps shim instead of waiting for it. See docs/arm-readiness-determinism-proof.md for the repeated-run proof and docs/configuration.md for the new FM_WATCH_ARM_NO_LOGIN_SHELL entry. Built directly on current main (45bd292); does not touch main's own prior timeout raise or its test_pi_session_transition_generation_owner fixture-ordering fix, both kept as-is. * no-mistakes(review): correct stale arm-readiness proof claims and restore expect_code diagnostics * no-mistakes(review): wait on observable arm rows and share gate shims * no-mistakes(review): restore fallback diagnostics and bound unretired arm counts * no-mistakes(review): correct lock-assertion cause attribution in proof doc * no-mistakes(review): restore two-cause attribution and note login-shell residual bound * no-mistakes(review): qualify stale readiness-window claim on login-shell test * no-mistakes(review): correct stale readiness-budget claims in suite header * no-mistakes(test): correct stale login-shell contention figures in determinism proof * no-mistakes(document): scope proof-doc timeout raise and idle figure to sources * no-mistakes(document): stamp determinism proof verification at a15d993 with provenance --------- Co-authored-by: joliverMI <joliver@sensibletech.biz>
…not a hardcoded origin A firstmate checkout can have "origin" pointing at the public template it was forked from while its own branch tracks a separate fork it actually develops on, and the two can genuinely diverge. bin/fm-update.sh (via bin/fm-ff-lib.sh's ff_target) hardcoded origin/<default> as the fast-forward base, so a home tracking a fork never advanced past the point the two remotes diverged. Extracts the "which remote does this checkout actually develop on" resolution into bin/fm-dev-remote-lib.sh (resolve_update_base: the branch's configured upstream when set, falling back to origin/<default> and saying so otherwise) and applies it everywhere a hardcoded origin assumption had the same failure mode: bin/fm-spawn.sh's pooled-worktree base refresh (this is what stranded this task's own worktree 28 commits behind on the wrong lineage, missing bin/fm-dashboard.sh and fm-spawn.sh --card entirely - direct evidence caught live), bin/fm-review-diff.sh's review base and PR-head fetch, bin/fm-lint.sh and bin/fm-test-run.sh's --changed diff base, and bin/fm-teardown.sh's landed-work fallback checks (ensure_commit_object, content_in_default). bin/fm-fleet-sync.sh, bin/fm-merge-local.sh, bin/fm-home-seed.sh, and the remote-home provisioning scripts were audited and left unchanged: they correctly always use "origin" for project clones, which have no fork/upstream split by design. Recommends (does not apply - push/PR targeting is the captain's call): repointing the remote code root's own "origin" at the fork it actually develops on, matching the primary checkout; and, for no-mistakes PR mergeability, keeping "origin" pointed at the fork for ordinary fleet work and reserving the CONTRIBUTING.md --fork-url external-contributor recipe (push fork, PR against origin) for a separate checkout used only for genuine upstream contributions. tests/fm-remote-update-follows-fork.test.sh proves the fix end to end: a remote code root tracking a fork updates from it and a fork-only file reaches the secondmate home, confirmed to fail against the unfixed code.
Author
|
Opened in error against upstream - this work belongs on our own fork and is being redirected there. Apologies for the noise. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Make the fleet's self-update follow the repository we actually develop in, not the upstream it was forked from. A firstmate checkout can have "origin" pointing at the public template it was forked from while its own branch tracks a separate fork it actually develops on (e.g. this checkout: origin=joliverMI/firstmate is the fork we develop on, upstream=kunchenguid/firstmate is the public template), and the two can genuinely diverge. bin/fm-update.sh compared against origin/main unconditionally and never advanced a home tracking a different upstream past the divergence point. Scope:
Constraints: fast-forward-only safety is absolute and unchanged - never force, never stash, never discard unlanded work, never merge; a dirty or genuinely diverged target is still skipped and reported, not forced. Do not touch anything under projects/. The PR targets joliverMI/firstmate, never kunchenguid/firstmate.
What Changed
bin/fm-dev-remote-lib.shas the single owner of development-remote resolution:resolve_update_basereads the default branch's configured upstream (%(upstream:remotename)/%(upstream:remoteref)) and falls back toorigin/<default>with an explicit note when no remote upstream is configured.bin/fm-ff-lib.shnow fetches and fast-forwards through it (per-remotefetch_oncekeying,default_branchtakes a remote argument), so/updatefirstmateadvances a checkout that tracks a fork instead of comparing against a divergedorigin/main; the fast-forward-only guards (no force, stash, merge, or dirty/diverged advance) are untouched.fm-review-diff.sh(base ref andrefs/pull/<n>/headfetch),fm-teardown.sh(cachedresolve_dev_remotefor PR-object and landed-content checks),fm-spawn.sh(pooled-worktree base freshening), and the changed-file base infm-lint.shandfm-test-run.sh.fm-fleet-sync.shis documented as deliberately origin-only because it walks project clones underprojects/, which have no fork/upstream split; push and PR targeting are unchanged.tests/fm-remote-update-follows-fork.test.sh, which drivesfm-remote-secondmate-control.sh updateagainst a fork/upstream fixture and asserts fork-only files land in the secondmate home, plus new cases in the lint, review-diff, test-run, and spawn-pool-base-freshen suites andfm-test-run.shselection entries for the new library and test. Docs and skills were reworded from "origin" to "configured upstream", anddocs/remote-secondmates.mdrecords the recommendation to set the remote code root's branch upstream rather than repoint itsorigin.Risk Assessment
✅ Low: The change is broad (one new shared resolver plus six consumers touching self-update, pooled-worktree provisioning, teardown's landed-work gate, and diff-base selection), but every resolution failure mode falls back to the prior
originbehavior or refuses rather than forcing, fast-forward-only safety is untouched, each fixed path gained a real behavioral regression test, and only two narrow fail-safe issues survived verification.Testing
Ran the change's targeted suites (fork-follow e2e, fm-update, fm-lint, fm-review-diff, fm-test-run, fm-spawn-pool-base-freshen, fm-gotmp, fm-teardown, fm-secondmate-sync, fm-bootstrap, fm-fleet-sync) — all pass — and, because passing unit tests alone do not show the fleet actually following the fork, drove the real remote-update entrypoint end to end against a fixture built from this repository's own history, A/B-ing the pre-fix and post-fix bin/ over identical fixtures: the old path reported "already current" and left the secondmate home without bin/fm-dashboard.sh or --card, while the fixed path reported "updated 4f0a4f5..a28663a"/"synced:" and delivered both, with --card proven by its real argument-validation error rather than help text. I also captured the explicit no-upstream origin/main fallback note and confirmed fast-forward-only safety (diverged and dirty roots skipped, nothing forced, stashed, or discarded). No visual artifact applies: this change is git plumbing behind a CLI update path with no rendered surface, so the reviewer-visible evidence is the CLI transcript. Two failures I hit are environment-only and pre-existing at the base commit (locale-sensitive coverage guard, missing ruby); the working tree is clean and all evidence lives in the evidence directory.
Evidence: Fleet update follows the fork: before/after CLI transcript (fork-only dashboard + --card reach the remote secondmate home)
Source: Fleet update follows the fork: before/after CLI transcript (fork-only dashboard + --card reach the remote secondmate home)
### Fixture (real firstmate history) public template kunchenguid/firstmate main = 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5) development fork joliverMI/firstmate main = a28663a no-mistakes(review): repair gotmp fixture and teardown test selection fork-only commits ahead of the template: 12 ### BEFORE the fix (bin/ at b98e098) fm-remote-secondmate-control.sh update ios (as /updatefirstmate dispatches it to the host) current: 4f0a4f544dc0051dfaa06436a7d5e8e6ea0b4689 home HEAD : 4f0a4f5 fix(bin): give every status board pointer a real link or an honest gap (#5) bin/fm-dashboard.sh : ABSENT fm-spawn.sh --card : ABSENT -> error: unknown harness '--card='; pass a raw launch command to use an unverified adapter ### AFTER the fix (bin/ at a28663a) fm-remote-secondmate-control.sh update ios (as /updatefirstmate dispatches it to the host) synced: a28663aa13d756fb7433f4cb130dc5c4b321ef04 home HEAD : a28663a no-mistakes(review): repair gotmp fixture and teardown test selection bin/fm-dashboard.sh : PRESENT -> fm-dashboard.sh - the ONLY way an agent touches the Admiral's Fleet Dashboard. fm-spawn.sh --card : PRESENT -> error: --card requires a non-empty value ### The update line the captain actually sees (bin/fm-update.sh against the same code root) before fix: firstmate: already current after fix: firstmate: updated 4f0a4f5..a28663a (instructions changed: AGENTS.md, bin, .agents/skills) ### No upstream configured: keeps origin/main and says so firstmate: no upstream configured for main; using origin/main firstmate: already current ### Fast-forward-only safety is unchanged diverged code root -> firstmate: skipped: diverged from fork/main diverged code root -> unlanded commit still at HEAD, nothing forced or discarded dirty code root -> firstmate: skipped: dirty working tree dirty code root -> uncommitted edit still in the worktree, nothing stashedEvidence: Reproduction script for the end-to-end fork-follow demo (builds the fixture from real repo history, A/Bs pre-fix vs fixed bin/)
Source: Reproduction script for the end-to-end fork-follow demo (builds the fixture from real repo history, A/Bs pre-fix vs fixed bin/)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
.agents/skills/afk/SKILL.md- branch carries 13 commit(s) that exist on your local main branch but were never pushed to origin/main; rebasing would bundle this unrelated work (140 file(s)) into the PR:Push main to origin, or rebase your branch onto origin/main, before gating.
bin/fm-review-diff.sh:110- The PR-head fetch was repointed fromoriginto the resolved development remote, butrefs/pull/<n>/headonly exists in the repository the PR was opened against — which by standing repo convention isorigin(CONTRIBUTING.md:18-30: "clone the parent repo or set your localoriginback to the parent" ... "it pushes the branch to your fork and opens the PR against the parent repo"). The intent says "Do not change push or PR targeting behavior in this task; that remains owned by standing repo convention and is currently working correctly." Concrete failure: in exactly the code-root shape this commit recommends and its own fixture builds (origin = upstream template,maintrackingfork/main), a firstmate self-dev task withpr=.../pull/9makesfetch_pull_headfetchrefs/pull/9/headfrom the fork.resolve_pr_head(bin/fm-review-diff.sh:120-136) prefers any successful fetch unconditionally and never cross-checks the fetched SHA against the recordedpr_head=, so an unrelated PR docs: clarify live supervision constraints #9 on the fork silently becomes the compare tip and the review shows the wrong diff with no warning. bin/fm-teardown.sh:826 has the same repoint, though it fails safe there (the recorded head SHA won't be present, soensure_commit_objectreturns 1 and teardown falls back to the content check). Update lineage and PR base repo are independent concepts that only coincide when origin==fork; recommend leaving the PR-ref fetch onorigin(or resolving it from theno-mistakespush/PR config) and keepingresolve_update_basefor base/diff refs only.bin/fm-dev-remote-lib.sh:31-resolve_update_basederives the remote by splitting%(upstream:short)on the first/, which is not a remote name in every case. When<default>tracks a local branch (branch.main.remote = ., e.g. aftergit branch --set-upstream-to=dev main),%(upstream:short)prints a bare branch name with no slash, so${upstream%%/*}and${upstream#*/}both yield that branch name andRESOLVE_BASE_REFbecomesdev/dev. A non-default fetch refspec that mapsrefs/heads/mainon remoteforktorefs/remotes/foo/barbreaks it the same way. Every caller then treats a non-existent remote as authoritative:ff_target(bin/fm-ff-lib.sh:314) reportsskipped: no dev remoteand the firstmate self-update stops advancing at all, andfreshen_spawn_worktree_base(bin/fm-spawn.sh:1775) hard-refuses to launch any pooled task worktree — both cases previously worked offorigin/<default>. Read the remote and branch directly instead:%(upstream:remotename)and%(upstream:remoteref)(striprefs/heads/), and fall back to the origin path whenremotenameis empty or..bin/fm-test-run.sh:996-bin/fm-dev-remote-lib.shis not enumerated infamilies_for_changed_path, so it falls through to thebin/*)branch and is mapped byfamilies_for_test_reference— a literal grep of every test for the basename. Only tests/fm-test-run.test.sh mentions the file (the threecplines added in this commit), so editing the shared resolver selectspure-contract-unitand nothing else. That silently skips every other suite that now depends on it: tests/fm-review-diff.test.sh and tests/fm-teardown.test.sh (pr-forge), tests/fm-update.test.sh (session-bootstrap), tests/fm-secondmate-sync.test.sh (secondmate), the spawn suites (backend-dispatch), and the new tests/fm-remote-update-follows-fork.test.sh (unclassified). A regression inresolve_update_basewould passbin/fm-test-run.sh --changed, which contradicts that map's own contract ("Conservative path → family map. Over-selects rather than under-selects"). Addbin/fm-dev-remote-lib.shto the explicit case list next tobin/fm-ff-lib.shwith the families of its real consumers.bin/fm-spawn.sh:1754-freshen_spawn_worktree_basestill hard-requiresoriginbefore it resolves the development remote: an unconditional fullgit fetch origin(line 1754) andgit remote set-head origin --auto(line 1758), both fatal on failure. The intent's scope item 2 asks to fix "every other place in bin/ that assumes origin is the development remote" under the same resolution, and this is the function the commit message names as the one that stranded the task worktree. Concretely: a checkout whose development remote isforkand whoseorigin(upstream template) is unreachable, renamed, or absent now refuses to launch any pooled task worktree even thoughfork/mainis fetchable — and in the healthy case it pays a redundant full fetch of a remote the base no longer comes from. Resolve the remote first, fetch/set-headthat remote, and only consultoriginfor the default-branch name when the resolved remote can't supply one.bin/fm-teardown.sh:823-ensure_commit_objectonly needs a remote, but now callsdefault_branchpurely to feedresolve_update_base, and returns 1 when it fails. In a project clone with norefs/remotes/origin/HEADand no localmain/master(bin/fm-teardown.sh:626-640), a merged PR that previously confirmed viafetch origin refs/pull/<n>/headcan no longer be confirmed;content_in_defaultthen fails on the same missing default, sowork_is_landedreturns false and teardown refuses work that has in fact landed. The direction is safe but the reachable set is newly narrowed. ResolveDEV_REMOTEonce at top level afterDEFAULTis already computed (the pattern bin/fm-review-diff.sh:78-80 uses) rather than re-deriving it per call inside bothensure_commit_objectandcontent_in_default.docs/scripts.md:90-bin/fm-dev-remote-lib.shis a new shared*-lib.shbut is not listed in docs/scripts.md, which README.md:218 calls "thebin/toolbelt reference" and which enumerates every other sourced helper includingfm-ff-lib.sh(line 90) andfm-tangle-lib.sh(line 86). Add a row so the new single owner of development-remote resolution is discoverable next to the fast-forward helper that sources it.docs/remote-secondmates.md:214- Intent item 5 requires proving "end to end, not just via code inspection" that "a fleet update must leave fork-only files (bin/fm-dashboard.sh and fm-spawn.sh's --card flag) present in a remote secondmate home that updates through the fixed path." The delivered proof is a synthetic fixture (tests/fm-remote-update-follows-fork.test.sh) that constructs a code root withforkadded andmainset to trackfork/main, using a stand-in file namedfm-dashboard.sh. The change itself deliberately does not repoint the real remote code root (intent item 4 forbids applying it), and the new doc line states the residual gap plainly: "a code root whose upstream resolves to a different repository than the primary checkout's own fork updates from the wrong lineage with no error." So the real remote home only receives fork-only files once its origin is repointed or itsmainupstream is configured out of band. Worth confirming with the captain whether the fixture proof satisfies item 5 or whether a live run against the real remote secondmate is expected before this lands.🔧 Fix: harden dev-remote resolution and its consumers
5 issues (1 warning, 4 infos) still open:
bin/fm-test-run.sh:990- The newbin/fm-dev-remote-lib.shchanged-path case claims to select "spawn's pooled worktree base" coverage via thebackend-dispatchfamily, buttests/fm-spawn-pool-base-freshen.test.shhas no entry infamily_for_basename(lines 137-235) and therefore resolves tounclassified, which no case emits. Concretely: editresolve_update_basesofreshen_spawn_worktree_basepicks the wrong remote, runbin/fm-test-run.sh --changed, andbackend-dispatchselects fm-spawn-batch/-dispatch-profile/-worktree-settle but never fm-spawn-pool-base-freshen.test.sh - the suite docs/architecture.md:162 names as the owner of the pooled-worktree base-freshness regression coverage, and the very suite this fix round extended withtest_fork_tracking_pool_refreshes_from_fork_not_origin. The regression passes --changed silently, contradicting that map's own "over-selects rather than under-selects" contract. Emitprintf '%s\n' "__script__:fm-spawn-pool-base-freshen.test.sh"from this case (the__script__:marker is already handled at line 862 and consumed at line 1107), or classify the suite intobackend-dispatchinfamily_for_basename.bin/fm-dev-remote-lib.sh:29- The header comment states that with a non-default fetch refspec%(upstream:...)"prints a tracking path whose first component is not a remote name at all" and that such a checkout therefore "take[s] the announced origin fallback". That is true of the old%(upstream:short)split but false of the code it now documents:%(upstream:remotename)returns the real remote name and%(upstream:remoteref)returnsrefs/heads/<branch>regardless of the fetch refspec (verified on git 2.43), soresolve_update_basereturns<remote>/<branch>with an empty NOTE. With e.g.remote.fork.fetch = +refs/heads/*:refs/remotes/mirror/*,ff_target's origin mode runs the plainfetch_once(populatingrefs/remotes/mirror/*), then fails the existence check at bin/fm-ff-lib.sh:331 and reportsskipped: fork/main does not existinstead of the fallback the comment promises. It fails safe, but the comment misdescribes the behavior a maintainer would rely on. Either correct the comment to say only the local-branch (.) and no-upstream cases fall back, or additionally guard the returned ref.docs/architecture.md:161- "no worker starts until its clean task worktree matches the fetched tip of origin's resolved default branch" is now stale:freshen_spawn_worktree_base(bin/fm-spawn.sh:1752-1799) resolves the remote from the default branch's configured upstream and can provision fromfork/mainwith origin never contacted. The change updated the "Self-updates stay safe" section of this same file (lines 327-328) and docs/remote-secondmates.md, so this line is an omission rather than a deliberate carve-out. Reword to "the development remote's resolved default branch (the configured upstream, falling back to origin)".bin/fm-test-run.sh:1011-bin/fm-ff-lib.shmaps to onlypure-contract-unit, and because that is an explicit case it shadows thebin/*reference scan. This change materially alteredff_target's base resolution,fetch_once's dedup key, anddefault_branch's signature - whose actual owning suites aretests/fm-secondmate-sync.test.sh(familysecondmate, sources fm-ff-lib.sh directly and asserts FF_STATUS/FF_INSTR) andtests/fm-update.test.sh(familysession-bootstrap). Neither is selected by--changedwhen only fm-ff-lib.sh is edited. The mapping is pre-existing, but the fix round corrected exactly this class for the new sibling lib while leaving the now-broader-blast-radius one wrong; addsecondmateandsession-bootstrapto this case.bin/fm-teardown.sh:656-resolve_dev_remote, and the rerouting ofensure_commit_object(line 852) andcontent_in_default(line 950) onto it, changed which remote decides whether a worktree's work has landed - the check that gates discarding unpushed commits. Every other consumer touched by this change got a fork-tracking regression case in the same commit (fm-review-diff, fm-spawn-pool-base-freshen, fm-lint, fm-test-run, fm-remote-update-follows-fork), buttests/fm-teardown.test.shwas not extended, so the diverged-forkcontent_in_defaultcomparison and the resolve-once caching (including theDEV_BRANCHempty fallback at line 951) have no coverage. The failure direction is safe (an unresolvable remote makescontent_in_defaultreturn non-zero and teardown refuse), which is why this is informational rather than blocking.🔧 Fix: correct dev-remote docs and test-selection coverage
3 issues (1 error, 2 infos) still open:
tests/fm-gotmp.test.sh:79-bin/fm-teardown.sh:195-196now sources$SCRIPT_DIR/fm-dev-remote-lib.shunconditionally underset -eu(line 160), buttests/fm-gotmp.test.shbuilds its fake root by symlinking each sourced sibling one by one and was not updated. Concrete failure:make_fake_root(line 47) doesln -s "$TEARDOWN" "$fake/bin/fm-teardown.sh"and symlinks fm-backend/fm-tmux-lib/fm-nm-run-lib/fm-lock-lib/fm-control-lib/fm-classify-lib/fm-wake-lib/fm-gate-refuse-lib/fm-pr-lib/fm-public-followup-lib/fm-x-lib/fm-secondmate-registry-lib/fm-secondmate-parent-lib but not fm-dev-remote-lib.sh;SCRIPT_DIRresolves to$fake/bin(dirname of BASH_SOURCE[0], the symlink path), so the.of the missing file returns 1 andset -eaborts.test_teardown_removes_tasktmp_dir(line 124) runsbash "$fake/bin/fm-teardown.sh" "$id" || fail "teardown exited non-zero with a valid tasktmp"and fails, as do the sibling cases. The second inline fixture builder (around line 157, used bytest_teardown_skips_gracefully_without_tasktmp) has the same omission. Fix: addln -s "$ROOT/bin/fm-dev-remote-lib.sh" "$fake/bin/fm-dev-remote-lib.sh"to both builders. This is the only fixture affected - every other test either uses$ROOT/bindirectly, symlinks the wholebindir (tests/fm-session-start.test.sh:602,649),cp -Rs it (tests/fm-teardown.test.sh:1667), or was already updated (the threecp "$RUNNER"sites in tests/fm-test-run.test.sh).bin/fm-test-run.sh:950- Thebin/fm-teardown.shchanged-path case emits onlypr-forgeanddashboard, buttests/fm-gotmp.test.sh- which runs the real teardown binary end to end against a hand-built fakebin/and is therefore the suite most sensitive to teardown's sourced-sibling set - is classifiedsession-bootstrap(line 183). So a teardown-only edit that adds or removes a sourced lib never selects the suite it breaks; that is exactly why the fixture drift above is invisible tobin/fm-test-run.sh --changedfor teardown edits. (This change happens to be caught only because the newbin/fm-dev-remote-lib.shcase at line 990 does emitsession-bootstrap.) The round-2 fix corrected this same class forbin/fm-ff-lib.sh(line 1011); addsession-bootstrapto the teardown case for the same reason.docs/scripts.md:90- Thefm-ff-lib.shrow still reads "Shared guarded fast-forward helper for origin pulls and local secondmate syncs", which this change makes inaccurate:ff_target'soriginbase mode now resolves the base throughresolve_update_baseand pulls from the checkout's configured upstream (e.g.fork/main), withorigin/<default>only as the announced fallback. The change updated the neighbouring prose in docs/architecture.md (lines 161, 327-328), docs/remote-secondmates.md, and .agents/skills/updatefirstmate/SKILL.md, and added thefm-dev-remote-lib.shrow directly beneath this one, so this is an omission rather than a deliberate carve-out. Reword to "for configured-upstream pulls (falling back to origin) and local secondmate syncs".🔧 Fix: repair gotmp fixture and teardown test selection
2 infos still open:
bin/fm-teardown.sh:950-resolve_dev_remoteresolves the remote from$PROJ(bin/fm-teardown.sh:660 ->resolve_update_base "$PROJ" "$name"), but both consumers then use that name as a remote in$WT:ensure_commit_object(line 852) runsgit -C "$WT" remote get-url "$DEV_REMOTE"andcontent_in_default(line 952) runsgit -C "$WT" fetch "$DEV_REMOTE" .... For task teardown those are the same repository, but for a secondmate they are not: fm-spawn.sh:634-636 recordsworktree=$homeandproject=$root, i.e. the persistent secondmate home and the code root are separate clones. Concrete failure: a code root whosemaintracksfork/mainwhile the secondmate home clone has onlyorigin. A secondmate home holding a local commit not on any remote reacheswork_is_landed->content_in_default;git -C "$WT" remote get-url forkfails, so it drops to theelifand compares against the home's possibly stale localrefs/heads/maininstead of a freshly fetched default, and teardown REFUSES work that has in fact landed on the fork. Previously the hardcodedoriginexisted in both repositories, so this path always fetched. The fail direction is safe (refuse, never discard), which is why this is informational. Resolve the remote from$WT- the repository whose objects are actually being fetched and compared - rather than from$PROJ, or fall back tooriginwhen$DEV_REMOTEis not a remote of$WT.bin/fm-spawn.sh:1788- The secondresolve_update_base "$worktree" "$default"is keyed on the REMOTE's default-branch name (line 1781, read fromrefs/remotes/$remote/HEADafterset-head --auto), not on the local branch whose upstream produced$remotein the first place. When those names differ,refs/heads/$defaulthas no upstream locally and the resolver falls back toorigin, silently undoing the fork resolution the first pass made. Concrete path: a checkout whose localmaintracksfork/mainwhilefork's own HEAD isdevelop- pass 1 resolvesfork,set-head fork --automakes$default=develop,refs/heads/developdoes not exist, soremoteflips tooriginand the pooled worktree isreset --hardtoorigin/develop(the diverged template lineage this change exists to stop using), or refuses if origin is unreachable.RESOLVE_BASE_NOTEis discarded here, so nothing in the error or the success path explains why origin was used. Suggest keeping pass 1's remote when pass 2 falls back tooriginwhile pass 1 resolved something else, and echoingRESOLVE_BASE_NOTEwhen it is non-empty.bin/fm-test-run.sh:606-bin/fm-test-run.sh --check-coverageexits 1 under this machine's LANG=en_US.UTF-8: itscomminvocations run in the ambient locale while their inputs are sorted withLC_ALL=C, so comm reports "file 2 is not in sorted order". This aborts tests/fm-test-run.test.sh mid-suite. It is NOT a regression — it reproduces identically from a clean checkout of base commit b98e098, and both base and target pass withLC_ALL=C bin/fm-test-run.sh --check-coverage. I re-ran the suite under LC_ALL=C to validate this change rather than editing the shared runner, since the bug is pre-existing and unrelated to the fork-follow intent; the author may want to pin LC_ALL=C on those comm calls in a separate change.tests/fm-test-run.test.sh:674- tests/fm-test-run.test.sh's CI-workflow YAML check hard-fails whenrubyis absent ("not ok - ruby is required to parse .github/workflows/ci.yml as YAML"). ruby is not installed in this sandbox and installing system packages is out of bounds here, so this one case cannot be exercised locally. It is environment-only, independent of the change under test, and every other case in that file passes.bash tests/fm-remote-update-follows-fork.test.sh— new e2e: remote code root tracking a fork updates from it and the fork-only file reaches the secondmate homebash tests/fm-update.test.sh— /updatefirstmate mechanics incl. off-default/detached/unsafe-home skipsbash tests/fm-lint.test.sh,bash tests/fm-review-diff.test.sh,bash tests/fm-spawn-pool-base-freshen.test.sh,bash tests/fm-gotmp.test.sh— the other resolve_update_base consumersLC_ALL=C bash tests/fm-test-run.test.sh— includes "--changed defaults to main's configured upstream, not a hardcoded origin/main" (17 ok; only the ruby-dependent CI-YAML case fails)bash tests/fm-fleet-sync.test.sh— confirms project-clone sync (deliberately origin-only) is unregressedLC_ALL=C bash tests/fm-teardown.test.sh,tests/fm-secondmate-sync.test.sh,tests/fm-bootstrap.test.sh— remaining fm-ff-lib consumersManual A/B e2e: fixture from real history (template=4f0a4f5 pre-dashboard, fork=a28663a), ranFM_ROOT_OVERRIDE=<code-root> FM_HOME=<home> bin/fm-remote-secondmate-control.sh update ioswith pre-fix bin/ (extracted from b98e098) vs fixed bin/, then checked the home forbin/fm-dashboard.sh --helpandfm-spawn.sh ... --card=flag behaviorManual:bin/fm-update.shagainst a code root withgit branch --unset-upstream mainto confirm the announced origin/main fallbackManual safety:bin/fm-update.shagainst a diverged code root (extra local commit) and a dirty code root, verifying skip messages and that neither the unlanded commit nor the uncommitted edit was touchedbash /tmp/fm-base-checkout/bin/fm-test-run.sh --check-coverage— reproduced the locale failure at base commit b98e098 to confirm it is pre-existingbin/fm-remote-secondmate-control.sh:261- bin/fm-remote-secondmate-control.sh:261 still dies with "remote code root did not complete a safe origin update" after the code root moved to configured-upstream resolution; the message names the wrong remote when a code root tracks a fork. I left it because it is a user-facing runtime string rather than documentation, and this phase may not change executable behavior. Suggested wording: "remote code root did not complete a safe update".🔧 Fix: reword remaining origin-as-dev-remote strings and comments
1 info still open:
tests/fm-secondmate-sync.test.sh:4- tests/fm-secondmate-sync.test.sh header/case comments (lines 4, 14, 236, 809) still describe the local-HEAD sync as "no origin fetch" and say standalone-clone homes converge through "/updatefirstmate's origin fetch". The assertions themselves are correct (they prove no fetch happens at all), so this is comment prose only. I left it because this phase is explicitly forbidden from editing tests; the equivalent wording in bin/ and the skill was updated to "remote fetch" / "fetch-based". Suggested follow-up in a phase allowed to touch tests: same one-line reword.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.