Conversation
…ng state (#2) * fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413) A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload. * fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary. * docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes * fix(calm): keep Pi's export confirmation visible (kunchenguid#2461) Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared. * feat(bin): render the captain's four-section status board from existing state fm-status-board.sh prints one HTML page (Needs you to continue, In progress, Waiting, Recently completed) rendered entirely from fm-bearings-snapshot.sh's already-correct cross-home classification and fm-fleet-snapshot.sh's untruncated per-item detail, with no second store for agents to keep in sync. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
Tearing down a task deletes state/<id>.meta, the record bin/fm-pr-merge.sh resolves the PR through. A branch can be fully pushed and reachable from a remote (so the existing dirty/unpushed/landed checks find nothing to refuse) while its own PR still sits open, stranding a mergeable PR with no guarded path to land it - this happened twice in one night. The new check runs after the existing dirty/unpushed/landed checks inside validate_worktree_teardown_safety, so their refusal messages take precedence and never contradict this one. It fires only when a pr= is actually recorded (scout, local-only, and secondmate teardowns never record one). A gh lookup error refuses loudly too, naming the PR firstmate could not confirm, rather than tearing down blind. --force already skips the dirty/landed checks and now explicitly skips this one too, since it is the same "captain says discard" escape hatch - the usage comment spells out what that means here (the PR needs merging or closing by hand, outside fm-pr-merge.sh's guarded path). Co-authored-by: joliverMI <joliver@sensibletech.biz>
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413) A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload. * fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary. * docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes * fix(calm): keep Pi's export confirmation visible (kunchenguid#2461) Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared. * fix(procevent): deliver captured results at most once A Lavish-captured answer reached a secondmate five times because forwarding it ran fm-send.sh and only afterward, manually, remembered to call `handled` - so every re-announcement of the still-unacknowledged capture (a restart, a compaction, a re-read wake) repeated the forward. The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves capture is exactly-once, but nothing paired a downstream effect with its acknowledgement atomically. Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`, which checks handled status, runs the command, and marks the generation handled as one call under the source lock, so a repeated delivery attempt against an already-delivered generation runs the command zero times and a failed command stays eligible for retry instead of being dropped. Point the process-event-sources skill and the runner's operating contract at it for exactly this class of forward. Five delivery attempts against one captured generation now run the downstream command once instead of five times - the read-and-discard cost of the other four is eliminated structurally. The historical incident's own token cost was not measured at the time and is not reconstructable after the fact. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
…#4) Adds a fifth board section between "Needs you to continue" and "In progress" for finished work that is deployed and usable, not merely merged. It renders only when firstmate has written a "ready: <where to see it>" note on a landed backlog row, confirming deployment and naming the surface - never inferred from merge/landed state alone, so a merged-but-not-yet-deployed change never appears here. The item shows its pointer text, never a PR link or repo path. Approval buttons (part 2) are not implemented in this PR: investigation found Lavish's artifact bridge (window.lavish.queuePrompt) can queue an exact, single-item approval string, but transmission still needs the existing manual "Send to Agent" tap by Lavish's own design, and the delivery path could not be confirmed end-to-end from this environment. Ship the verified section now and report the button investigation as a plan. Co-authored-by: joliverMI <joliver@sensibletech.biz>
#5) The board named locations the captain could not reach from his phone: a pull request he does not review, a raw data/<id>/report.md path, a "fuller detail lives in the X home record" cross-home line with nothing to tap, and a bearings-truncated in-progress summary with no full text behind it. Pull requests and GitHub links are now never rendered anywhere on the page. A report links only when its backlog note carries an explicit report-url: marker firstmate wrote after actually publishing it (the same convention as the existing ready: marker) - investigated whether Lavish's existing route could serve reports directly and confirmed it cannot: Lavish is a publish-one-artifact-at-a-time surface, not a path-mapped static file server, so there is no passive route to piggyback on. A Ready item's "where to see it" text is flagged honestly instead of shown verbatim when it looks like a filesystem path or a pull-request URL. A cross-home item's bounded detail says plainly that the rest is not reachable from the page, instead of naming an unreachable home record. An in-progress row now prefers its fuller current-state text over the shortened one-line summary. Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488) * feat(stow): persist the open records a session is holding /stow curated memory and captured session knowledge, but never touched record state, while AGENTS.md called it an "unfinished-work sweep" and the receipt declared the session "safe to reset" - wording that implied a record-correctness guarantee stow does not make. A shipped PR with no backlog item, a queued umbrella whose phases had merged, and four decision holds left open after their answers shipped all survived repeated stows. Add a bounded pass that files record state from the same volatile input the rest of stow already uses: the open threads in context, minutes before the reset destroys them. It creates a record for an unfiled thread and corrects one the session knows is wrong, through the owning path, and states its boundary as part of the contract - it never enumerates the backlog, lists holds, or queries a forge, because it cannot be a reconciliation and must not be read as one. Correct the wording in AGENTS.md and the completion receipt so reset-safe means what it actually guarantees: nothing this session knew was lost. * no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi * no-mistakes(document): note /stow open-record persistence in README command catalog * refactor(stow): state open-record persistence as principle, not procedure The first version enumerated triggers, named commands, and prescribed an ordered procedure. That is too rigid for an agent skill: it invites literal execution of a checklist instead of judgment, and every enumerated example is a way for the guidance to go stale. Reduce it to the intent - before a reset, the important open work you are holding in context must end up durably recorded rather than dying with the session, filing what is unfiled and correcting what is stale - and let the agent judge importance, the record, and the owning write path. Keep the scope bound, since it is a decided contract and not a mechanic: this covers the open work the session is holding, never a reconciliation of durable records against repository or forge reality. The wording corrections in AGENTS.md and the completion receipt are unchanged. * fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490) * fix(decisions): close captain holds at answer time Firstmate had two "a decision is open" ledgers with asymmetric closing mechanics. The live status-log ledger closes atomically at answer time, because bin/fm-send.sh --resolve-key makes answering a decision be the act that closes it. The durable backlog hold ledger had no such coupling: answering and recording were two separate acts, and only the first was forced by the workflow. That asymmetry lost four real captain decisions. Their answers were captured durably to disk, keyed character for character by the hold decision keys, acknowledged, and even implemented and shipped, yet the holds stayed open for two days and the captain was asked to re-answer decisions already on his own disk. Give the hold ledger the same answer-time-closure property: - bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's counterpart to --resolve-key. It shares one unrouted close implementation with `decline`, so it carries every existing guard - the captain decision file, the active-hold requirement, retry identity, and the refusal to release still-routed work - and differs only in the resolution mode it records. `decline` keeps its stronger meaning that the answer routes no follow-up work at all. - bin/fm-procevent-lavish.sh wires the channel that actually carried the lost answers. `arm --decisions-origin` binds a deck to the origin whose holds it carries, `answers` reads the structured choices out of a captured poll result, `close-decisions` maps each key to its hold and closes it through the command above, and `autohandle` lets the runner apply that at capture time. Safety is preserved rather than traded away. Only rows tagged `choice` are read, so freeform captain prose cannot forge a decision key. Closure is confined to the one bound origin. The decision text is a pure function of the captured result, so a replayed capture is idempotent. A hold that is absent, already closed, or still blocking routed work is skipped and left for `resolve`, never forced. A deck armed without the binding touches no hold at all. And autohandle deliberately never reports full handling, because recording an answer is transcription while acting on it is firstmate's judgement - so the check wake still reaches the handler. fm-send --resolve-key is untouched. * no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory * refactor(decisions): make keyed-answer closure one general capability The previous pass gave holds answer-time closure but built it as bespoke Lavish wiring: the review adapter carried the source-to-origin binding, mapped keys to hold identities, wrote decision records, decided what to skip, and closed holds itself. That treated a review deck as a special decision source. It is not - it is an ephemeral discussion format that happens to carry answers. Collapse it into ONE general capability with one owner. bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its matching hold": - `answers <origin> --source <provenance>` is the channel-agnostic intake. It reads key/answer/label lines on stdin, maps each key to its hold, and closes it through the same `answer` path, so every guard applies identically whatever channel the answer came from. --source is provenance recorded in the decision, never a behavior switch; there is no per-channel branch and no knowledge of chat, decks, or transports. - `bind`/`unbind`/`binding` own the source-to-origin binding for any channel whose answers arrive detached from their origin. Every channel is now an ordinary caller that only turns what it received into keyed lines: - bin/fm-send.sh (chat) feeds the intake for a key that names an active hold. This also fixes a real gap: once `complete` transfers a decision to its hold it closes the live status copy, so --resolve-key alone could never answer a transferred decision. - bin/fm-procevent.sh feeds it generically. A bound source's captured result goes to `<adapter> answers <result-file>` and whatever that prints is piped into the intake. The runner names no adapter, parses no result, and carries no decision rule, so any future adapter with an `answers` command works with no change here. - bin/fm-procevent-lavish.sh keeps only `answers`, which reports the structured choices a review captured and stops. It maps nothing to a hold and closes nothing; it lost ~160 lines of decision logic. Feeding is independent of handling, so it never acknowledges a result and never suppresses a wake - recording an answer is transcription, acting on it stays firstmate's judgement. The regression that proves closure now drives a FIXTURE adapter that is not the review adapter, so what is proven is that any bound channel reaches the intake rather than that one channel is wired specially. A new regression drives the real fm-send over a stubbed transport for the chat side. Every prior guarantee still holds, and fm-send's status-log behavior is unchanged. * no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression * fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512) A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md. The installer now creates and migrates to a recoverable two-line pointer file. * fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515) * feat(dashboard): add the Admiral's Fleet Dashboard task board Replaces the generated Lavish status page with a purpose-built board: a zero-dependency Python backend (stdlib http.server + sqlite3), a vendored React frontend styled to match Spectra, an agent-facing CLI (bin/fm-dashboard.sh) that is the only way agents touch it, and the fleet-dashboard skill teaching the fleet auditor and other agents how to use it. The board owns its own persistent records rather than mirroring any backlog - its six-status vocabulary and four review tabs don't map onto tasks-axi state, and the fleet auditor needs a separate claim to check against live reality. See docs/dashboard.md for the full architecture, the deploy story, and the drift-risk tradeoff that choice implies. Link posting rejects GitHub/PR links and local-only hosts structurally (standing order 17), not just by agent memory. The page shows a loud, persistent banner when it cannot reach its own API, and the auditor's "never run" state is rendered distinctly from a clean sweep, so an absence of confirmation is never shown as a confirmed-good state. --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: joliverMI <joliver@sensibletech.biz>
Testing and needs-attention were conflated: both put a finished-enough card in front of the Admiral, but only needs-attention actually blocks on him. A card can now be set to needs-attention with a --reason that renders directly on the card, sorts above every other status, and gets its own loud always-visible section on the page. The fleet auditor's sweep now treats a needs-attention card's age as a finding, since that status - unlike testing - is not supposed to sit unanswered. Existing databases migrate in place; the new column is nullable and existing cards, history, and statuses are untouched. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…orce Audit button (#8) The auditor's 15-minute cadence never actually ran: it registered a durable check in its own home, but nothing polled it because that home had no live supervision session. Host cron (already running independent of any chat session) now drives bin/fm-fleet-audit-tick.sh, which heartbeats every tick and sweeps via bin/fm-fleet-audit-sweep.sh once the dashboard's own interval setting says it's due, self-healing a lock abandoned by a crashed sweep. The Force Audit button uses the identical sweep script through a detached subprocess so a scheduled and a forced run can never stack. Also fixes the "unassigned" agent noise on every card (shows the captain instead) and adds a Go to card jump from each discrepancy log entry. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…ard (#9) Every dashboard card's agent and backlog_ref sat empty, so status transitions only happened when someone remembered to make them by hand - worst of all at teardown, the exact moment the worker that would have remembered is being removed. fm-spawn.sh --card <card-id> now records dashboard_card=<id> in the task's own state/<id>.meta and best-effort links the card's ref/agent to the task identity, advancing a not_started card to working. fm-teardown.sh reads that identity back and advances the card to testing once a real (non-forced) landing succeeds, leaving scout, secondmate, and already-complete cards untouched. No card is the normal case and stays a complete no-op. A bad card id or an unreachable dashboard never fails the spawn or the teardown; it warns loudly and best-effort records fm-dashboard.sh audit-log --fleet so a failed link is visible to the fleet auditor rather than silently dropped. All 21 live cards still carry no agent/ref - 20 belong to a project this task has no access to verify task identities for, and the one firstmate-owned card is already complete with no live task to link - so none were backfilled; see the PR description for the per-card list. Co-authored-by: joliverMI <joliver@sensibletech.biz>
…card Captain-routed cards had no mechanism to advance at all: bin/fm-spawn.sh --card and bin/fm-teardown.sh only fire off local task metadata, and bin/fm-backlog-handoff.sh had no dashboard integration whatsoever, so a card handed to a secondmate froze at not-started forever. bin/fm-backlog-handoff.sh --card <card-id> now gives a single handed-off item the same best-effort link: once it has genuinely landed in the secondmate's backlog, it sets the card's ref/agent and advances a not_started card to working. Since a handed-off item has no local task metadata to hold dashboard_card=, the id is instead recorded as a `dashboard_card:` body line on the item itself, via `tasks-axi update --body-file`, so it survives the move and any remote delivery retry. Reaching testing still depends on the secondmate later passing that same --card id to its own fm-spawn.sh/fm-teardown.sh cycle against the shared board - the same reliance every other --card use already has, now also covering the handoff case rather than a new, weaker guarantee.
…und dashboard calls
…ry delivered card
joliverMI
force-pushed
the
fm/fm-dashboard-card-link-on-handoff
branch
from
August 19, 2026 02:37
2dc80c5 to
acf91c4
Compare
Author
|
Opened in error by our own tooling against the wrong base repository - this work is for our fork and was never intended as an upstream contribution. Closing; apologies for the noise. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What Changed
bin/fleet-dashboard/, driven by the newbin/fm-dashboard.sh(cards, seven statuses includingneeds-attention, per-card tabs, fleet audit log), plusbin/fm-fleet-audit-sweep.sh/bin/fm-fleet-audit-tick.shto reconcile each card's claimed state against live crew/backlog state, andbin/fm-status-board.shfor the captain's four-section board.fm-spawn.sh --cardwritesdashboard_card=intostate/<id>.meta, sets the card'sbacklog_ref/agent, and advancesnot_started→working;fm-teardown.shreads that back and advances the card totestingafter a non---forceteardown succeeds (never downgrading acompletecard);fm-backlog-handoff.sh --cardlinks a single handed-off item to a secondmate throughstate/handoff-cards/<secondmate-id>. Every dashboard call is best-effort — failures warn on stderr and record aaudit-log --fleetentry rather than blocking a spawn, teardown, or handoff.fm-teardown.shnow refuses while a task's recorded PR is still open (and on an unconfirmableghlookup);fm-procevent.shgains adeliversubcommand that runs a paired external effect at most once under the source lock and feeds keyed captain answers intofm-decision-hold.sh, as doesfm-send.sh --answer;CLAUDE.mdbecomes a real@AGENTS.mdpointer file instead of a symlink, withfm-ensure-agents-md.shand the CI pointer check updated to match;fm-test-run.shgains adashboardtest family,LC_ALL=Ccoverage-guard comparisons, and a 20-minute cap on the real-Herdr CI step; new and extended shell suites plusdocs/dashboard.mdand thefleet-dashboardskill document the above.Risk Assessment
✅ Low: The round-4 redesign removed the whole body-rewrite hazard class by moving the pairing into firstmate-local state, every board call is bounded and non-fatal by construction, the guarded/unguarded relink and co-staged link semantics are correct on all five call sites I traced, and the seven handoff tests assert real board end-state on the previously-uncovered non-empty-destination, non-empty-outbox, resume, and card-less-flush paths; the one residual issue is a narrowly-reachable exit-status leak.
Testing
Ran the targeted dashboard/handoff/teardown/test-runner suites (all green, including the previously-failing locale and CI-YAML-parser cases now fixed in-tree), then demonstrated the intent as an operator would see it: against a real dashboard server, a
--cardhandoff moved a not_started card to working with refdesign:phone-status-pageand agentdesignwhile leaving the handed-off item byte-identical, a re-run refused to re-claim the link, an unresolvable card id warned loudly and landed a fleet discrepancy finding without blocking the move, and a landed teardown advanced its linked card to testing unattended; screenshots of the rendered board capture both transitions and the discrepancy log. Temp homes, the demo server, and the browser session were cleaned up and the worktree is unchanged.Evidence: CLI transcript: handoff --card links the card end-to-end, is idempotent on re-run, and is loud on an unresolvable card id
Source: CLI transcript: handoff --card links the card end-to-end, is idempotent on re-run, and is loud on an unresolvable card id
$ bin/fm-backlog-handoff.sh design phone-status-page --card make-the-status-page-readable-on-his-pho-absa handed off 1 item(s) to design: phone-status-page dashboard: linked card make-the-status-page-readable-on-his-pho-absa to phone-status-page (ref=design:phone-status-page, agent=design, status not_started -> working) $ bin/fm-dashboard.sh show make-the-status-page-readable-on-his-pho-absa status: working agent: design ref: design:phone-status-page $ bin/fm-backlog-handoff.sh design phone-status-page --card <same card> # re-run dashboard: card ... already links to design:phone-status-page (agent=design); left unchanged $ bin/fm-backlog-handoff.sh design unrelated-thing --card no-such-card warning: dashboard card link failed for unrelated-thing -> card no-such-card (ref): fm-dashboard.sh: server refused (404): no such task: 'no-such-card' handoff exit: 0 (the item still moved - a dashboard problem never blocks the handoff)Evidence: CLI transcript: teardown consumes dashboard_card= from task meta and advances the card to testing
Source: CLI transcript: teardown consumes dashboard_card= from task meta and advances the card to testing
$ bin/fm-teardown.sh phone-status-page-impl teardown phone-status-page-impl complete (window firstmate:fm-phone-status-page-impl, worktree .../wt) dashboard: advanced card rebuild-the-board-for-one-handed-reading-xmmb to testing for phone-status-page-impl teardown exit: 0 $ bin/fm-dashboard.sh show rebuild-the-board-for-one-handed-reading-xmmb status: testingPipeline
Updates from git push no-mistakes
⏭️ **intent** - skipped
✅ No issues found.
⏭️ **Rebase** - skipped
.agents/skills/firstmate-coding-guidelines/SKILL.md- branch carries 9 commit(s) that exist on your local main branch but were never pushed to origin/main; rebasing would bundle this unrelated work (69 file(s)) into the PR:Push main to origin, or rebase your branch onto origin/main, before gating.
bin/fm-backlog-handoff.sh:757- The idempotent "nothing to move" path re-links the card unconditionally, overwriting a newer and more precise link. Sequence:fm-backlog-handoff.sh design item-x --card C1sets C1's ref=design:item-x, agent=design, status=working. The secondmate then runsfm-spawn.sh --card C1in its own home, which sets ref=<sm-home>:<task-id>and agent=<task-id>(bin/fm-spawn.sh:2888-2896). A re-run of the same handoff command - documented as idempotent, and the reason this branch exists - falls into the ALREADY branch and calls handoff_dashboard_link, resetting ref/agent back to the coarse handoff identity while the real task is running. The board then shows a stale agent, which is the exact failure shape docs/dashboard.md says the mechanical link must not produce. Line 629 has the same problem on the remote path whenalreadyis non-empty andto_moveis empty (a resend). Suggested guard: on the already-present path, skip the ref/agent write when the card is no longernot_started, or when its current ref/agent already point somewhere other than this handoff identity.bin/fm-backlog-handoff.sh:344- annotate_dashboard_card only short-circuits when the body already contains the SAME card id; otherwise it appends a newdashboard_card:line rather than replacing any existing one. Sequence:fm-backlog-handoff.sh design item-x --card C1, then laterfm-backlog-handoff.sh design item-x --card C2(card recreated or corrected) - the ALREADY path at line 757 appends a second line, so the item's body now carriesdashboard_card: C1anddashboard_card: C2. On the remote path this is worse: outbox_dashboard_card_pairs (line 368) emits one pair per matching line, so resume_remote_outbox (line 670) links BOTH cards to the same item, setting ref/agent on a card that no longer serves it. The docs and the script header both describe this as the single durable identity link (the equivalent of state/<id>.meta's onedashboard_card=), so the write should drop any pre-existingdashboard_card:line from$existingbefore appending the new one.bin/fm-backlog-handoff.sh:670- The new dashboard calls are made with no bounded timeout while a cross-process lock is held. dash_call in bin/fm-dashboard.sh usescurl -sSwith no--max-time/--connect-timeout, and dashboard_link_card issues up to five such calls (ref, agent, show, status, audit-log). resume_remote_outbox runs under with_remote_route_locks, which holds both ACTIVE_REGISTRY_LOCK and ACTIVE_HANDOFF_LOCK, and fm_lock_acquire_wait (bin/fm-wake-lib.sh:827) spins forever against a live holder. bin/fm-bootstrap.sh:734 calls--resume-pendingsynchronously (|| truedoes not help with a hang). Failure: the dashboard URL points at a tailnet host (the documented deployment) that is powered off, so TCP SYNs are dropped rather than refused; each call blocks for the OS SYN timeout (~130s), so one pending outbox can stall bootstrap for minutes and block every other handoff for that secondmate. The unreachable-dashboard test uses port 1, which gets an immediate ECONNREFUSED, so it does not exercise this. Fix by adding--connect-timeout/--max-timeto dash_call, or by moving the link out of the locked region (it needs delivery confirmation, not the lock). Line 629 is exposed the same way via remote_handoff.bin/fm-backlog-handoff.sh:400- dashboard_link_card is a near-verbatim copy of spawn_dashboard_link (bin/fm-spawn.sh:2881-2920): same ref/agent/show/status call sequence, samefailedaccumulator, same warning wording, same audit-log --fleet fallback. Only the ref format and the log subject differ. bin/fm-teardown.sh:977 holds a third variant of the same shape. The repo already carries the shared-helper convention heavily (bin/fm-*-lib.sh), so abin/fm-dashboard-link-lib.shtaking (card, ref, agent, subject) would give the card-link semantics one owner instead of three, which matters because docs/dashboard.md describes these as one mechanism with three entry points.tests/fm-dashboard-card-link.test.sh:331- All three new tests exercise only the local-move path. The remote paths carry the intricate new logic and have no coverage: the staged-outbox annotation plus delivery-confirmed link (bin/fm-backlog-handoff.sh:622, :629), and the crash-recovery link in resume_remote_outbox (:655-671), which is the sole consumer of outbox_dashboard_card_pairs (:368) - that function is currently never executed by any test. The single-item--cardrefusal (:132) is also unexercised. tests/fm-backlog-handoff.test.sh already has remote-route scaffolding to build on.🔧 Fix: guard handoff card relink, dedupe annotation, bound dashboard calls
2 issues (1 error, 1 warning) still open:
bin/fm-backlog-handoff.sh:336- annotate_dashboard_card silently destroys part of the handed-off item's body. backlog_key_body_lines ends capture atcapturing { exit }(:336) for any line that is neither empty, nor^-indented, nor a##heading - which includes a whitespace-only line that is not exactly empty (one space, or a tab). annotate_dashboard_card (:356-369) then writes that truncated capture as the item's ENTIRE new body throughtasks-axi update --body-file, with no--archive-body, so every body line after the offending line is lost unrecoverably.Reproduced concretely. Body:
backlog_key_body_lines returns only
intent: do the thing; the desired body becomesintent: do the thing\ndashboard_card: <card>andmore:/owner:are written away. Critically, the pre-move refusal guard backlog_key_noncanonical_body_lines (:298-311) does NOT flag that line, because it requires/[^[:space:]]/- so this passes validation and fires on the ordinary local--cardmove path (:866).Two further paths run annotate with no canonical check at all, so there even a tab- or single-space-indented content line truncates: the already-present path (:810, annotating $SUB_BACKLOG, which is never validated) and the remote
already-outbox path (:667, annotating items recovered from a prior staging).The fix-round test test_handoff_card_annotation_replaces_a_stale_card_id asserts
intent: keep this body linesurvives, but its fixture body has no whitespace-only or non-2-space line, so it passes with the bug present.Suggested fix: treat any whitespace-only line as blank in backlog_key_body_lines, and refuse to rewrite (returning failure, which already produces the loud non-fatal
annotatewarning) when the capture stopped before the item's real end - i.e. run the noncanonical check, extended to whitespace-only lines, as a precondition on whichever file is about to be rewritten. Adding--archive-bodywould also make any future clobber recoverable.bin/fm-backlog-handoff.sh:677- remote_handoff links only the current run's--card, but the delivery it performs lands every item staged in the outbox - so a card annotated by an earlier failed run is silently orphaned, which is exactly the freeze-at-not-started bug this change exists to fix.Concrete sequence:
fm-backlog-handoff.sh remote-sm item-A --card C1stages item A into data/handoff/remote-sm.outbox.md and annotates itdashboard_card: C1; the remote is unreachable, so remote_deliver_outbox fails, the outbox is preserved, and C1 is correctly left unlinked (the fix round's own test asserts this). Later,fm-backlog-handoff.sh remote-sm item-B --card C2(or with no --card at all) runs against a reachable remote. to_move=(item-B), so remote_deliver_outbox (:557-574) transfers the WHOLE outbox - A and B - to the secondmate and thenrm -f -- "$outbox". The link block at :677 fires only for${requested[0]}->$CARD_ARG, i.e. B/C2 (or nothing at all when no --card was passed). Item A has now genuinely landed in the secondmate's backlog, but C1 is never advanced from not_started, and because the outbox is deleted,--resume-pendingcan never recover it either - the annotated item text was the only surviving record.The invariant "once an annotated item's delivery is confirmed, its card is linked" is already enforced on the crash-recovery path: resume_remote_outbox (:705-728) reads outbox_dashboard_card_pairs and links every pair. The immediate remote path is the same delivery boundary and should read the same pairs (captured after annotation, before remote_deliver_outbox) and link them all, keeping the existing preserve guard for the pair matching the current --card.
Flagged as ask-user rather than auto-fix because the mechanical fix means a handoff invoked with no --card would touch the dashboard when it flushes a card-bearing outbox, which docs/dashboard.md currently states is a complete no-op - that line needs the author's call.
🔧 Fix: stop card annotation truncating bodies, link every delivered card
3 issues (1 error, 2 infos) still open:
bin/fm-backlog-handoff.sh:371- The round-3 round-trip guard compares the ENTIRE file byte-for-byte, so it refuses the annotation on the ordinary local handoff path whenever the secondmate's backlog already holds a queued item - the durabledashboard_card:line the whole mechanism rests on is then never written.Root cause:
tasks-axi mvappends the moved item directly against the following##heading when the destination section is non-empty (verified: destination## Queuedwith one existing item yields...\n intent: fresh\n## Done, no blank separator; an EMPTY destination yields a blank line and round-trips clean, which is why every test fixture passes).annotate_body_roundtripsthen writes the unchanged body onto a probe copy, tasks-axi re-inserts the missing blank separator, the whole-file diff at :371 differs, andannotate_dashboard_card(:400) refuses - even though that rewrite is provably lossless.Reproduced end-to-end with the real script (main home + secondmate home, sub backlog already containing one queued item, unreachable board):
and
grep -c dashboard_card sub/data/backlog.md-> 0.Blast radius: (a) every second-and-later
--cardhandoff to the same secondmate emits two alarming warnings on an operation that actually succeeded; (b) the item's single durable card identity - which docs/dashboard.md:67 states "survives the move and any later delivery retry" - is not recorded; (c) on the remote path the same shape hits any item staged into an outbox that already holds one, so a crash before delivery leaves that card unrecoverable by--resume-pending, i.e. the freeze-at-not-started bug this change exists to fix stays reachable for the co-staged case. Only the fallback pair appended at :758 keeps the immediate remote link working.The existing tests cannot catch this: every fixture hands off into an empty destination
## Queued, the one shape that round-trips identically.Fix: compare content, not bytes - e.g. diff the trailing-whitespace-stripped, blank-line-stripped line sequences of the two files, or scope the comparison to the target item's own region. Content loss (the case the guard exists for) still shows up as a missing non-blank line, and the tab-indented case stays safe because tasks-axi preserves lines it does not read as body (verified separately).
tests/fm-dashboard-card-link.test.sh:573-test_handoff_already_present_annotation_keeps_a_noncanonical_bodycannot fail on the pre-fix code, so it is not a regression test for anything.tasks-axistops reading body at a tab-indented line exactly as the reader does, andupdate --body-fileleaves the lines it did not read in place (verified: writing a one-line body to an item followed by\tmore: ...\n owner: ...preserves both). The old reader would therefore also have kept them.Worse, on the current code this fixture silently exercises the refusal path from the finding above: the annotation is rejected and no
dashboard_card:line is written, and the test never notices because it only asserts the three pre-existing body lines are still present. (The sibling whitespace-only test at :536 is a genuine regression test - it does fail on the old reader.)Once the guard is content-based, this test should also assert that the durable
dashboard_card: $cardline was actually written alongside the surviving non-canonical line; otherwise it asserts an outcome that a broken annotation satisfies.bin/fm-backlog-handoff.sh:807- Bounded, but still serialized: with an unreachable board, each pair inlink_delivered_card_pairscosts onecard_already_linkedshow plus two failing writes plus one audit-log, each up to the 5s connect timeout - roughly 20s per pair - while both ACTIVE_REGISTRY_LOCK and ACTIVE_HANDOFF_LOCK are held and bin/fm-bootstrap.sh:734 drives--resume-pendingsynchronously. An outbox carrying five annotated items stalls bootstrap and every other handoff for that secondmate for ~100s where it previously cost nothing.No action required for realistic one-or-two-item outboxes; noting it because the round-2 fix framed the timeout as removing the stall rather than capping it. If it ever matters, one reachability probe that short-circuits the remaining pairs once the board is proven unreachable bounds the whole loop at a single timeout.
🔧 Fix: hold handoff card pairing in state, never the item body
3 issues (1 warning, 2 infos) still open:
bin/fm-backlog-handoff.sh:478- consume_handoff_card_record deletes the record unconditionally, including when every link attempt just failed - destroying the one durable statement at exactly the moment it was needed. link_delivered_card_pairs (:437) and dashboard_link_card (:377) always return 0 (by design: never fail the handoff), so :478 runs whether the board answered or not.Concrete silent-orphan sequence, entirely through supported paths:
fm-backlog-handoff.sh remote-sm item-A --card C1stages item-A into data/handoff/remote-sm.outbox.md and writes state/handoff-cards/remote-sm; the remote is unreachable, delivery fails, outbox + record are correctly preserved and C1 stays not_started (the branch's own test asserts this).fm-backlog-handoff.sh --resume-pending >/dev/null 2>&1 || true. The remote is back, so remote_deliver_outbox succeeds and deletes the outbox.|| true.End state: item-A delivered, C1 frozen at not_started forever, nothing on stderr (discarded), nothing in the fleet audit log, and no recovery path - the outbox is gone, so --resume-pending can never see it again, and re-running the handoff fails with "no backlog or pending outbox item matched: item-A" because item-A is no longer in the main backlog or the outbox. That is the freeze-at-not-started failure this change exists to eliminate, now reachable with no trace.
The local paths do not have this problem: after a failed link the item is 'already present', so re-running the same command re-puts the record and retries (guarded, so it cannot clobber). Only the remote/resume paths are one-shot, because delivery consumes the outbox.
The required invariant is 'a recorded pairing is discarded only once the board has actually accepted it'. Earliest supported boundary: have link_delivered_card_pairs report which pairs failed, and have consume_handoff_card_record rewrite the record with the survivors instead of calling handoff_card_record_clear unconditionally. Every surviving pair is already guarded by card_already_linked, so a later retry cannot overwrite a link the secondmate has since made more precise, and the next handoff or --resume-pending for that secondmate completes it. Flagged ask-user rather than auto-fix because the current behaviour is a stated design choice (the comment at :431-436 and docs/dashboard.md:67 both assert this is 'the last moment they can be linked at all'), so retaining failed pairs also means those two claims need the author's revision.
tests/fm-dashboard-card-link.test.sh:9- The file header still states the handoff pairing is "recorded as adashboard_card:body line on the item itself". That store was deleted in this same branch and replaced by state/handoff-cards/<secondmate-id>; nodashboard_card:line is ever written to a backlog item now. The tests immediately below it assert the opposite (test_handoff_card_leaves_the_item_body_byte_identical pins that the item is not rewritten at all), so the header actively misdescribes what the suite proves and would send the next reader looking for a mechanism that no longer exists. The in-suite section comment at :347-352 is already correct; the header just was not updated with it.bin/fm-test-run.sh:904- The new handoff card-link behaviour lives entirely in tests/fm-dashboard-card-link.test.sh, which family_for_basename does not classify (it falls through tounclassifiedat :226). families_for_changed_path maps bin/fm-backlog-handoff.sh explicitly to thesecondmatefamily only (:903-907), and that explicit case shadows thebin/*fallback at :1006 that would otherwise have found the suite by reference scan. So a future edit to bin/fm-backlog-handoff.sh selects thesecondmatefamily and never runs any of the seven handoff card tests added here - the coupling this change introduces is invisible to the selector.This branch itself is safe (the test file is in the diff, so the
__script__:marker at :853 selects it directly); the gap only bites the next change. Either add the suite's basename to a family that bin/fm-backlog-handoff.sh already emits, or add that family to the :903 case arm.🔧 Fix: map dashboard suites into changed-file test selection
1 info still open:
bin/fm-backlog-handoff.sh:478- consume_handoff_card_record ends on handoff_card_record_clear's barerm -f -- "$record", with no explicitreturn 0, so a bookkeeping cleanup failure becomes the handoff's own exit status - the precise outcome the function's neighbouring comments ("this function must never turn that success into a failure") and dashboard_link_card's explicitreturn 0are written to prevent.Reachable when
$STATE/handoff-cardsis not writable by the caller (created under another uid, or the record path replaced by a non-empty directory).rm -fthen exits 1, and it is the last command of handoff_card_record_clear, which is the last command of consume_handoff_card_record, which is the last command in every one of its four call sites:rc=1and the script exits non-zero for a handoff that fully landed - and the script's own vocabulary for that exit tells the operator to preserve the outbox and rerun, which no longer exists.--resume-pendingexits 1 after a successful delivery. bin/fm-bootstrap.sh drives that path.set -ethis aborts before the followingexit 0.tasks-axi mvreported as a failed handoff.Fix: end consume_handoff_card_record (or handoff_card_record_clear) with an explicit
return 0.🔧 **Test** - 2 issues found → auto-fixed ✅
bin/fm-test-run.sh:659-bin/fm-test-run.sh --check-coverageaborts under a UTF-8 locale, sotests/fm-test-run.test.shfails on a developer machine withLANG=en_US.UTF-8.run_coverage_guardbuilds its lane files withLC_ALL=C sortbut then callscommin the ambient locale;comm -12 shards_union serialreportsfile 2 is not in sorted orderand exits non-zero, andset -ekills the guard with no diagnostic. Not introduced by this branch — I reproduced the identical abort with the base commit's copy of the runner over the base-era test inventory, and everything passes underLC_ALL=C. Fixing it means addingLC_ALL=Cto thecommcalls inrun_coverage_guard, which is outside this change's scope, so it is the author's call whether to fold it in here.tests/fm-test-run.test.sh:634-test_herdr_ci_family_run_has_a_step_timeout, added by this branch, hard-failed (not ok - ruby is required to parse .github/workflows/ci.yml as YAML) because ruby is not installed here and no other test in the suite depends on it. Fixed: the YAML is now parsed with python3 + PyYAML (python3 was already required two lines later for the JSON read), with the original ruby parse kept as a fallback and an honest skip only if neither parser exists. The assertion itself is unchanged and still verified against a mutatedci.yml, and the suite now passes.bash tests/fm-dashboard-card-link.test.sh— 21/21 pass, including the handoff--cardlink, the byte-identical-item guarantee, the no-overwrite re-run, and the remote outbox/--resume-pendingcard recoverybash tests/fm-dashboard.test.sh— the board's own server/CLI suitebash tests/fm-backlog-handoff.test.sh— handoff regression suitebash tests/fm-remote-backlog-handoff.test.sh— remote delivery regression suiteLC_ALL=C bash tests/fm-test-run.test.sh— lane/family selection incl. the newdashboardfamily mapping; failed on the new ruby-dependent test, passes after the fixLC_ALL=C bash bin/fm-test-run.sh --check-coverage→FM_TEST_COVERAGE ok total=150 parallel=24 serial=114 serial_shards=4 herdr=12(all 4 new dashboard suites are lane-covered)Manual end-to-end: started a real dashboard server (bin/fm-dashboard.sh starton a free port), created 3 cards, ranbin/fm-backlog-handoff.sh atlas login-redirect-loop --card <id>against a seeded secondmate home, and capturedbin/fm-dashboard.sh show <id> --jsonbefore/afterManual browser capture ofhttp://127.0.0.1:<port>/board view and card-detail view before and after the handoff (chrome-devtools-axi open/screenshot)Manual idempotency check: replaced the card link with a preciseatlas-home:task-4471identity, then re-ran the same handoff and confirmed the board was left unchangedMutation check on the new CI-timeout test: parsing a copy ofci.ymlwithtimeout-minutes: 20changed to40yieldsstep_timeout: 40, so the assertion still detects a broken workflow contractConfirmed the--check-coveragelocale failure is pre-existing by running the base commit'sbin/fm-test-run.sh(fromgit show 6789876:bin/fm-test-run.sh) restricted to the base-era test inventory — samecomm: file 2 is not in sorted orderabort🔧 Fix: run coverage guard's comm under LC_ALL=C
✅ Re-checked - no issues remain.
bin/fm-test-run.sh tests/fm-dashboard-card-link.test.sh(all 21 cases, real dashboard server)bin/fm-test-run.sh tests/fm-backlog-handoff.test.sh tests/fm-dashboard.test.sh tests/fm-teardown.test.shbin/fm-test-run.sh tests/fm-remote-backlog-handoff.test.sh tests/fm-test-run.test.shManual E2E: started a real server viabin/fm-dashboard.sh start, created a card, ranbin/fm-backlog-handoff.sh design phone-status-page --card <card>, verified ref/agent/status viabin/fm-dashboard.sh show, and confirmed the item body landed unchanged in the secondmate backlogManual E2E: re-ran the same handoff (idempotent, link left unchanged) and ranbin/fm-backlog-handoff.sh design unrelated-thing --card no-such-cardto confirm the loud stderr warning, the non-blocking exit 0, and the fleet finding inbin/fm-dashboard.sh audit-status --jsonManual E2E:bin/fm-teardown.sh <id>on a landed local-only task whosestate/<id>.metacarrieddashboard_card=, confirming the card advanced totesting(run with the repo's documentedFM_GATE_REFUSE_BYPASS=1harness hatch against a /tmp sandbox fleet)Browser capture of the rendered board and card detail viachrome-devtools-axi open/screenshotbin/fm-test-run.sh --list --changed --base 3c8f796to confirm the new changed-file mapping selects the dashboard suitesdocs/configuration.md:517- Two hand-maintained inventories in this repo drift silently and both were touched by this change's surface area, but consolidating them is out of scope here. (1) docs/configuration.md's "Environment variables" block documents 107 of the 282 FM_/FMX_ variables actually read by bin/.sh; I left the six new FM_DASHBOARD_ variables out of it on purpose, since docs/dashboard.md is their operator-facing owner and is already in readmeSetupTargets, so copying them here would be synchronization rather than ownership. (2) AGENTS.md's state/ tree omits pre-existing directories such as state/terminal-outcomes/ (bin/fm-inactive-reconcile.sh:48) and state/remote-replies/ (bin/fm-procevent-remote-reply.sh:68); I added only state/handoff-cards/, the one this change introduced. A follow-up could decide whether either inventory should be generated from its authoritative source and drift-checked (the way bin/fm-doc-audience-check.sh already guards the audience inventory) instead of hand-copied.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.