Skip to content

feat(dashboard): add the Admiral's Fleet Dashboard and split testing from review - #2712

Closed
joliverMI wants to merge 23 commits into
kunchenguid:mainfrom
joliverMI:fm/fm-dashboard-review-status
Closed

joliverMI wants to merge 23 commits into
kunchenguid:mainfrom
joliverMI:fm/fm-dashboard-review-status

Conversation

@joliverMI

Copy link
Copy Markdown

Intent

Split the Admiral's Fleet Dashboard's testing status into two: review (done, ready for the Admiral's review — what testing meant until now, optional to him, nothing lost if he never looks, he marks it complete from here, and must NOT be age-checked by the fleet auditor) and testing (the fleet is actively testing it right now — work in flight, not waiting on him; a testing card with no live crew activity behind it is a genuine discrepancy the auditor must flag, the same way a stale working card is). Scope: bin/fm-dashboard.sh and its HTTP API accept and store both statuses, keeping CLI spelling consistent with existing statuses (hyphens on the CLI, underscores in the store). Migrate every existing testing card to review exactly once, idempotently, proving what changed. bin/fleet-dashboard/web/ lists both statuses with sensible labels/ordering and working filters, with review reading as the one awaiting him. .claude/skills/fleet-dashboard/SKILL.md's seven-status section documents both, keeps the not-a-finding age rule attached to review, and adds that a testing card with no live crew activity is a finding. The fleet auditor's --status testing sweep instructions in that skill now verify live crew activity instead of treating testing as inert. docs/dashboard.md is updated everywhere statuses are enumerated. Constraints: do not invent additional statuses or rename any of the other five; do not change what needs-attention means; the board is live (used from the Admiral's phone) so the running server must not break and no card may end up with a status the UI cannot render.

What Changed

  • Adds the Admiral's Fleet Dashboard: a bin/fm-dashboard.sh lifecycle/CLI wrapper over a stdlib-Python server (bin/fleet-dashboard/server/, SQLite-backed) serving a vendored-React web board (bin/fleet-dashboard/web/), alongside the render-from-state bin/fm-status-board.sh page and the fleet auditor's sweep/tick scripts (bin/fm-fleet-audit-sweep.sh, bin/fm-fleet-audit-tick.sh).
  • Splits the board's testing status into review (done, awaiting the Admiral, exempt from the auditor's age check) and testing (fleet actively testing, where a card with no live crew activity is a finding), with CLI hyphen/store underscore spelling, a one-time idempotent startup migration of existing testing cards to review that reports what it moved, cache-busted asset reload so a restarted board never renders an unknown status, and updated filters/labels in the web UI, docs/dashboard.md, and the fleet-dashboard skill.
  • Wires cards to real work mechanically through fm-spawn.sh --card, fm-teardown.sh, and fm-backlog-handoff.sh --card (with a pending-link record and --resume-pending recovery), and carries adjacent fixes: teardown refuses while a recorded PR is still open, procevent delivers captured results at most once, tmux endpoint targets resolve exactly for liveness probes and sends, watch-arm re-evaluates coalescing when the fleet lock changes, and a trust-registered custom check now counts as supervision-needed — each covered by new or extended shell tests under tests/.

Risk Assessment

✅ Low: The change is well-bounded and satisfies every source-verifiable intent criterion, and its one genuinely irreversible element — the live board's testing→review migration — is gated behind a successful socket bind, a settings marker, and an idempotent INSERT OR IGNORE, with a real cross-restart regression test proving it runs at most once and never rewrites a post-split testing card.

Testing

Ran the three targeted suites that own this surface (fm-dashboard, fm-fleet-audit, fm-dashboard-card-link) — all pass — then demonstrated the intent end-to-end against a real dashboard server on a seeded pre-split database: the one-time migration rewrote exactly the 3 pre-existing testing cards to review and reported them by id, left an audit trail and updated_at intact, did not re-run across a live restart, and left a genuinely in-flight testing card alone; the CLI and HTTP API accept both statuses and refuse invented ones; the auditor flags a testing card with no live crew and never age-checks review; and phone-width browser screenshots show both chips in order with distinct colours, working filters, Mark Complete only on review, and a real click moving a migrated card to complete. No findings.

Evidence: Evidence index (what each file proves)

Source: Evidence index (what each file proves)

# Testing/review split — end-to-end evidence

A real `bin/fm-dashboard.sh` server was started against a seeded **pre-split**
SQLite database (3 `testing` cards plus one card in each other status), then
driven exactly the way an agent, the fleet auditor, and the Admiral's phone
drive it: the CLI, the HTTP API, the fleet-auditor sweep, and the rendered
board in Chrome.

| File | What it shows |
|---|---|
| `01-migration-startup.txt` | `fm-dashboard.sh start` against a pre-split DB; the server's startup report names the 3 cards it rewrote. |
| `02-migrated-state-and-audit-trail.txt` | Persisted state after migration: 3 cards now `review`, each with a single `testing -> review` `status_history` entry carrying the reason, and `updated_at` untouched. |
| `03-cli-and-api-accept-both.txt` | CLI and HTTP API accept and store both `testing` and `review`; an invented status is refused by both (CLI exit 1, API HTTP 400). |
| `04-restart-idempotence.txt` | A live `restart` does **not** re-run the migration; a genuinely in-flight `testing` card stays `testing`. No card carries a status the web bundle cannot render. |
| `05-auditor-sweep.txt` | The auditor flags a `testing` card whose linked crew reads `blocked`, flags none of the 4 `review` cards, and produces no finding once a live crew is behind the same testing card. |
| `06-board-review-and-testing.png` | The rendered board at phone width: `Testing (1)` and `Review (4)` chips in order, distinct colours, `Mark Complete` only on `review` cards. |
| `07-filter-review.png` | The `Review` filter chip active — the 4 awaiting-him cards, each with `Mark Complete`, still showing their original ages (4d/5d/6d) because the migration left `updated_at` alone. |
| `08-filter-testing.png` | The `Testing` filter chip active — the one in-flight card, with the auditor's discrepancy for it visible in the Discrepancy Log below. |
| `09-card-overlay-status-dropdown.png` | A `testing` card's overlay: the status dropdown offers all eight statuses including `Review`, and a `testing` card gets no `Mark Complete`. |
| `10-mark-complete-from-review.png` | After clicking `Mark Complete` on the migrated "Re-rig the mainsail halyard" card in the browser: `Review (4) -> (3)`, `Complete (1) -> (2)`, and the API confirms `review -> complete` persisted. |
- Evidence: [Rendered board: Testing (1) and Review (4) chips, Mark Complete only on review cards](https://github.com/joliverMI/firstmate/blob/f8fc46b26f625671f245b072b69c80ae36963720/.no-mistakes/evidence/fm/fm-dashboard-review-status/06-board-review-and-testing.png) - Evidence: [Review filter active: the 4 cards awaiting the Admiral, original ages preserved](https://github.com/joliverMI/firstmate/blob/f8fc46b26f625671f245b072b69c80ae36963720/.no-mistakes/evidence/fm/fm-dashboard-review-status/07-filter-review.png) - Evidence: [Testing filter active: the one in-flight card, with the auditor's discrepancy for it below](https://github.com/joliverMI/firstmate/blob/f8fc46b26f625671f245b072b69c80ae36963720/.no-mistakes/evidence/fm/fm-dashboard-review-status/08-filter-testing.png) - Evidence: [Card overlay: status dropdown offers Review; a testing card gets no Mark Complete](https://github.com/joliverMI/firstmate/blob/f8fc46b26f625671f245b072b69c80ae36963720/.no-mistakes/evidence/fm/fm-dashboard-review-status/09-card-overlay-status-dropdown.png) - Evidence: [After clicking Mark Complete on a migrated review card: Review 4 -> 3, Complete 1 -> 2](https://github.com/joliverMI/firstmate/blob/f8fc46b26f625671f245b072b69c80ae36963720/.no-mistakes/evidence/fm/fm-dashboard-review-status/10-mark-complete-from-review.png)
Evidence: Startup migration report against a pre-split database

Source: Startup migration report against a pre-split database

$ fm-dashboard.sh --help | grep statuses statuses: needs-attention not-started working paused waiting testing review complete $ fm-dashboard.sh start fleet dashboard started (pid 2965092) - http://127.0.0.1:8477/

$ cat $FM_HOME/state/dashboard.log dashboard: migrated 3 card(s) from testing to review (testing/review split): card-rigging, card-charts, card-galley fleet dashboard listening on http://127.0.0.1:8477

$ fm-dashboard.sh --help | grep statuses
statuses: needs-attention not-started working paused waiting testing review complete

$ fm-dashboard.sh start
fleet dashboard started (pid 2965092) - http://127.0.0.1:8477/  log: /tmp/fm-dash-demo-2955262/state/dashboard.log

$ cat $FM_HOME/state/dashboard.log   # the server's own startup migration report
dashboard: migrated 3 card(s) from testing to review (testing/review split): card-rigging, card-charts, card-galley
fleet dashboard listening on http://127.0.0.1:8477  (db: /tmp/fm-dash-demo-2955262/data/dashboard.db)
Evidence: Migrated state + per-card audit trail, updated_at untouched

Source: Migrated state + per-card audit trail, updated_at untouched

$ curl -s $B/api/tasks/card-rigging | jq -c '.status_history[]' {"id":1,"task_id":"card-rigging","from_status":"testing","to_status":"review","changed_at":"2026-08-21T01:29:33Z","note":"migrated: testing/review split - this card meant ready for his review"} # updated_at deliberately untouched, so a mechanical relabel does not reorder his board: card-galley review 2026-08-16T08:15:00Z card-charts review 2026-08-15T11:30:00Z card-rigging review 2026-08-14T09:00:00Z

$ curl -s $BOARD/api/tasks | jq -r '.tasks[] | "\(.status)  \(.id)"' | sort
complete	card-lantern	Replace the stern lantern housing
needs_attention	card-mutiny	Crew rotation conflict needs a call
not_started	card-compass	Compass deviation card re-swing
paused	card-signal	Signal-flag protocol refresh
review	card-charts	Correct the coastal chart overlays
review	card-galley	Galley inventory reconciliation
review	card-rigging	Re-rig the mainsail halyard
waiting	card-anchor	Anchor winch bearing replacement
working	card-hull	Hull scan for the dry-dock window

# every migrated card carries its own audit trail (status_history over the HTTP API):
$ curl -s $BOARD/api/tasks/card-rigging | jq -c '.status_history[]'
{"id":1,"task_id":"card-rigging","from_status":"testing","to_status":"review","changed_at":"2026-08-21T01:29:33Z","note":"migrated: testing/review split - this card meant ready for his review"}
$ curl -s $BOARD/api/tasks/card-charts | jq -c '.status_history[]'
{"id":2,"task_id":"card-charts","from_status":"testing","to_status":"review","changed_at":"2026-08-21T01:29:33Z","note":"migrated: testing/review split - this card meant ready for his review"}
$ curl -s $BOARD/api/tasks/card-galley | jq -c '.status_history[]'
{"id":3,"task_id":"card-galley","from_status":"testing","to_status":"review","changed_at":"2026-08-21T01:29:33Z","note":"migrated: testing/review split - this card meant ready for his review"}

# updated_at deliberately untouched, so a mechanical relabel does not reorder his board:
card-galley	review	2026-08-16T08:15:00Z
card-charts	review	2026-08-15T11:30:00Z
card-rigging	review	2026-08-14T09:00:00Z
Evidence: CLI and HTTP API accept both statuses, refuse invented ones

Source: CLI and HTTP API accept both statuses, refuse invented ones

$ fm-dashboard.sh status card-compass bogus-status fm-dashboard.sh: unknown status 'bogus-status' - valid: needs-attention not-started working paused waiting testing review complete exit=1 $ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"testing"}' {"id":"card-compass","status":"testing"} $ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"review"}' {"id":"card-compass","status":"review"} $ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"in-review"}' {"error": "unknown status: 'in-review'. Valid: needs_attention, not_started, working, paused, waiting, testing, review, complete"} <- HTTP 400

# --- CLI: hyphens on the CLI, underscores in the store --------------------
$ fm-dashboard.sh --help | grep '^statuses'
statuses: needs-attention not-started working paused waiting testing review complete

$ fm-dashboard.sh list --status testing
ballast-trim-sea-trial-yrag	testing	captain_dj	-	Ballast trim sea trial
$ fm-dashboard.sh list --status review
bilge-pump-swap-md8i	review	firstmate	-	Bilge pump swap
card-galley	review	captain_river	-	Galley inventory reconciliation
card-charts	review	captain_dj	-	Correct the coastal chart overlays
card-rigging	review	firstmate	-	Re-rig the mainsail halyard

$ fm-dashboard.sh status card-compass bogus-status    # no statuses invented
fm-dashboard.sh: unknown status 'bogus-status' - valid: needs-attention not-started working paused waiting testing review complete
exit=1

# --- HTTP API: this is what the board on his phone calls -------------------
$ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"testing"}' | jq -c '{id,status}'
{"id":"card-compass","status":"testing"}
$ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"review"}' | jq -c '{id,status}'
{"id":"card-compass","status":"review"}
$ curl -sX POST $B/api/tasks/card-compass/status -d '{"status":"in-review"}'   # must be refused
{"error": "unknown status: 'in-review'. Valid: needs_attention, not_started, working, paused, waiting, testing, review, complete"} <- HTTP 400

$ curl -s '$B/api/tasks?status=review' | jq -r '.tasks[].id'
card-compass
bilge-pump-swap-md8i
card-galley
card-charts
card-rigging
$ curl -s '$B/api/tasks?status=testing' | jq -r '.tasks[].id'
ballast-trim-sea-trial-yrag

# the status change is persisted with its own history entry, same as any other:
{"id":"card-compass","status":"review","status_history":[{"id":6,"task_id":"card-compass","from_status":"not_started","to_status":"not_started","changed_at":"2026-08-21T01:29:57Z","note":null},{"id":7,"task_id":"card-compass","from_status":"not_started","to_status":"testing","changed_at":"2026-08-21T01:30:16Z","note":null},{"id":8,"task_id":"card-compass","from_status":"testing","to_status":"review","changed_at":"2026-08-21T01:30:16Z","note":null}]}
Evidence: Live restart does not re-run the migration; no unrenderable status

Source: Live restart does not re-run the migration; no unrenderable status

$ fm-dashboard.sh restart stopped (pid 2965092) fleet dashboard started (pid 2988796) - http://127.0.0.1:8477/&#10;&#10;$ cat $FM_HOME/state/dashboard.log # no second migration report fleet dashboard listening on http://127.0.0.1:8477&#10;&#10;$ fm-dashboard.sh list --status testing # after restart - still testing ballast-trim-sea-trial-yrag testing captain_dj - Ballast trim sea trial cards on the live board with an unrenderable status: none

# The live board is upgraded with 'restart'. The migration must not fire again,
# or the genuinely in-flight testing card would be silently relabelled review.

$ fm-dashboard.sh list --status testing        # before restart
ballast-trim-sea-trial-yrag	testing	captain_dj	-	Ballast trim sea trial

$ fm-dashboard.sh restart
stopped (pid 2965092)
fleet dashboard started (pid 2988796) - http://127.0.0.1:8477/  log: /tmp/fm-dash-demo-2955262/state/dashboard.log

$ cat $FM_HOME/state/dashboard.log             # no second migration report
fleet dashboard listening on http://127.0.0.1:8477  (db: /tmp/fm-dash-demo-2955262/data/dashboard.db)

$ fm-dashboard.sh list --status testing        # after restart - still testing
ballast-trim-sea-trial-yrag	testing	captain_dj	-	Ballast trim sea trial

$ curl -s $B/api/tasks | jq -r '.tasks[]|[.status,.id]|@tsv' | sort | uniq -c -w16
      1 complete
      1 needs_attention
      1 not_started
      1 paused
      4 review
      1 testing
      1 waiting
      1 working

# and no card is left carrying a status the UI cannot render:
statuses the web bundle can render: ['complete', 'needs_attention', 'not_started', 'paused', 'review', 'testing', 'waiting', 'working']
cards on the live board with an unrenderable status: none
Evidence: Auditor: flags a testing card with no live crew, never age-checks review

Source: Auditor: flags a testing card with no live crew, never age-checks review

$ bin/fm-crew-state.sh sea-trial-crew state: blocked · source: status-log · the trial berth is unavailable $ FM_AUDIT_STALE_NEEDS_ATTENTION_MINUTES=0 bin/fm-fleet-audit-sweep.sh --forced {"tasks_checked": 5, "discrepancies_found": 1, "forced": 1} discrepancies: [{"task_id": "ballast-trim-sea-trial-yrag", "text": "card claims testing, but the linked crew reads: state: blocked · source: status-log · the trial berth is unavailable"}] # none of the 4 review cards flagged # with a live crew behind the same testing card: state: working · source: status-log · running the ballast trial {"tasks_checked": 5, "discrepancies_found": 0}

# The fleet auditor now treats 'testing' as a live-crew claim it must corroborate,
# while 'review' is optional to the Admiral and is never age-checked.

# board going in:
ballast-trim-sea-trial-yrag	testing	captain_dj	-	Ballast trim sea trial
bilge-pump-swap-md8i	review	firstmate	-	Bilge pump swap
card-galley	review	captain_river	-	Galley inventory reconciliation
card-charts	review	captain_dj	-	Correct the coastal chart overlays
card-rigging	review	firstmate	-	Re-rig the mainsail halyard

$ bin/fm-crew-state.sh sea-trial-crew   # the crew behind the testing card
state: blocked · source: status-log · the trial berth is unavailable

$ FM_AUDIT_STALE_NEEDS_ATTENTION_MINUTES=0 bin/fm-fleet-audit-sweep.sh --forced
exit=0

$ curl -s $B/api/audit/status | jq '{last_run, discrepancies: [.log[]|select(.kind=="discrepancy")|{task_id,text}]}'
{
  "last_run": {
    "id": 1,
    "started_at": "2026-08-21T01:30:53Z",
    "completed_at": "2026-08-21T01:30:53Z",
    "duration_seconds": 0.0,
    "tasks_checked": 5,
    "discrepancies_found": 1,
    "forced": 1
  },
  "discrepancies": [
    {
      "task_id": "ballast-trim-sea-trial-yrag",
      "text": "card claims testing, but the linked crew reads: state: blocked · source: status-log · the trial berth is unavailable"
    }
  ]
}

# Not one of the four review cards was flagged - review is his to look at whenever,
# nothing is lost if he never does. The testing card WAS flagged, because the fleet
# is not actually exercising it.

# --- and the converse: put a live crew behind that same testing card -------
$ bin/fm-crew-state.sh sea-trial-crew
state: working · source: status-log · running the ballast trial
$ bin/fm-fleet-audit-sweep.sh --forced
exit=0
$ curl -s $B/api/audit/status | jq '.last_run | {tasks_checked, discrepancies_found}'
{
  "tasks_checked": 5,
  "discrepancies_found": 0
}
# genuinely-in-flight testing is corroborated and produces no finding.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Push main to origin, or rebase your branch onto origin/main, before gating.

🔧 **Review** - 2 issues found → auto-fixed (3) ✅
  • ⚠️ bin/fleet-dashboard/server/store.py:163 - The one-time, irreversible testing->review migration is committed inside Store.__init__, and serve() constructs the Store (bin/fleet-dashboard/server/api.py:417) before it binds the listening socket (api.py:419). The live board's upgrade path is bin/fm-dashboard.sh restart (fm-dashboard.sh:550), which runs cmd_server_stop -> kill &#34;$pid&#34; and returns immediately without waiting for the old process to exit (fm-dashboard.sh:500-508) - the exact race the author had to fix in the test harness by adding the polling stop_dashboard_server helper (tests/fm-dashboard.test.sh:24-38). Failing sequence: restart signals the old pre-split server; the new process starts, runs the migration and commits it (marker set, every testing card now review); ThreadingHTTPServer((host, port)) then raises EADDRINUSE because the old process is still listening; the new process dies; cmd_server_start reports "failed to start". The database is now permanently migrated while the still-running pre-split server serves review cards whose status its own API rejects (400 unknown status) and whose pills its bundled app.js cannot render - the state the intent's "the running server must not break and no card may end up with a status the UI cannot render" constraint is about. A retry of start recovers, but the window is unattended. Earliest shared boundary that makes the invariant hold: have cmd_server_stop wait for the pid to actually exit before returning (the same poll the tests already implement), so restart cannot overlap two servers; binding the socket before constructing the Store would additionally ensure the migration only commits once the new server is definitely the one serving.
  • ℹ️ bin/fleet-dashboard/server/store.py:211 - The migration marker is written with a bare INSERT INTO settings(key, value), unlike the INSERT OR IGNORE used for audit_interval_minutes 48 lines above, and settings.key is a PRIMARY KEY (store.py:84). The check-then-insert is not atomic across processes: if two Store constructions race against the same pre-split database (e.g. two overlapping fm-dashboard.sh start invocations during the very first upgrade), both read the marker as absent and the loser raises sqlite3.IntegrityError: UNIQUE constraint failed: settings.key, which propagates out of Store.__init__ through serve() and aborts server startup with a bare traceback in state/dashboard.log. Narrow window (only on the one start that actually migrates), but INSERT OR IGNORE makes it free to close and matches the adjacent line.

🔧 Fix: bind dashboard socket before migrating, make stop synchronous
1 warning still open:

  • ⚠️ bin/fm-dashboard.sh:557 - restart cannot bring the board back up whenever cmd_server_stop takes its die path. die is exit 1 (fm-dashboard.sh:92), so restart) cmd_server_stop 2&gt;/dev/null; cmd_server_start ;; terminates the script before cmd_server_start ever runs, and 2&gt;/dev/null swallows the reason. Concrete sequence: the dashboard process crashes (or is killed outside this script), leaving state/dashboard.pid with a dead pid; the operator runs bin/fm-dashboard.sh restart to bring the live board back; cmd_server_stop hits kill -0 &#34;$pid&#34; || { rm -f &#34;$pf&#34;; die &#34;recorded pid $pid is not running&#34;; } (line 504) and exits 1 with no output at all - the board stays down, and only a separately-typed start recovers it. The same silent abort now also applies to the two die paths that exist in stop: no pidfile (line 502) and the new 15-second pid did not exit timeout this run added (line 510). Fix is mechanical: run the stop in a subshell so its exit cannot kill the parent and a failed stop is not fatal to the restart, e.g. restart) ( cmd_server_stop ) &gt;/dev/null 2&gt;&amp;1 || true; cmd_server_start ;;. If the process genuinely refuses to die, cmd_server_start then fails loudly on EADDRINUSE with failed to start - see .../dashboard.log, which is an honest report rather than silence.

🔧 Fix: keep restart starting the board when stop refuses
4 issues (1 warning, 3 infos) still open:

  • ⚠️ bin/fm-dashboard.sh:509 - The new wait loop escalates to kill -9 on whatever pid state/dashboard.pid names, without ever confirming that pid is the dashboard server. Line 504 only proves the pid is alive, not that it is ours. The stale-pidfile-after-crash state is a first-class supported input here - this run's own test_restart_recovers_from_a_crashed_or_stopped_board constructs it deliberately by SIGKILLing the server so it cannot remove its own pidfile. Concrete sequence: the dashboard is OOM-killed or dies with its tmux session, leaving a pidfile with a dead pid; the host's pid counter later wraps and that pid is reused by an unrelated long-lived process (a build, another agent's crew, a daemon); the operator runs bin/fm-dashboard.sh restart to bring the board back. cmd_server_stop sees kill -0 succeed, SIGTERMs the innocent process, and - this is the part this run added - now blocks 10 seconds and SIGKILLs it, so a process that deliberately handles or ignores SIGTERM (which previously survived) is destroyed. Under restart (line 557) the whole thing is on discarded stderr, so nothing is reported. Suggested fix: before signalling, confirm the pid's command line actually names the dashboard entrypoint (e.g. ps -p &#34;$pid&#34; -o command= matching fleet-dashboard/server/main.py) and treat a mismatch like the dead-pid case - drop the stale pidfile and refuse, rather than killing. Related, same function: line 514 prints stopped (pid %s) identically whether the server exited on SIGTERM or had to be force-killed at line 509, so a server that had to be SIGKILLed reads exactly like a clean shutdown - worth saying which happened.
  • ℹ️ bin/fm-dashboard.sh:557 - restart) ( cmd_server_stop ) 2&gt;/dev/null || true; cmd_server_start ;; discards stop's stderr, which now also swallows the 15-second timeout die this run introduced at line 510. Sequence: the old server is wedged and survives both SIGTERM and SIGKILL; stop dies at line 510 leaving the pidfile in place (deliberately); cmd_server_start then sees a pidfile whose pid is still live and dies with already running (pid N) - see: fm-dashboard.sh server-status. The operator who just typed restart is told the board is already running, when the actual fact is that the old process refused to exit - the one message that would explain it was thrown away. Dropping 2&gt;/dev/null (the || true alone is what keeps restart alive) makes the real reason visible without changing the recovery behavior.
  • ℹ️ docs/dashboard.md:13 - docs/dashboard.md still says only "bin/fm-dashboard.sh server-status reports the process state and whether the API answers; stop and restart round out the lifecycle." The fix rounds landed after the documentation commit (035403a) and changed operator-visible lifecycle behavior that this page is the owner of: stop is now synchronous and can block up to ~15 seconds, escalates to SIGKILL after ~10, and on timeout intentionally leaves the pidfile in place; restart now starts the board even when stop refuses. The migration/reload notes at lines 94-96 were documented; this half was not, and the page already explains why stop must be synchronous (the migration ordering) without saying what that made stop and restart actually do.
  • ℹ️ bin/fleet-dashboard/web/app.js:22 - The status label is the bare word Review, next to Testing. The intent asks for "sensible labels ... with review reading as the one awaiting him," and this change's own SKILL.md (line 46) states that testing and review "are also easy to conflate, since both put a finished-enough card in front of a reviewer - but the reviewer differs." On the board itself, Review does not say which reviewer it means; a fleet that also does code review can read it as the fleet reviewing. The surrounding signals do carry the meaning - it inherits the old testing gold (--st-review #ffd700), sorts just before Complete, and is the only status rendering the Mark Complete button - so this is polish, not a defect. Flagging for the author's call rather than changing it: a label like "Ready for You" or "For Review" would carry the asymmetry in the word itself.

🔧 Fix: validate recorded pid before signalling the dashboard
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-dashboard.test.sh — 18 checks pass, including testing and review are separate, independently settable statuses, the testing-to-review split migration converts only pre-existing testing cards, and runs at most once, a start that cannot bind leaves the migration pending, and stop waits for the port, restart brings the board back from a crash, from stopped, and from healthy, the lifecycle commands refuse a recycled pid instead of signalling it
  • bash tests/fm-fleet-audit.test.sh — 11 checks pass, including the sweep checks testing cards and never flags one it cannot verify and the sweep flags a working or testing card whose linked crew is demonstrably not working
  • bash tests/fm-dashboard-card-link.test.sh — 42 checks pass, including teardown advances a linked card to review once landed cleanup actually succeeds and teardown never downgrades a card the Admiral already marked complete
  • Manual: seeded a pre-split SQLite DB via store.SCHEMA with 3 testing cards plus one card in each other status, then bin/fm-dashboard.sh start — captured the startup migration report naming all 3 rewritten cards
  • Manual: curl -s $B/api/tasks and curl -s $B/api/tasks/&lt;id&gt; | jq &#39;.status_history&#39; — confirmed 3 cards now review, each with exactly one testing -&gt; review entry carrying the migration note, and updated_at unchanged from its pre-migration value
  • Manual: bin/fm-dashboard.sh restart against the already-migrated live board, then list --status testing and cat state/dashboard.log — no second migration report, in-flight testing card unchanged
  • Manual: bin/fm-dashboard.sh add --status testing / --status review, list --status testing, list --status review, status &lt;id&gt; bogus-status (exit 1), and curl -X POST $B/api/tasks/&lt;id&gt;/status with testing, review, and in-review (HTTP 400)
  • Manual: cross-checked every live card's status against the STATUSES array the web bundle actually renders — zero unrenderable statuses
  • Manual: FM_AUDIT_STALE_NEEDS_ATTENTION_MINUTES=0 bin/fm-fleet-audit-sweep.sh --forced with a blocked crew behind the testing card, then again with a working crew — 1 discrepancy then 0, and none of the 4 review cards flagged in either run
  • Manual (browser, chrome-devtools-axi at 480px): screenshotted the live board, clicked the Review and Testing filter chips, opened a testing card's overlay to inspect the status dropdown, and clicked Mark Complete on the migrated card-rigging card — verified the resulting review -&gt; complete transition through the API
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

joliverMI and others added 23 commits August 16, 2026 18:28
…ng state (#2)

* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(bin): render the captain's four-section status board from existing state

fm-status-board.sh prints one HTML page (Needs you to continue, In
progress, Waiting, Recently completed) rendered entirely from
fm-bearings-snapshot.sh's already-correct cross-home classification and
fm-fleet-snapshot.sh's untruncated per-item detail, with no second store
for agents to keep in sync.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
Tearing down a task deletes state/<id>.meta, the record bin/fm-pr-merge.sh
resolves the PR through. A branch can be fully pushed and reachable from a
remote (so the existing dirty/unpushed/landed checks find nothing to
refuse) while its own PR still sits open, stranding a mergeable PR with no
guarded path to land it - this happened twice in one night.

The new check runs after the existing dirty/unpushed/landed checks inside
validate_worktree_teardown_safety, so their refusal messages take
precedence and never contradict this one. It fires only when a pr= is
actually recorded (scout, local-only, and secondmate teardowns never
record one). A gh lookup error refuses loudly too, naming the PR firstmate
could not confirm, rather than tearing down blind. --force already skips
the dirty/landed checks and now explicitly skips this one too, since it is
the same "captain says discard" escape hatch - the usage comment spells
out what that means here (the PR needs merging or closing by hand,
outside fm-pr-merge.sh's guarded path).

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* fix(ci): fail hung Herdr behavior runs in 20 minutes (kunchenguid#2413)

A wedged family-run step was occupying the runner until the 75-minute
job cap; bound that step so cleanup and timing artifacts still upload.

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* fix(procevent): deliver captured results at most once

A Lavish-captured answer reached a secondmate five times because
forwarding it ran fm-send.sh and only afterward, manually, remembered to
call `handled` - so every re-announcement of the still-unacknowledged
capture (a restart, a compaction, a re-read wake) repeated the forward.
The retry loop was ours, not lavish-axi's: fm-procevent.sh already proves
capture is exactly-once, but nothing paired a downstream effect with its
acknowledgement atomically.

Add `fm-procevent.sh deliver <source-id> <sequence> -- <command>...`,
which checks handled status, runs the command, and marks the generation
handled as one call under the source lock, so a repeated delivery attempt
against an already-delivered generation runs the command zero times and a
failed command stays eligible for retry instead of being dropped. Point
the process-event-sources skill and the runner's operating contract at it
for exactly this class of forward.

Five delivery attempts against one captured generation now run the
downstream command once instead of five times - the read-and-discard
cost of the other four is eliminated structurally. The historical
incident's own token cost was not measured at the time and is not
reconstructable after the fact.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
…#4)

Adds a fifth board section between "Needs you to continue" and "In
progress" for finished work that is deployed and usable, not merely
merged. It renders only when firstmate has written a "ready: <where
to see it>" note on a landed backlog row, confirming deployment and
naming the surface - never inferred from merge/landed state alone, so
a merged-but-not-yet-deployed change never appears here. The item
shows its pointer text, never a PR link or repo path.

Approval buttons (part 2) are not implemented in this PR: investigation
found Lavish's artifact bridge (window.lavish.queuePrompt) can queue an
exact, single-item approval string, but transmission still needs the
existing manual "Send to Agent" tap by Lavish's own design, and the
delivery path could not be confirmed end-to-end from this environment.
Ship the verified section now and report the button investigation as a
plan.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
#5)

The board named locations the captain could not reach from his phone: a
pull request he does not review, a raw data/<id>/report.md path, a "fuller
detail lives in the X home record" cross-home line with nothing to tap, and
a bearings-truncated in-progress summary with no full text behind it.

Pull requests and GitHub links are now never rendered anywhere on the page.
A report links only when its backlog note carries an explicit report-url:
marker firstmate wrote after actually publishing it (the same convention as
the existing ready: marker) - investigated whether Lavish's existing route
could serve reports directly and confirmed it cannot: Lavish is a
publish-one-artifact-at-a-time surface, not a path-mapped static file
server, so there is no passive route to piggyback on. A Ready item's
"where to see it" text is flagged honestly instead of shown verbatim when
it looks like a filesystem path or a pull-request URL. A cross-home item's
bounded detail says plainly that the rest is not reachable from the page,
instead of naming an unreachable home record. An in-progress row now
prefers its fuller current-state text over the shortened one-line summary.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515)

* feat(dashboard): add the Admiral's Fleet Dashboard task board

Replaces the generated Lavish status page with a purpose-built board:
a zero-dependency Python backend (stdlib http.server + sqlite3), a
vendored React frontend styled to match Spectra, an agent-facing CLI
(bin/fm-dashboard.sh) that is the only way agents touch it, and the
fleet-dashboard skill teaching the fleet auditor and other agents how
to use it.

The board owns its own persistent records rather than mirroring any
backlog - its six-status vocabulary and four review tabs don't map
onto tasks-axi state, and the fleet auditor needs a separate claim to
check against live reality. See docs/dashboard.md for the full
architecture, the deploy story, and the drift-risk tradeoff that
choice implies.

Link posting rejects GitHub/PR links and local-only hosts structurally
(standing order 17), not just by agent memory. The page shows a loud,
persistent banner when it cannot reach its own API, and the auditor's
"never run" state is rendered distinctly from a clean sweep, so an
absence of confirmation is never shown as a confirmed-good state.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: joliverMI <joliver@sensibletech.biz>
Testing and needs-attention were conflated: both put a finished-enough
card in front of the Admiral, but only needs-attention actually blocks
on him. A card can now be set to needs-attention with a --reason that
renders directly on the card, sorts above every other status, and gets
its own loud always-visible section on the page. The fleet auditor's
sweep now treats a needs-attention card's age as a finding, since that
status - unlike testing - is not supposed to sit unanswered.

Existing databases migrate in place; the new column is nullable and
existing cards, history, and statuses are untouched.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…orce Audit button (#8)

The auditor's 15-minute cadence never actually ran: it registered a durable
check in its own home, but nothing polled it because that home had no live
supervision session. Host cron (already running independent of any chat
session) now drives bin/fm-fleet-audit-tick.sh, which heartbeats every tick
and sweeps via bin/fm-fleet-audit-sweep.sh once the dashboard's own interval
setting says it's due, self-healing a lock abandoned by a crashed sweep. The
Force Audit button uses the identical sweep script through a detached
subprocess so a scheduled and a forced run can never stack. Also fixes the
"unassigned" agent noise on every card (shows the captain instead) and adds
a Go to card jump from each discrepancy log entry.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…ard (#9)

Every dashboard card's agent and backlog_ref sat empty, so status
transitions only happened when someone remembered to make them by
hand - worst of all at teardown, the exact moment the worker that
would have remembered is being removed.

fm-spawn.sh --card <card-id> now records dashboard_card=<id> in the
task's own state/<id>.meta and best-effort links the card's ref/agent
to the task identity, advancing a not_started card to working.
fm-teardown.sh reads that identity back and advances the card to
testing once a real (non-forced) landing succeeds, leaving scout,
secondmate, and already-complete cards untouched.

No card is the normal case and stays a complete no-op. A bad card id
or an unreachable dashboard never fails the spawn or the teardown; it
warns loudly and best-effort records fm-dashboard.sh audit-log
--fleet so a failed link is visible to the fleet auditor rather than
silently dropped.

All 21 live cards still carry no agent/ref - 20 belong to a project
this task has no access to verify task identities for, and the one
firstmate-owned card is already complete with no live task to link -
so none were backfilled; see the PR description for the per-card list.

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…nd sends (#12)

* fix(bin): tmux endpoint liveness reports dead targets as alive

fm_backend_target_exists's tmux branch trusted display-message's exit
status alone. tmux silently falls back to the caller's own active pane
(or an unrelated live session/window whose name is an unambiguous
PREFIX of the missing target) while still exiting 0, so a dead tmux
endpoint reads alive whenever the caller itself runs inside tmux -
which firstmate always does - or shares a socket with any other live
session/window. Pin the tmux probe to an exact target: session:window
pairs use list-panes -t "=$session:=$window" (the same exact-match
form fm_backend_tmux_kill already uses), bare %N pane ids / @n window
ids / $N session ids pass straight through since list-panes already
resolves them exactly, a bare (colon-less) target is checked against
exact session and window-name inventory since tmux's `=` pin does not
make a bare target exact the way it does the session component of a
session:window target, an all-digit window component resolves exactly
against the session's own window-index inventory (FM_SUPERVISOR_TARGET
_DEFAULT is "firstmate:0", an index not a name), and a dotted window
component is tried first as an exact window name and, if that fails,
falls back to pane-qualified addressing (window.pane or
window-index.pane) - both are valid, competing readings of the same
string in real tmux.

fm-crew-state.sh's pane_readable duplicated the same probe and now
defers to the fixed primitive. A task record naming a remote host is
no longer judged by the local tmux server at all in either caller:
fm-session-start.sh's digest reports those as an explicit "not checked
from here" verdict instead of probing a session that can never resolve
locally, and fm-crew-state.sh skips the local probe entirely for a
remote_host record so its status-log fallback - previously reached
only by accident, through the old probe's own bug - still runs.
fm-fleet-snapshot.sh, fm-control.sh, and fm-send.sh were checked and
already guard remote_host before reaching this probe.

herdr, zellij, orca, and cmux already address their target by id
through a structured inventory query, so they do not share this
defect shape.

Adds a real-tmux regression suite that runs the probe from inside a
live tmux pane (the one condition that reproduces the false positive,
since the fallback needs a live session to fall back to) against
nonexistent, prefix-colliding, dotted, pane-qualified, bare-name, and
window-index targets, and a crew-state case proving a remote
secondmate's status is still read when the local probe always fails.
All fail against the pre-fix code and pass after. Updates fake-tmux
test fixtures that modeled only display-message's semantics to model
the pinned probe instead.

Known gaps, deliberately not fixed here and tracked in one follow-up,
fm-tmux-agent-state-session-prefix-match: the recovery-grade sibling
fm_backend_tmux_agent_state (bin/backends/tmux.sh) still resolves its
session component by unpinned prefix match, same defect class,
unfixed; and fm_backend_tmux_kill can silently destroy a DIFFERENT
live window (not merely fail to remove the target one) when given a
dotted window name - verified live, also unfixed. No task id in this
fleet currently contains a dot, so nothing live is exposed by either
gap today. The follow-up will converge every tmux target-resolution
helper in this repo onto one shared resolver rather than each parsing
target strings independently.

Delivery note: this branch was built by hand directly on this fork's
own current main tip (git apply --3way of the reviewed diff) rather
than by rebasing, because this fork's history has diverged from the
upstream template repository it was created from and a rebase would
pull in commits that exist only on that upstream - this repo's
standing rule is that nothing is ever proposed to it.

* no-mistakes(review): refuse tmux key sends to targets that resolve inexactly

* no-mistakes(review): guard and exact-pin every tmux input-delivering send

* no-mistakes(review): resolve tmux probe and send targets through one resolver

* no-mistakes(review): refuse ambiguous bare tmux names, correct stale disclosures

* no-mistakes(test): add list-panes arm to shared secondmate fake tmux

* no-mistakes(document): align endpoint-verdict and supervisor-target docs with exact tmux resolution

* no-mistakes: apply CI fixes

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
* feat(dashboard): mechanically link handed-off backlog items to their card

Cards for work routed to a secondmate had no mechanism to advance: status
advances only through bin/fm-spawn.sh --card and bin/fm-teardown.sh, both
keyed off local task metadata that bin/fm-backlog-handoff.sh never creates,
so a card handed to a secondmate froze at not-started forever.

bin/fm-backlog-handoff.sh --card <card-id> gives a single handed-off item
the same best-effort link: once it lands in the secondmate's backlog, it
sets the card's ref/agent and advances a not_started card to working. The
durable identity lives in this home's own state/handoff-cards/<id> record
(never in the backlog item, which the handoff does not own) and is retired
per pair only once the board genuinely confirms that pair's link - never on
a merely-attempted one, so a failed link is retried by the next arrival
instead of silently orphaning the card. bin/fm-dashboard.sh's own curl calls
are now bounded so an unreachable board cannot stall a held handoff lock.

Includes a fix to the shared dashboard-card-link test's own server-start
wait: it previously trusted bin/fm-dashboard.sh's 1-second process-liveness
sleep as proof the server was ready, which flakes on a busier CI runner
(confirmed pre-existing on main, not introduced here). It now polls the
real /api/health reachability check before running any test.

* no-mistakes(review): bound dashboard curl calls, fix card-link identity, 404, and record races

* no-mistakes(review): confirm card links only on read state, keep unlinkable pairs

* no-mistakes(review): fail card-link ownership check closed, never link superseded cards

* no-mistakes(document): fix stale usage line for multi-item handoff

* fix(tests): show captured output on a failed exit-code assertion

expect_code only ever reported "expected exit 0, got 1", with no way to
see why - the same missing-evidence shape as the CI failure this is
meant to help diagnose. Accept the command's own captured output as an
optional argument and print it on failure, matching how
assert_contains/assert_not_contains already show theirs. Wired into
every expect_code call in tests/fm-dashboard-card-link.test.sh.

Also fold in the changed-file test-family mapping this branch's rebuild
had dropped: bin/fm-backlog-handoff.sh, bin/fm-spawn.sh, and
bin/fm-teardown.sh (the three entry points to the mechanical card link)
plus bin/fleet-dashboard/* now select the dashboard family, so the
suites protecting this change actually run when the files they protect
change.

* no-mistakes(review): sweep local card records on --resume-pending, reject zero timeout

* no-mistakes(review): hold back only still-staged pairs on --resume-pending

* no-mistakes(document): document handoff card record write and resume sweep

* no-mistakes(review): fail loudly on card record rewrite errors, clear stale warnings ledger

* no-mistakes(review): commit superseded marks after record write, gate ledger reset

* no-mistakes(review): reject separator-bearing --card ids before they reach the record

* no-mistakes(review): drop durable warnings ledger for in-process report suppression

* no-mistakes(review): gate unreadable-card warning through reason-keyed report mark

* no-mistakes(review): replace bash-4 assoc array with 3.2-portable report set

* no-mistakes(document): scope outbox recovery claims, fix card-link report bound

* fix(dashboard): stop flooding the fleet audit log for unretirable pairs

An unretirable pending pair (the board says a card id does not exist)
was writing a durable audit-log --fleet entry every time it was swept.
bin/fm-bootstrap.sh's --resume-pending sweep now runs unattended on
every session start, so a single permanently-unlinkable pair would
grow that entry on a cadence no operator controls - the exact
always-red-marker failure the fleet auditor's own design already
guards against.

The stderr warning stays exactly as built, still bounded in-memory per
process; only the fleet-log write is removed. Surfacing a permanently
unlinkable pair to the Admiral belongs to the fleet auditor's own
sweep, raised once as a finding, not to this handoff path on every boot.

* no-mistakes(review): handle unterminated final lines in handoff card state

* no-mistakes(document): explain unterminated-line handling in handoff card state

* fix(dashboard): never write the fleet audit log from --resume-pending

The no-such-card case already stopped writing to the fleet audit log
from bin/fm-bootstrap.sh's unattended --resume-pending sweep. This
applies the same rule to the general transport-failure case in
dashboard_link_card: a ref/agent/status write that fails still warns
on stderr everywhere, but only an operator-initiated invocation (no
--resume-pending) may write it to audit-log --fleet, and at most once
per invocation.

* no-mistakes(review): surface pending handoff card links at session start

* no-mistakes(document): document spawn/teardown fleet-log writes and card-link resume gating

* no-mistakes(document): note superseded pairs as a persisting card-link cause

* no-mistakes(ci): drain the denied arm attempt before flipping the lock

tests/fm-pi-watch-extension.test.sh's OpenCode session-lock case wrote a
foreign pid, fired session.idle, slept a fixed 120ms, then took the lock and
fired again. The event hook does not await its arm attempt, and that attempt
walks the process tree with one `ps` per level, so on a loaded runner the
denied attempt was still in flight at 120ms. ensureArm coalesces onto
launchInFlight, so the second event joined the stale read-only answer and
never armed - the shard's "expected exit 0, got 1".

Wait on the plugin's own coordinator handle instead of the clock: ensureArmed
joins an in-flight launch, so once it resolves no attempt is outstanding, and
its status is the denial under test. Also pass the node output to expect_code
so a future failure here shows why rather than an exit code alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#13)

* fix(bin): count a trust-registered custom check as supervision-needed

fm_supervision_status only looked for state/*.meta, state/x-watch.check.sh,
and state/procevent/*.source. A home whose only duty is a check registered
through bin/fm-check-register.sh (a state/<id>.check-trust binding paired
with its state/<id>.check.sh) matched none of those, so FM_SUP_NEEDED stayed
false and the Claude Stop auto-arm never armed a watcher for it.

Count a state/<id>.check-trust file paired with its state/<id>.check.sh as
supervision-needed, mirroring the existing presence-only style used for
process-event sources and the X-mode relay poll. Update the turn-end guard
and pull-warning banners to name a registered check specifically instead of
falling through to the X-mode wording, and update docs/turnend-guard.md and
its cross-references (docs/architecture.md, AGENTS.md, docs/scripts.md,
docs/subagent-guard.md, docs/supervision-protocols/grok.md, and
bin/fm-claude-stop-autoarm.sh's own header) to point at docs/turnend-guard.md
as this predicate's single owner rather than restating its input list.

* no-mistakes(document): note check-predicate rationale, generalize stale guard comments

* no-mistakes(document): record check-predicate verification evidence entry

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
…14)

* fix(watcher): re-evaluate arm coalescing when the fleet lock changes

Two independent causes made the arm-readiness suite fail a different
assertion almost every run:

- ensureArm() in the OpenCode watch plugin reused a still-resolving
  earlier caller's beginArm() result unconditionally. Every ordinary
  session.idle produces two callers, so when the fleet lock was
  reacquired while an earlier attempt was mid-flight, the later
  caller inherited that attempt's stale read-only verdict and never
  armed. Fixed with premise-validated coalescing: a caller shares an
  in-flight attempt only while the lock file content it captured is
  still current; otherwise it evaluates fresh. Two callers on an
  unchanged lock still coalesce into one subprocess walk, so this
  costs nothing on the ordinary turn a serialized-everything fix
  would have doubled.

- Both adapters spawn their arm child through a login shell, which
  sources /etc/profile in addition to the account's own profile
  files; the system-wide half is not relocatable via HOME. main
  already raised the readiness timeout to absorb that cost (250ms to
  2000ms) and its own measurement shows a worst case of ~1740ms
  under contention against that budget - narrower headroom, not a
  removed confound. This adds FM_WATCH_ARM_NO_LOGIN_SHELL, a
  test-only opt-out that spawns the arm child under plain bash -c,
  removing the confound instead of padding around it. Production
  keeps the login shell as the unconditional default.

Also fixes three unhandled-EPIPE crash sites found while proving this
change under load: child.stdin.end() in fm-primary-turnend-guard.js,
fm-primary-turnend-guard.ts, and fm-operational-input.js raises an
unhandled 'error' event on the stdin stream (not the ChildProcess)
when the child exits before the write lands, which was crashing the
whole session process. Each site now no-ops that stream error since
the child's own close/error handlers already drive resolution.

Adds two regression tests (opencode coalescing on an unchanged lock,
login-shell default vs. opt-out for both adapters) and one EPIPE
regression test, and rewrites the existing OpenCode lock test to
force the stale-verdict race deterministically via a gated ps shim
instead of waiting for it. See docs/arm-readiness-determinism-proof.md
for the repeated-run proof and docs/configuration.md for the new
FM_WATCH_ARM_NO_LOGIN_SHELL entry.

Built directly on current main (45bd292); does not touch main's own
prior timeout raise or its test_pi_session_transition_generation_owner
fixture-ordering fix, both kept as-is.

* no-mistakes(review): correct stale arm-readiness proof claims and restore expect_code diagnostics

* no-mistakes(review): wait on observable arm rows and share gate shims

* no-mistakes(review): restore fallback diagnostics and bound unretired arm counts

* no-mistakes(review): correct lock-assertion cause attribution in proof doc

* no-mistakes(review): restore two-cause attribution and note login-shell residual bound

* no-mistakes(review): qualify stale readiness-window claim on login-shell test

* no-mistakes(review): correct stale readiness-budget claims in suite header

* no-mistakes(test): correct stale login-shell contention figures in determinism proof

* no-mistakes(document): scope proof-doc timeout raise and idle figure to sources

* no-mistakes(document): stamp determinism proof verification at a15d993 with provenance

---------

Co-authored-by: joliverMI <joliver@sensibletech.biz>
`testing` on the Admiral's Fleet Dashboard carried two different meanings:
done and ready for his review, and the fleet actively exercising it right
now. Split it into `review` (the old meaning - optional to him, never
age-checked by the auditor) and a redefined `testing` (live fleet activity,
corroborated by the fleet auditor exactly like `working`).

Every card in `testing` before this change meant the old thing, so
bin/fleet-dashboard/server/store.py migrates them to `review` once, gated
by a settings marker so a later genuine `testing` card is never rewritten.
bin/fm-teardown.sh now advances a landed card's mechanical link to `review`
instead of `testing`. The web board, CLI, HTTP API, fleet-dashboard skill,
and docs/dashboard.md are updated to match.
@joliverMI

Copy link
Copy Markdown
Author

Opened in error against upstream - this work belongs on our own fork and is being redirected there. Apologies for the noise.

@joliverMI joliverMI closed this Aug 21, 2026
@joliverMI
joliverMI deleted the fm/fm-dashboard-review-status branch September 6, 2026 02:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants