Skip to content

feat(bin): add Bridge captain dashboard with SQLite-backed answer queue - #3689

Closed
wonder-media wants to merge 9 commits into
kunchenguid:mainfrom
wonder-media:fm/fm-captain-board
Closed

wonder-media wants to merge 9 commits into
kunchenguid:mainfrom
wonder-media:fm/fm-captain-board

Conversation

@wonder-media

Copy link
Copy Markdown

Intent

Build phase 1 of Bridge, the captain dashboard: a self-updating, SQLite-backed fleet board served by one always-on local daemon, replacing the hand-regenerated .lavish/fleet-board-gen.py and master-board.html pair. The page title and header wordmark are Bridge; bin/fm-board.* and CLI names stay unchanged.

The final plan at /Users/patrick/firstmate/data/fm-captain-board/plan.md is the contract. This intent consolidates its v2 plus all 14 accepted v2.1 points in their current accepted form. Phase 1 includes database, ingest excluding GitHub, server, decisions, answer round trip, UI/UX, files, and relevant verification. GitHub mirroring, GitHub calls, brief-scaffold producer rule, secondmate charter producer changes, deploying production configuration/launchd, and retiring old Lavish boards/pollers are later phases, explicitly excluded. Published reads Not connected yet in phase 1. Never mutate .lavish/, config/, state/, or launchd in the primary checkout. The old primary .lavish/fleet-board-gen.py is read-only reference for tasks-axi held parsing, atomic writes, and JSONL ts/home/task/choice/note plus key/answer_id. Do not reuse directory-mtime looping, merged/LIVE-substring classification, or wildcard CORS. No LLM tokens anywhere in daemon or handler; no gh calls in this phase.

Deliver all shared files tracked through this PR:

  • bin/fm-board.py daemon, Python 3.14 stdlib only (including sqlite3/http.server/json/threading/subprocess/uuid). Header owns config/board.json schema, all HTTP routes, health fields and SSE contract. Do not assume /usr/bin/python3 (3.9); use actual Python 3.14, such as /opt/homebrew/bin/python3.
  • bin/fm-board.sh shellcheck-clean CLI: serve [--config ]; ingest --once [--home ]; decision --project --title --option "A: ..." [--option ...] [--rec A] [--why ]; live --url --env [--evidence ]; answered marks consumed; refresh touches FM_HOME/state/board-refresh; arm-answers registers board-answers process-event source through bin/fm-procevent.sh; backup uses VACUUM INTO, keeps 7. Validate every subcommand's input; bad input exits nonzero with one-line reason; all subcommand --help exits 0. Read fm-procevent --help and process-event-sources skill before working on registration.
  • bin/board/board-answers-source.py ported out of .lavish with env/config paths, not hardcoded; bin/board/board-answers-handle.sh deterministic burst handler. Python implementation helper and dedicated process-event adapter are permitted. Registered exact (home,task,key) answers route with bin/fm-send.sh --resolve-key '' for live workers or bin/fm-decision-hold.sh answers for backlog captain holds, then fm-board.sh answered . Keyed answer with note sends ' (note: )' to worker, never firstmate solely because note exists. Unknown/unkeyed/failed route yields bounded structured packet for firstmate.
  • bin/board/dashboard.html one static file, no build/CDN, dark only.
  • documented bin/board/board.example.json and bin/board/com.wondermedia.firstmate-board.plist.example: absolute Python 3.14 path placeholder, WorkingDirectory, explicit FM_HOME/config, KeepAlive, ThrottleInterval 10, StandardOut/ErrorPath under state/logs. Firstmate installs actual config/plist/bookmark after merge.
  • tests/fm-board.test.sh on existing runner, colocated tests/.test.sh convention; run bin/fm-lint.sh before done.
  • docs/configuration.md short Captain dashboard entry pointing at fm-board.py header. AGENTS.md section 2 only five one-line layout entries in neighboring style: config/board.json, state/board.sqlite, state/board-inbox/, state/backups/, state/logs/. Nothing else in AGENTS.md. Tracked Markdown one sentence per line, plain dashes. No agent co-author. Follow firstmate-coding-guidelines skill.

Database: main home owns one private state/board.sqlite, WAL, schema_version/migrations at startup, per-thread/call connections, busy_timeout, short transactions. Secondmate homes are ingest sources. projects(tag,home_id,repo,board_url) uses config repo_tags by repository, never home: WOK lives in MF home; CES and CSLS-OG share home. tasks keyed home/task with ELI5 backlog title, kind,current_state,worker,pr_url,updated_at,deleted_at tombstone. decisions keyed home/task/key/revision with question, options JSON, recommendation, source worker/firstmate/hold, open/queued/sent/consumed/failed/closed states, asked_at/closed_at. Captain holds actionable only when captain_actionable and no unresolved blockers. events append status/merged/built/live with project, live requires URL/environment/verified_at/evidence; prune after 30 days nightly. answers UUID and decision identity/revision/choice/note/device/received_at/exported_at/consumed_at/error, uniqueness identity+revision+choice+note hash prevents duplicate phone/desktop submits; stale revision 409 returns current decision. board_items schema placeholder for later GitHub mirror itemid/fetched_at/stale. ingest_runs per-home last_ok,last_error,duration_ms. Nightly VACUUM INTO state/backups/board-.sqlite, keep7, test restore.

Ingest: every5s explicit per-home file manifest fingerprints mtime_ns,size,inode for state/.meta, state/.status, data/backlog.md,data//report.md. NEVER directory mtimes. Ignore watcher internals .count-,.hash-,.last-watcher-beat,.stale-*, locks, DB,inbox,logs. Only changed ids refreshed via fm-classify-lib.sh status_open_decisions and bin/fm-crew-state.sh ; tasks-axi list held and in-flight for backlog. Batch per-home checks into few invocations, nice -n 10; HTTP independent of ingest. No full snapshot in5s path. Full bin/fm-fleet-snapshot.sh --json only startup/wake,15min reconcile or board-refresh, hard90s timeout (live home measured66s). Per-id subprocess hard10s timeout. Single flight/debounce: ticks while ingest running only dirty flag/coalesced next pass, no second process. Failure retains last-good rows, records ingest_error; full successful pass tombstones vanished .meta; partial pass unknown. Rev monotonic once per visible committing write, no timestamp-only bump. SSE changed carries rev/generated_at. Never classify merged or done or LIVE substring as green. Verified status live convention or explicit live registration requires URL/env/verified_at. Dashboard ingest/parse errors only ingest_runs,/healthz,red page connection state, never wake firstmate.

Server: Python3.14 stdlib, ingest/HTTP threads, explicit routes only: GET /, GET /api/state?project=TAG includes rev/per-home freshness/connection, GET /events, GET /healthz, POST /answer, POST /answers, plus explicitly documented answer controls and loopback POST /internal/reload. Every API request Authorization: Bearer secret. Initial ?k= bootstrap stored in localStorage then immediately history.replaceState to bare path. Secret required on every POST. Validate Host and Origin against lan_host, no wildcard CORS, cap requests, data rendered text nodes, validate URL schemes. Plain private LAN HTTP accepted, HTTPS later. LAN accessible not loopback-only main page. /internal/reload secret-required and bound127.0.0.1; decision/live CLI POST it to push SSE within1s. /api/state ETag rev; If-None-Match following SSE =>304 unchanged. SSE15s heartbeat,boundedclientqueues,sockettimeouts,BrokenPipe handled,noheldSQLtransaction. health ok,ingest_age_s perhome,last_snapshot_ms,sse_clients,db_ok,outbox_backlog,answers_armed. GracefulSIGTERM,reconcile startup/wake,daemon size-based logrotation. Source armed startup/rearmed afterhandledfire, answers_armed honest.

Decisions: options registered through schema-validated CLI JSON, never || status delimiter and no brief-scaffold rule in this phase. Card projectbadge first,plainELI5title,one-sentenceconsequence,fullwidth44px radio option rows, no defaultselection. Recommendation inline on option row with label/reason, never green. Legacy prose decisions: summary, Send custom answer (note required), Request concrete options, never invent Approve/Hold. Request concrete options is EXACT fixed handler steer: Captain requested structured options for decision : register 2-4 distinct alternatives with bin/fm-board.sh decision <home> <task> <key> --option '...' and stop.

Answer round trip (v2.1 current contract supersedes original normal-path JSONL): SQLite answers row IS durablequeue. POST validates secret,identity,revision,option and commits row; never markanswered onPOST. Daemon directly runs bin/board/board-answers-handle.sh forburst single-flight/idempotentanswer_id with consumed_at/error writtenback. Preserve15s queued undo before irreversible delivery. Replay rows consumed_atNULL/errorNULL aftercrash beforehandler. JSONL ONLY exceptions (unkeyed,unknown,failedroute), no dualwrite normalpath. Exceptions retain bytecompatible ts,home,task,choice,note,key,answer_id; atomic/idempotentexport byanswer_id with crashbetweenpublication/exportmark replay exactlyonce. One JSONLburst and ONE process-eventfire forbatch3exceptions. Source wakesfirstmate onlyexceptions, daemonrearms eachhandledfire. Keyed+note routedworker. Open->queued->sent->consumed/failed; consumed means actualhandlingack, card collapses24h answered drawer; failedredreason. Afterconsumption Request correction only. Unknown/unkeyed/failed leftforfirstmate structuredpacket. Dashboard errors themselves never exceptionwakes.

UI: landing Waiting on you fleetwide; chiprow All/WOK/CES/MF/CSLS-OG/JVP/WM/FM/Charlier, horizontalphone scroll, unreaddecisioncounts, filterALL3bands. Threebands Waiting on you(decisions,captainholds,yourcheck), Happening(inflight,externalwaits,built not on site amber), Published(phase1 Not connected yet). Compactwaitingcount visibleeveryband. Activework neutralborder/outline notfourthlight; greenonlyverifiedliveclaim,amberbuiltunverified,redcaptainwaitorconfirmedfailure,agestextnotcolor,everylighttextlabel. Maxwidth1280,gutters24desktop16phone,8pointspacing,16body1.5,13meta,19cardtitle,24pagetitle,systemfont,nativecontrols,CSSvars,minimalinlineSVG. Palette --bg#0e1116 --panel#161b22 --tx#e6edf3 --dim#9aa7b4 --grn#2ea043 --amb#d29922 --red#da3633. Noemoji,CDN,lighttheme,5columnkanban,19stepwizard,secondmatecoordinatorchatter,blue/gold/slatestatuscolors.

Everycard alwaysvisible labeledcompactone-line note expandsfocus, nohidden toggle. Selection editsdraft; explicit Confirm shows chosenwording. Batchlabel Submit drafted (M of N), disabledno selections, submitsONLYexplicitdraftselections. Undo15s whilequeued/unconsumed. PersistentlocalStorage drafts home|task|key|revision, serverwins; revisionchange retainsolddraftwithreviewflag. SSE patchcardidentity preservingfocus/caret/selection/scroll/order, newitems politeannouncement nevermoveeditedcard, neverlocation.reload. HeaderFleetchecked relativeage/absolutePT hover plusperhomestale, GitHubnotconnectedphase1, freshness separateconnectiondot (greenheartbeat<20s,amberfallback,redPOSTfailure). No15sJSONpoll onceSSEworks; fallbackwhenbroken. Losskeepslastpaint, localStorageconfirmsqueue/retry, Disconnected showing last known state. FailedloadneverfalseNothingwaiting. Specificempty Nothing is waiting on you in CES. / No verified live change today. Phonecards/optionsstack44pxtargets,noteaboveConfirm,nostickykeyboardoverlay. visibilitychangevisible ifheartbeatolder25s close/refetch/reconnect. Keyboard arrowschips,SpaceEnteractivate,j/kcards,1-9option,nnote,Enterconfirm,uundo,visiblefocus.

Verification required: scratchFM_HOME/scratchconfig daemon page/api reflecttasks,held/openkeyeddecisions,healthok/answersarmed; statusappend =>SSEcardwithin10s no reload; CLI options/recappear; confirms desktop/390phone seconddevicequeued,stale409,dedupe,SQLite row/sent thenconsumed (normalrouteNOJSONL perv2.1; exceptionrouteJSONLonly). Tests:fixtureingest,tombstones,revonce/notimestamps,singleflight tick/no2ndprocess,timeoutlastgoodingest_error,decisionCLIoptions,POSTmissingsecret403/stalerev409currentdecision/duplicatesdedupe,batch3exceptionsoneJSONLburstonefire,killbetweenSQLite/exceptionexportrestartonce,SSEchangedrev,healthfields,backuprestore,allsubcommandhelp0. Keyedhandlerfm-sendresolvekey inclnote,heldroute,fixedrequestoptions. MergedwithoutdeployPublishedempty/Happeningamber. ExerciseEVERYcontrol realchrome-devtools-axi desktop AND390pxphone,keyboards only pass,refreshmidnotepreservesnote. Tests green,fm-lint clean,docs updated. Launchctlproductionkickstart andGitHubcallratemeasurement deferred because this phase cannot installlaunchd/callGitHub; validate scratchserve andexampleplist instead.

Later explicit environment safety requirement: keep env-isolation fix and regression proving daemon, runner Popen, and every test cleanup only ever touch fixture FM_HOME. Strip FM_STATE_OVERRIDE, FM_DATA_OVERRIDE, FM_PROJECTS_OVERRIDE at all external launch/cleanup boundaries. Assert sweep-home is NEVER invoked against nonfixture state dir. Tests poison all3overrides at a sentinel home and assert unchanged; copied fixture helper boundary refuses nonfixtureFM_HOME beforemutation. Deliberate narrowly scoped exception inside testguard: existing generic procevent sweep recursively invokes retire with explicit FM_STATE_OVERRIDE equal to that SAME fixture state; permit that internal recursive retire only, while top-level register/start/sweep must have all3overrides stripped. Do not weaken protection or touch primary home during validation. Re-arm/retirement tests must remain fixture-scoped. Exception-source durablecapture test and browser acceptance completed beforecommit.

Validation context (not replacements for requirements): implementation committed5becd05; boardtest passed, fm-lint clean using worktree-local pinnedShellCheck0.11/actionlint1.7.12, nestedhandlerlint clean; docscoverage/pycompile/plutil clean. Scratchbrowser evidence private .no-mistakes/board-check/verification.md and bridge-desktop.png/bridge-phone.png. Allscratchdaemons/sources stoppedafteracceptance. Additionaltracked CI setup-python3.14 ensuresrequiredtestruntime. Shared no-mistakes daemon must never be restarted/upgraded/stopped. Driver must not hand-edit whilepipelineactive; applyfixes throughpipeline. Ask-user findings belongtofirstmate; no --yes. Stopatchecks-passed withoutmerging.

What Changed

  • Add the Bridge fleet dashboard: bin/fm-board.py is a Python 3.14 stdlib daemon that ingests per-home task state into a private WAL SQLite database, serves the static dark-only bin/board/dashboard.html, exposes bearer-secret HTTP routes (/api/state, /events SSE, /healthz, /answer, /answers, loopback /internal/reload), and is driven by the shellcheck-clean bin/fm-board.sh CLI (serve, ingest, decision, live, answered, refresh, arm-answers, backup).
  • Add the answer round trip: captain answers are committed as durable SQLite rows, then bin/board/board-answers-handle.sh routes keyed answers to live workers via fm-send.sh --resolve-key or to backlog holds via fm-decision-hold.sh, while unkeyed, unknown, or failed routes are exported as a JSONL exception burst that wakes firstmate once through the board-answers process-event source. Review fixes dropped the empty (note: ) suffix, added backoff and honest answers_armed reporting when another owner holds the source, pinned the handler to the daemon's Python, and folded the project filter into the /api/state ETag.
  • Add tests/fm-board.test.sh and map bin/board/ paths to it in bin/fm-test-run.sh, ship board.example.json and a launchd plist example, add a setup-python 3.14 step to CI, and document the board in docs/configuration.md, AGENTS.md, and the process-event-sources skill. The branch also carries five previously landed local-main commits (treehouse teardown and slot rebind safety, Claude credential .gitignore entries, and the herdr project-spaces, brief-evidence, and autocompact landings) that were not yet on origin/main.

Risk Assessment

✅ Low: All five previously accepted findings are fixed correctly and narrowly within the phase-1 contract, each with focused regression coverage, the default /api/state ETag and the exact request-options steer are preserved, and the fixture FM_HOME guards plus stripping of FM_STATE_OVERRIDE/FM_DATA_OVERRIDE/FM_PROJECTS_OVERRIDE remain intact in the daemon, the runner Popen, and test cleanup.

Testing

Completed 1 recorded test check.

  • Outcome: ⏭️ skipped across 3 runs (2h16m56s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Push main to origin, or rebase your branch onto origin/main, before gating.

⚠️ **Review** - 1 info
  • ⚠️ bin/board/board-answers-handle.py:96 - answer_text = f&#39;{wording} (note: {note})&#39; appends the note suffix unconditionally, so an answer submitted with no note is delivered as Ship it (note: ). The dashboard makes the note optional for every non-custom choice (chosen() in bin/board/dashboard.html only requires a note when choice === 'custom'), so this is the common path. Failure scenario: captain selects option A on a registered decision and leaves Note blank -> POST /answer with note:'' -> route() sends fm-send.sh &lt;task&gt; --resolve-key &lt;key&gt; &#39;Ship it (note: )&#39; to the worker. The user intent specifies the format only for a keyed answer with a note; append the suffix only when the normalized note is non-empty.
  • ⚠️ bin/fm-board.py:899 - service_loop re-arms with no backoff whenever the process-event runner is not alive: every ~0.5 s it calls arm(), which runs a blocking fm-procevent.sh register subprocess and then Popens fm-procevent.sh start. cmd_start_public -> cmd_start prints already owned and exits 0 immediately when another home/watcher already holds the source claim (bin/fm-procevent.sh:335), and fm-procevent.sh reconcile is documented to start a runner for any registered source with no live owner. Failure scenario: the board daemon's runner exits after a handled fire, the watcher's reconcile claims board-answers first, and from then on the daemon spawns two subprocesses per second and rewrites the registration file (mv -f, new inode/identity) indefinitely until the other owner's runner exits. Health also reports answers_armed=false and ok=false for that whole period even though the source is genuinely armed, contradicting the 'answers_armed honest' requirement. Add a backoff when the runner exits quickly, and treat an 'already owned' exit as armed.
  • ⚠️ bin/board/board-answers-handle.py:103 - The hold route builds a tab-separated line f&#39;{a[&#34;decision_key&#34;]}\t{answer_text}\tCaptain dashboard\n&#39;, but only note is whitespace-normalized (line 95); wording comes straight from the stored option label. The decision CLI matches options with re.fullmatch(r&#39;([A-Z0-9]):\s*(.+)&#39;, option, re.S) (bin/fm-board.py:340), and text() explicitly permits '\n' and '\t', so an option label may contain either. Failure scenario: an option registered as --option $&#39;A: Approve\nwith conditions&#39; is chosen on a captain hold -> stdin becomes two lines -> fm-decision-hold.sh command_answers reads the second line, its first field fails the slug check and hits continue (bin/fm-decision-hold.sh:771), so the remainder of the captain's answer is silently dropped and is not counted in skipped, leaving exit status 0. Normalize wording the same way as note before building the line.
  • ⚠️ bin/board/board-answers-handle.sh:11 - exec &#34;${FM_BOARD_PYTHON:-python3}&#34; falls back to whatever python3 is first on PATH, with none of the >=3.14 enforcement that bin/fm-board.sh applies, and Board.run() (bin/fm-board.py:157) does not propagate FM_BOARD_PYTHON=sys.executable to child processes. Failure scenario: the daemon is started as /opt/homebrew/bin/python3 bin/fm-board.py serve without FM_BOARD_PYTHON (launchd PATH resolves python3 to /usr/bin/python3, 3.9); route_loop then runs board-answers-handle.sh under 3.9, and route() in turn calls bin/fm-board.sh answered &lt;aid&gt; which exits 1 with 'Python 3.14 or newer is required' after fm-send.sh already delivered the answer to the worker -> the delivered answer is recorded as failed and escalated to firstmate as an exception. The daemon should export FM_BOARD_PYTHON=sys.executable in run(), and/or handle.sh should apply the same version guard as fm-board.sh.
  • ℹ️ bin/fm-board.py:1017 - The /api/state ETag is f&#39;&#34;{state[&#34;rev&#34;]}&#34;&#39;, derived from rev alone, but the response body is filtered by the project query parameter. Failure scenario: a client fetches /api/state (All), then fetches /api/state?project=CES with If-None-Match set to the same rev -> 304 Not Modified, and the client keeps the unfiltered payload. The shipped dashboard always requests /api/state with no project and filters client-side, so this is not reachable from bin/board/dashboard.html today; it only bites a future or third-party client. Include the project in the ETag.

🔧 Fix: Fix answer text, re-arm churn, handler Python, project ETag
1 info still open:

  • ℹ️ bin/fm-board.py:838 - armed_now() reports the external-owner case from cached state (self.armed and self.armed_elsewhere), which maintain_arm() only refreshes when the backoff timer expires; arm_delay is not reset on the 'live' branch, so it climbs to 60 s and stays there. Failure scenario: another home's runner owns board-answers, the board reports answers_armed=true, that runner dies, and /healthz keeps reporting answers_armed=true (and ok=true) for up to 60 s before the next source_owner() check re-arms locally. This is a bounded staleness window in a polled indicator and is a large improvement over the previous permanent false negative, so it is recorded as an acknowledged tradeoff rather than a defect.
⏭️ **Test** - skipped
  • 🚨 tests failed with exit code 2
  • bin/fm-test-run.sh --changed --exclude-family real-herdr-gated

🔧 Fix: Map bin/board paths to the fm-board changed-test suite
1 error still open:

  • 🚨 tests failed with exit code 1
  • bin/fm-test-run.sh --changed --exclude-family real-herdr-gated

🔧 Fix: Classify eight changed-suite failures as pre-existing Bash 3.2 environment
1 error still open:

  • 🚨 tests failed with exit code 1
  • bin/fm-test-run.sh --changed --exclude-family real-herdr-gated
✅ **Document** - passed

✅ No issues found.

⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 127)
✅ **Push** - passed

✅ No issues found.

wonder-media and others added 9 commits August 17, 2026 21:27
* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix(ci): keep CLAUDE.md pointer check valid (kunchenguid#2515)

* ci: gate GitHub workflows with pinned actionlint (kunchenguid#2517)

* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation

* fix: install pinned lint tools across supported platforms (kunchenguid#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers

* docs: reconcile test-evidence docs with store_in_repo: true (kunchenguid#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.

* docs: clarify test evidence branch storage (kunchenguid#2549)

* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes

* feat: add worker auto-compaction launch config

* no-mistakes(document): docs: add crew-autocompact to AGENTS.md config file inventory

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
#2)

* fix: keep the public promise reachable when work is routed to a second mate (kunchenguid#2457)

The lightweight Relay follow-up link lives in the answering home's own
state/<task-id>.meta, so it can only bind work that home owns. When a
Relay-linked request is routed to a second mate, the task record lives in the
second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and
nothing else picked the promise up: only the soft acknowledgement was ever
posted. The typed promised-final path already supports --work-home
secondmate:<id>; the playbook simply never chose it.

- fmx-respond now states the routing rule crisply: a task in this home takes the
  lightweight link, and second-mate-routed work takes a promised-final
  commitment bound to that home, registered up front with the brief command
  carried into the routed worker's instructions.
- fm-x-link.sh refuses a task with no local record by naming the registered
  second mate whose home actually holds it and printing the promised-final
  registration command, with the exact --work-home when the match is
  unambiguous. A home with no registered second mates keeps the plain error.
- fm-backlog-handoff.sh reports, after a successful move, any moved key that
  still owes a public reply bound to main/<key>, since that binding no longer
  names the home owning the work. The move itself is never blocked.

Docs and the secondmate handoff prose follow the same rule. Tests cover the
refusal, its scoping, the unchanged local-link path, and both handoff outcomes
at the script boundary.

* docs(skills): add remote-secondmate recovery hint for false-negative verdicts (kunchenguid#2456)

* fix(skills): hint that remote secondmate liveness verdicts false-negative

fm-crew-state and fm-send routinely misreport a live remote secondmate
as dead; confirm against the pane before relaunching, and relaunch only
through fm-spawn.sh, never raw herdr pane surgery.

* no-mistakes: apply CI fixes

* fix(calm): keep Pi's export confirmation visible (kunchenguid#2461)

Pi 0.83.0 added a status line to every tool-expansion change, and Pi
updates the previous status line in place when two status messages
arrive back to back. Calm's post-export redraw cycled tool expansion on
the macrotask right after Pi printed "Session exported to: <path>", so
both expansion status lines coalesced over that confirmation and the
captain was left with no record of where their export landed.

Calm now repaints only the tool rows it presents, by invalidating each
row through the render context Pi hands its render slots, and requests
the surrounding redraw through setStatus. Neither appends to the
transcript. The repaint is still needed because Pi can re-render a row
asynchronously - the built-in edit row invalidates itself once its diff
is ready - and that re-render can land inside the window where /export
forces stock rendering.

The real-terminal /export case now asserts the confirmation is still on
screen after the redraw has settled, and that the redraw restored every
Calm-hidden row, instead of only racing the moment the confirmation
first appeared.

* feat(stow): add open-record persistence to /stow before reset (kunchenguid#2488)

* feat(stow): persist the open records a session is holding

/stow curated memory and captured session knowledge, but never touched
record state, while AGENTS.md called it an "unfinished-work sweep" and the
receipt declared the session "safe to reset" - wording that implied a
record-correctness guarantee stow does not make. A shipped PR with no
backlog item, a queued umbrella whose phases had merged, and four decision
holds left open after their answers shipped all survived repeated stows.

Add a bounded pass that files record state from the same volatile input the
rest of stow already uses: the open threads in context, minutes before the
reset destroys them. It creates a record for an unfiled thread and corrects
one the session knows is wrong, through the owning path, and states its
boundary as part of the contract - it never enumerates the backlog, lists
holds, or queries a forge, because it cannot be a reconciliation and must
not be read as one.

Correct the wording in AGENTS.md and the completion receipt so reset-safe
means what it actually guarantees: nothing this session knew was lost.

* no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi

* no-mistakes(document): note /stow open-record persistence in README command catalog

* refactor(stow): state open-record persistence as principle, not procedure

The first version enumerated triggers, named commands, and prescribed an
ordered procedure. That is too rigid for an agent skill: it invites literal
execution of a checklist instead of judgment, and every enumerated example
is a way for the guidance to go stale.

Reduce it to the intent - before a reset, the important open work you are
holding in context must end up durably recorded rather than dying with the
session, filing what is unfiled and correcting what is stale - and let the
agent judge importance, the record, and the owning write path.

Keep the scope bound, since it is a decided contract and not a mechanic:
this covers the open work the session is holding, never a reconciliation of
durable records against repository or forge reality. The wording
corrections in AGENTS.md and the completion receipt are unchanged.

* fix(decisions): close decision holds at answer time via one general keyed-answer path (kunchenguid#2490)

* fix(decisions): close captain holds at answer time

Firstmate had two "a decision is open" ledgers with asymmetric closing
mechanics. The live status-log ledger closes atomically at answer time,
because bin/fm-send.sh --resolve-key makes answering a decision be the
act that closes it. The durable backlog hold ledger had no such coupling:
answering and recording were two separate acts, and only the first was
forced by the workflow.

That asymmetry lost four real captain decisions. Their answers were
captured durably to disk, keyed character for character by the hold
decision keys, acknowledged, and even implemented and shipped, yet the
holds stayed open for two days and the captain was asked to re-answer
decisions already on his own disk.

Give the hold ledger the same answer-time-closure property:

- bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's
  counterpart to --resolve-key. It shares one unrouted close
  implementation with `decline`, so it carries every existing guard - the
  captain decision file, the active-hold requirement, retry identity, and
  the refusal to release still-routed work - and differs only in the
  resolution mode it records. `decline` keeps its stronger meaning that
  the answer routes no follow-up work at all.
- bin/fm-procevent-lavish.sh wires the channel that actually carried the
  lost answers. `arm --decisions-origin` binds a deck to the origin whose
  holds it carries, `answers` reads the structured choices out of a
  captured poll result, `close-decisions` maps each key to its hold and
  closes it through the command above, and `autohandle` lets the runner
  apply that at capture time.

Safety is preserved rather than traded away. Only rows tagged `choice`
are read, so freeform captain prose cannot forge a decision key. Closure
is confined to the one bound origin. The decision text is a pure function
of the captured result, so a replayed capture is idempotent. A hold that
is absent, already closed, or still blocking routed work is skipped and
left for `resolve`, never forced. A deck armed without the binding
touches no hold at all. And autohandle deliberately never reports full
handling, because recording an answer is transcription while acting on it
is firstmate's judgement - so the check wake still reaches the handler.

fm-send --resolve-key is untouched.

* no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory

* refactor(decisions): make keyed-answer closure one general capability

The previous pass gave holds answer-time closure but built it as bespoke
Lavish wiring: the review adapter carried the source-to-origin binding,
mapped keys to hold identities, wrote decision records, decided what to
skip, and closed holds itself. That treated a review deck as a special
decision source. It is not - it is an ephemeral discussion format that
happens to carry answers.

Collapse it into ONE general capability with one owner.

bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its
matching hold":
- `answers <origin> --source <provenance>` is the channel-agnostic
  intake. It reads key/answer/label lines on stdin, maps each key to its
  hold, and closes it through the same `answer` path, so every guard
  applies identically whatever channel the answer came from. --source is
  provenance recorded in the decision, never a behavior switch; there is
  no per-channel branch and no knowledge of chat, decks, or transports.
- `bind`/`unbind`/`binding` own the source-to-origin binding for any
  channel whose answers arrive detached from their origin.

Every channel is now an ordinary caller that only turns what it received
into keyed lines:
- bin/fm-send.sh (chat) feeds the intake for a key that names an active
  hold. This also fixes a real gap: once `complete` transfers a decision
  to its hold it closes the live status copy, so --resolve-key alone
  could never answer a transferred decision.
- bin/fm-procevent.sh feeds it generically. A bound source's captured
  result goes to `<adapter> answers <result-file>` and whatever that
  prints is piped into the intake. The runner names no adapter, parses
  no result, and carries no decision rule, so any future adapter with an
  `answers` command works with no change here.
- bin/fm-procevent-lavish.sh keeps only `answers`, which reports the
  structured choices a review captured and stops. It maps nothing to a
  hold and closes nothing; it lost ~160 lines of decision logic.

Feeding is independent of handling, so it never acknowledges a result
and never suppresses a wake - recording an answer is transcription,
acting on it stays firstmate's judgement.

The regression that proves closure now drives a FIXTURE adapter that is
not the review adapter, so what is proven is that any bound channel
reaches the intake rather than that one channel is wired specially. A
new regression drives the real fm-send over a stubbed transport for the
chat side. Every prior guarantee still holds, and fm-send's status-log
behavior is unchanged.

* no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression

* fix(memory): emit a real @AGENTS.md pointer instead of a CLAUDE.md symlink (kunchenguid#2512)

A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md.
The installer now creates and migrates to a recoverable two-line pointer file.

* fix: tighten generated brief evidence discipline

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…policy) (#3)

* feat(herdr): add per-project workspaces

* no-mistakes(review): preserve journal recovery guard; path-hash project bindings

* no-mistakes(document): add herdr-project-spaces to secondmate inherited-config lists

* no-mistakes(review): run presentation-journal recovery guard whenever a journal exists

* no-mistakes(test): register live agent so guard e2e refusal check fires

* no-mistakes(document): sync architecture herdr layout summary; fix stale section pointers

* no-mistakes(document): document herdr project-space state binding in AGENTS inventory

* fix(herdr): consume project workspace outputs
* Fix treehouse teardown path and slot rebind safety

* Harden treehouse teardown occupancy safety

* Lock pooled relaunch slot adoption

* Restore slot owner after aborted relaunch

* Make Treehouse test doubles capability-explicit

* Fix Zellij teardown Treehouse fixtures
@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown

Confidence Score: 4/5

The PR appears safe to merge, with one non-blocking LAN availability-hardening issue in the HTTP server.

The implemented data and answer flows have no accepted correctness blocker, but the LAN listener permits unbounded pre-authentication request threads and can be made unavailable by sufficiently many concurrent connections.

Files Needing Attention: bin/fm-board.py

Security Review

The LAN server bounds individual socket lifetimes but does not bound concurrent pre-authentication connections, allowing a LAN client to exhaust server threads and availability. How this was verified: The external listener uses ThreadingHTTPServer without admission control, while authentication runs only inside each allocated request thread.

Reviews (1): Last reviewed commit: "no-mistakes(document): Document board-an..." | Re-trigger Greptile

Comment thread bin/fm-board.py
Comment on lines +997 to +998
with contextlib.suppress(ProcessLookupError):
os.killpg(self.runner.pid, signal.SIGTERM)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 security Unbounded LAN request threads

If a LAN client opens many concurrent slow or idle connections, ThreadingHTTPServer allocates an unrestricted thread per connection before authentication; the 20-second socket timeout limits duration but not concurrency, so sustained connections can exhaust process resources and make the dashboard unavailable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant