Skip to content

fix(bearings): support closing and dropping captain decisions - #8

Merged
prandelicious merged 31 commits into
mainfrom
fm/firstmate-upstream-sync-r2
Aug 22, 2026
Merged

prandelicious merged 31 commits into
mainfrom
fm/firstmate-upstream-sync-r2

Conversation

@prandelicious

@prandelicious prandelicious commented Aug 22, 2026 •

Copy link
Copy Markdown
Owner

Intent

Synchronize prandelicious/firstmate with kunchenguid/firstmate upstream main on branch fm/firstmate-upstream-sync-r2. Preserve the fork's seven unique commits and intended public-skill behavior while incorporating upstream's current commits, resolving the listed conflicts semantically rather than choosing one side blindly. Keep upstream's newer fleet-board decision-card close/drop (drop) and always-on freeform behavior when it is the complete contract, while retaining fork-only intent not covered upstream. Exclude the parked firstmate-guidelines work. Validate every conflict resolution with relevant tests and run the repository's full required validation. Deliver a green PR against prandelicious/firstmate main only; never target kunchenguid/firstmate and never merge the PR. Preserve all prior no-mistakes pipeline-fix commits. The approved CI recovery adds workflow_dispatch and explicit pull_request activity types to .github/workflows/ci.yml; Actions is enabled, PR #8 is open, and the fresh CI run must register and pass on the current branch head before the pipeline returns checks-passed.

What Changed

  • Add Close / drop controls and always-on supplementary freeform input to Bearings decision cards.
  • Handle reserved __drop__ answers by declining the matching hold while preserving dependent queued work, with regression coverage and updated lifecycle guidance.
  • Require authenticated no-mistakes evidence for contributor PRs and restore explicit CI triggers for PR activity and manual dispatch.

Risk Assessment

🚨 High: The required compliance gate can accept stale pipeline evidence and directly contradicts the intent to preserve the prior current-head attestation fix.

Testing

Inspected the target delta, then exercised the executable Bearings board and decision-hold lifecycle end-to-end: strict decision choices, freeform-only credential cards, binding/arming order, Close/drop through captured feedback, durable declined resolution, Captain’s Call removal, idempotent replay, routed-work release, and public decline guards all passed; transcripts were captured and the worktree remained clean.

Evidence: Bearings board end-to-end transcript

Source: Bearings board end-to-end transcript

FM_TEST_BEGIN 2026-08-22T06:12:26Z tests/fm-bearings-board.test.sh family=pure-contract-unit expected_gate_skip=none
ok - path prints the stable home-scoped board location
ok - build refuses malformed payloads before touching the board
ok - build keeps freeform-only credential cards valid
ok - build injects the payload, binds any-origin, then arms the source
not-autohandled: lavish-05085ce67434b1cf (left for the handler; still unacknowledged)
ok - registration can consume answers only after any-origin binding exists
ok - build establishes the Lavish session before binding and arming
ok - rebuild refreshes the board in place without double-arming
ok - build refuses a template without exactly one data slot
ok - a reserved close/drop answer declines the hold and leaves Captain's Call
FM_TEST_END 2026-08-22T06:12:37Z tests/fm-bearings-board.test.sh exit=0 duration_ms=10519 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=10588
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=10519 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-bearings-board.test.sh duration_ms=10519
Evidence: Decision-hold lifecycle transcript

Source: Decision-hold lifecycle transcript

FM_TEST_BEGIN 2026-08-22T06:12:41Z tests/fm-decision-hold-lifecycle.test.sh family=pure-contract-unit expected_gate_skip=none
ok - report-only unresolved decision is reproduced and completion refuses before loss
ok - non-forced scout teardown always requires durable inventory verification
ok - a declined decision closes with a recorded answer and no routed work
ok - a decision closed outside the script is repairable and then clears teardown
ok - an unanswered decision still blocks completion and resists both unrouted close paths
ok - captain holds are idempotent, distinct, teardown-safe, Bearings-visible, and durably routed before close
ok - completion and verification validate origins before constructing paths
ok - ended visual review follows the same decision-hold completion owner
ok - resolved findings and decision-like prose do not create false holds
ok - terminal single-owner stale status decisions do not block empty inventory
ok - main-home and secondmate-home captain holds remain correctly routed
ok - resolve matches first/middle/last in quoted blocked_by and rejects a genuinely absent id
ok - a bound channel's captured answers close their captain holds at answer time
ok - a channel source with no decision binding closes nothing
ok - an any-origin bound source closes full-identity holds across origins
ok - the answer path keeps every guard the unrouted close path already had
ok - the chat channel feeds the same keyed-answer intake a captured review does
FM_TEST_END 2026-08-22T06:14:17Z tests/fm-decision-hold-lifecycle.test.sh exit=0 duration_ms=95942 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=96011
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=95942 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-decision-hold-lifecycle.test.sh duration_ms=95942

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 error
  • 🚨 .github/workflows/no-mistakes-required.yml:52 - Intent requires “Preserve all prior no-mistakes pipeline-fix commits,” but this lookup regresses 90ea8ed’s current-head binding: it accepts the newest evidence commit whose message names the branch without reading the evidence record or comparing its attested head SHA to github.event.pull_request.head.sha. After a subsequent push, stale evidence from an earlier successful run remains reachable and the compliance check can pass without validating the current PR head. Restore head-SHA validation at this evidence-verification boundary.
✅ **Test** - passed

✅ No issues found.

  • git diff --stat 166111ccbe102b0f9f3b79de2df4e9f1c7103c57..f82aefecc0cb8d8bc479392ec9ec1a79ce6036b3 and targeted diff inspection
  • bin/fm-test-run.sh tests/fm-bearings-board.test.sh
  • bin/fm-test-run.sh tests/fm-decision-hold-lifecycle.test.sh
  • git status --short
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 29 commits August 17, 2026 13:34
* fix(lint): catch malformed GitHub workflows before merge

A self-broken ci.yml cannot report its own breakage, so parse every
workflow in the local lint path that no-mistakes already runs.

* fix(lint): pin actionlint instead of Ruby for workflow lint

A self-broken ci.yml still has to fail in the local lint path, and the
named tool for that gate is actionlint, not a new Ruby runtime.

* no-mistakes(document): Clarify pinned workflow lint documentation
…d#2546)

* fix: install pinned shellcheck and actionlint on macOS and linux arm64

The installers were hardcoded to linux amd64 and sha256sum, so a Mac
dev could not satisfy the refuse-on-mismatch lint gate. Select the
official per-platform archive and checksum, and fall back to shasum -a 256.

* no-mistakes(document): Document cross-platform pinned lint installers
…uid#2548)

.no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but
CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described
the old policy of keeping evidence out of the repo in a temp directory.

The current no-mistakes behavior for store_in_repo: true is to publish each run's
test evidence to the orphan no-mistakes/evidence branch and link it from the PR
body. That branch shares no history with code branches, so evidence never enters
a pushed feature branch or the default branch, and CI's tracked personal fleet
paths rule stays accurate.

Docs only. No change to .no-mistakes.yaml or any workflow.
* docs: correct test evidence storage comment in .no-mistakes.yaml

* no-mistakes: apply CI fixes
…nchenguid#2563)

Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.
…chenguid#2570)

* fix(bin): report remote secondmate delivery and state truthfully

A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send
leg whose unconfirmed submit read-back (verdict=pending, typically a busy
mate whose harness queues the steer) was flattened into exit 1, so the
parent printed "error: text not submitted" / "error: text not sent" and
discarded the pending-reply expectation for a steer that had actually
landed. fm-send now carries the verdict across the ssh boundary as a
documented delivered-unconfirmed exit 3: the parent reports the steer as
delivered with confirmation pending, exits 0, keeps the expectation armed
(awaiting_report), and closes --resolve-key decisions, while transport
loss (ssh 255) and real remote failures keep failing loudly with the
remote leg's stderr attached. A local unconfirmed submit now also exits 3
with an honest non-error message and still never closes a decision key.

fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint
through local probes (which misreported a healthy mate as "worktree gone"
/ "can't find session: remote"): both now use the true remote source over
fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote,
never as gone or dead.

* no-mistakes(document): Document remote delivery and state truth

* no-mistakes: apply CI fixes
* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch.

quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake.

* no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts

* no-mistakes(document): Correct quota verification evidence chronology

* Collapse quota-array-dispatch onto TOON-first spendPriority ranking.

Decide from quota-axi's default TOON; keep --json as a rare defensive fallback.
Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers.

* no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures

* no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty

* no-mistakes(document): Document TOON-first quota dispatch evidence
* docs: add GROK_BOT.md Grok Bot system prompt

* docs: amend GROK_BOT.md with charter report-back and delegation marker

* docs: classify GROK_BOT.md as public-product

* docs: make GROK_BOT.md the plain Grok Bot system prompt
Refine language for clarity and consistency in instructions.
…#2595)

* fix(bin): guarantee inactive-reconcile scan progress under second quantization

The inactive-outcome scan computed its aggregate deadline in whole seconds,
so a 1-second budget's effective value lands anywhere in (0,1]; a scan
starting just before a wall-clock second boundary rounded its whole budget
away mid-scan and exited having visited no child, while the durable cursor
had already advanced past the never-examined child. This is the CI flake
behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not
resume with the following child' (watcher-wake-lock family, portable
serial 2, seen on the PR kunchenguid#2590 run).

Every scan now visits at least its first due child with the per-child
state-read bound floored at one second, so no invocation can be a zero-work
no-op. The outer process-group kill moves to budget+1s: the scan's own
deadline enforces the budget, and the kill is a backstop for a scan wedged
in an unbounded wait instead of a racer that routinely preempts the clean
bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s);
the previously flaky assertion is unchanged.

* no-mistakes(document): Document inactive-reconcile deadline backstop
Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.
Clarified guidelines for handing off work to crewmates and managing secrets.
…tions deterministic (kunchenguid#2617)

Three assertions in tests/fm-procevent.test.sh depended on a detached runner
having finished work that the command starting it does not wait for.

reconcile's replacement runner is started through detach_runner, which only
forks: reconcile returns and counts the start before that runner has claimed
its source or exec'd its child. Any assertion taken straight after reconcile
therefore samples a race.

- The publish-before-apply recovery section left its always-ready /bin/echo
  source registered across the recovery reconcile, so that reconcile launched
  a competing detached poll (observed: started=1) that then raced every later
  assertion for the source claim, the next capture sequence, and this home's
  applied record, and outlived the section holding a live claim. It is now
  retired before that reconcile - re-announcement is proven from the durable
  inbox alone and needs no registration - and started=0 is asserted so a
  competing poll cannot be reintroduced unnoticed. This is the same
  retire-before-reconcile discipline the self-announcing section already
  carries; that section acquired it after the identical race made its
  "not-autohandled: self-src" assertion read "already owned: self-src".

- The crashed-leader replacement section snapshotted the replacement's claim
  file and execution log behind a fixed 0.5s settle window. On a loaded
  machine that window expires first, which is the CI flake behind "a
  replacement runner started without recording its own claim" and "reconcile
  did not start exactly one replacement source". Both effects are now waited
  for with the suite's bounded wait helpers; the exact one-replacement count
  is still asserted afterwards, unchanged.

- The duplicate-start section slept 0.5s for reconcile's runner to record
  ownership before asserting that a second start loses to it. It now waits
  for that claim.

Also tighten one assertion that could not fail as written: "autohandled:
self-src" is a substring of "not-autohandled: self-src", so the applied path
was accepted even when the runner reported the capture left for the handler.

Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced
6 failing runs, all in the crashed-leader section. On the fixed suite, 216
full runs under the same load produced none. Reverting the self-announcing
section's retire-before-reconcile line reproduces "already owned: self-src"
on the first iteration, confirming the shared mechanism.
)

* fix(bin): keep pending-reply expectations honest on both send legs

Two related asymmetries let the parent-owned secondmate reply guard drop or
nag requests it should not have.

Local delivered-unconfirmed dropped the expectation. A marked request whose
submit read-back stayed unconfirmed (verdict=pending) is the same
not-a-failure outcome the remote leg reports as delivered, but fm-send
discarded the parent's pending-reply record for it, so a request that very
likely landed stopped being tracked entirely. The record now stays armed on
its unconfirmed-delivery marker: a correlated report still resolves it, and
an unanswered one still surfaces through the library's own reconciliation.
Exit 3 and the local rule that an unconfirmed answer never closes a decision
key are unchanged.

Remote replies were nagged for a repost they did not need. A remote mate's
report reaches the parent's status log only through the asynchronous mirror
in fm-procevent-remote-reply.sh, yet the guard read an absent correlated
line as proof the mate never reported - even while the answer was still in
flight, which is the common case because the mirror's poll window is
comparable to the recovery grace. The mirror now publishes one caught-up
watermark from a quiet window, and the guard admits a missing report as
evidence only once that watermark passes the turn that should have produced
it. A genuinely missed report still gets exactly one repost, and a channel
that is behind, unarmed, or broken leaves the request durably open and
un-nagged rather than nagging blind; the mirror escalates its own continuity
failures as before.

Tests: a local unconfirmed secondmate send keeps its expectation armed and
resolvable; a mirrored correlated remote reply resolves with no repost; a
stale or absent watermark withholds the repost while a fresh one still
releases it; a quiet remote window publishes the watermark and retirement
clears it.

* no-mistakes(review): Distinguish preempted polls from quiet windows

* no-mistakes(document): Clarify remote reply channel freshness

* no-mistakes(lint): Annotate shared remote preemption exit constant
…d#2619)

* fix(watch): honor a declared pause on a busy pane's completed-turn bound

A worker that declares an external wait (`paused:`) and then blocks in one
long foreground call - a review-hosting scout parked in a single blocking
`lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane
BUSY, so the stale path that already honors declared pauses never ran for it.
The busy-pane completed-turn bound instead routed it straight into
wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past
the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long
as the review stayed open.

busy_turn_bound_check now owns which absorber takes a crossed bound: a crew
whose own last status line declares an external wait or a verified captain-held
transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else
keeps the unchanged wedge timer. The discriminator is the declaration together
with liveness (the caller has already confirmed the pane is busy), never a
blanket silencing - a crew that declared nothing, or whose pane is not live,
escalates exactly as before, and a declared pause still re-surfaces once per
long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched:
the daemon owns pause triage there and already reads the same vocabulary.

The two call sites also no longer clear pause bookkeeping in the same poll the
pause cadence recorded it, which would have erased the re-surface throttle and
turned the long cadence back into a per-poll re-surface.

Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence
recheck, and the restored wedge escalation once the declaration is lifted on the
same busy over-age pane.

Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded
machine: fixed liveness budgets were reaping watchers mid-startup, so assertions
on post-poll state passed vacuously or failed spuriously. Waits that describe a
poll's outcome now wait for a completed poll cycle via the liveness beacon, the
heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit
budget is the uniform 10s already used elsewhere in the file.

* no-mistakes(review): Fail poll-cycle waits on timeout

* no-mistakes(review): Prevent poll timeout test hangs

* no-mistakes(document): Clarify paused busy-pane supervision
Added guidelines for decision communication to the captain.
Clarify communication protocols with crewmates regarding task delegation and reporting.
* fix(herdr): confirm local steers that native agent-state misses

Herdr can leave agent_status idle for a landed Claude turn and can keep
queued Enter text visible while busy, so fm-send was reporting false
swallows. Confirm those cases through the shared queued-Enter verdict
and a cleared composer, and keep a genuine idle pending composer as
unconfirmed.

* no-mistakes(review): Stop Herdr Enter retries on unreadable composers

* no-mistakes(review): Reject queued delivery when all Herdr Enter sends fail

* no-mistakes(review): Prevent confirmation after failed Herdr Enter

* no-mistakes(review): Pace Herdr retries and clarify submit fallback

* no-mistakes(review): Align Herdr submit docs with idle fallback

* no-mistakes(document): Correct Herdr submit-confirmation documentation
* feat(bin): accept any-origin decision bindings with full-identity keys

An aggregation surface (the bearings board) carries captain answers for holds
across origins, but a binding was one-origin-per-source and the Lavish adapter
capped question keys at 64 chars while real full hold identities measure 69-81.

- fm-decision-hold.sh: bind <source-id> --any-origin records the (any) marker;
  binding prints it verbatim and answers accepts it, so the runner's feed seam
  carries an any-origin source with no runner change. In any-origin mode each
  key is a full hold identity <origin>-decision-<key>, split at its first
  -decision-; a key with no separator (merge/dispatch instructions) is skipped
  and feeds nothing, keeping non-decision answers out of the hold ledger by
  construction. Every existing close guard applies unchanged.
- fm-procevent-lavish.sh: raise the question-key cap 64 -> 128 so a full hold
  identity fits; the slug-shape security property is unchanged.
- tests: cross-origin closure through the real runner seam, an 81-char
  identity through the adapter, cap and shape refusals, routed-work skips,
  nonexistent-identity skips, and idempotent replay.

* feat(bearings): add the /bearings lavish interactive fleet board

/bearings lavish renders the bearings snapshot onto a shipped, reusable board
template and arms it as a Lavish process-event source, so the captain answers
Captain's Call items on the board and firstmate is woken by an ordinary check
wake - no conversational turn ever blocks on a poll.

- .agents/skills/bearings/assets/board-template.html: the shipped template
  (myfirstmate design system inlined, one fm-bearings-board.v1 JSON slot,
  fail-closed schema guard that renders an error card instead of an empty
  fleet). Per-invocation agent work is composing the payload only.
- bin/fm-bearings-board.sh: build/refresh owner - fail-closed payload
  validation, slot injection with a round-trip check and \u003c escaping,
  stable board path, any-origin bind ALWAYS before arm, arm-if-absent.
- bearings SKILL.md: the lavish invocation option, board composition rules,
  board-wake handling, and the captain-ruled merge-click authorization with
  its mandatory safeguards (PR resolved from the task's own meta record,
  wake-time green re-verification, never a red or changed PR, merges only
  through bin/fm-pr-merge.sh, chat echo with the full PR URL).
- process-event-sources SKILL.md: one-line board-wake routing trigger.
- tests: payload refusals, injection round-trip, bind-before-arm, idempotent
  re-arm, and template slot integrity.

Fleet pickup: homes receive this after merge plus a firstmate self-update;
landing timing is coordinated with the main firstmate.

* no-mistakes(review): Harden bearings board validation and wake handling

* no-mistakes(review): Require HTTPS for bearings board PR links

* no-mistakes(review): Fail closed and bound bearings board answers

* no-mistakes(review): Enforce UTF-8 byte limits for board answers

* no-mistakes(review): Serve bearings board before arming and reject empty actions

* no-mistakes(review): Prove bind-before-arm ordering through live answer consumption

* no-mistakes(document): Document bearings board and cross-origin answers
…enguid#2707)

* fix(bearings): always show decision options and a close/drop control

Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call.

* no-mistakes(review): Fix drop closure and decision-only option validation

* no-mistakes(review): Preserve answerability for non-decision cards

* no-mistakes(document): Clarify decision drop documentation
Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.
…am-sync-r2

# Conflicts:
#	.agents/skills/bearings/SKILL.md
#	.agents/skills/bearings/assets/board-template.html
#	bin/fm-bearings-board.sh
#	bin/fm-decision-hold.sh
#	bin/fm-test-run.sh
#	docs/decision-hold-lifecycle.md
#	tests/fm-bearings-board.test.sh
@prandelicious prandelicious changed the title fix(bearings): support closing and dropping captain decisions fix(bearings): support closing and dropping captain decisions [CI trigger] Aug 22, 2026
@prandelicious prandelicious changed the title fix(bearings): support closing and dropping captain decisions [CI trigger] fix(bearings): support closing and dropping captain decisions Aug 22, 2026
@prandelicious prandelicious changed the title fix(bearings): support closing and dropping captain decisions fix(bearings): support closing and dropping captain decisions [CI refresh] Aug 22, 2026
@prandelicious prandelicious changed the title fix(bearings): support closing and dropping captain decisions [CI refresh] fix(bearings): support closing and dropping captain decisions Aug 22, 2026
@prandelicious
prandelicious merged commit 4db874e into main Aug 22, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants