Merge upstream/main into fork integration line (catch-up 4, 107 commits) - #72
Merged
Merged
Conversation
A wedged family-run step was occupying the runner until the 75-minute job cap; bound that step so cleanup and timing artifacts still upload.
…d mate (kunchenguid#2457) The lightweight Relay follow-up link lives in the answering home's own state/<task-id>.meta, so it can only bind work that home owns. When a Relay-linked request is routed to a second mate, the task record lives in the second mate's home, fm-x-link.sh failed with a bare "no such task ...meta", and nothing else picked the promise up: only the soft acknowledgement was ever posted. The typed promised-final path already supports --work-home secondmate:<id>; the playbook simply never chose it. - fmx-respond now states the routing rule crisply: a task in this home takes the lightweight link, and second-mate-routed work takes a promised-final commitment bound to that home, registered up front with the brief command carried into the routed worker's instructions. - fm-x-link.sh refuses a task with no local record by naming the registered second mate whose home actually holds it and printing the promised-final registration command, with the exact --work-home when the match is unambiguous. A home with no registered second mates keeps the plain error. - fm-backlog-handoff.sh reports, after a successful move, any moved key that still owes a public reply bound to main/<key>, since that binding no longer names the home owning the work. The move itself is never blocked. Docs and the secondmate handoff prose follow the same rule. Tests cover the refusal, its scoping, the unchanged local-link path, and both handoff outcomes at the script boundary.
…verdicts (kunchenguid#2456) * fix(skills): hint that remote secondmate liveness verdicts false-negative fm-crew-state and fm-send routinely misreport a live remote secondmate as dead; confirm against the pane before relaunching, and relaunch only through fm-spawn.sh, never raw herdr pane surgery. * no-mistakes: apply CI fixes
Pi 0.83.0 added a status line to every tool-expansion change, and Pi updates the previous status line in place when two status messages arrive back to back. Calm's post-export redraw cycled tool expansion on the macrotask right after Pi printed "Session exported to: <path>", so both expansion status lines coalesced over that confirmation and the captain was left with no record of where their export landed. Calm now repaints only the tool rows it presents, by invalidating each row through the render context Pi hands its render slots, and requests the surrounding redraw through setStatus. Neither appends to the transcript. The repaint is still needed because Pi can re-render a row asynchronously - the built-in edit row invalidates itself once its diff is ready - and that re-render can land inside the window where /export forces stock rendering. The real-terminal /export case now asserts the confirmation is still on screen after the redraw has settled, and that the redraw restored every Calm-hidden row, instead of only racing the moment the confirmation first appeared.
…nguid#2488) * feat(stow): persist the open records a session is holding /stow curated memory and captured session knowledge, but never touched record state, while AGENTS.md called it an "unfinished-work sweep" and the receipt declared the session "safe to reset" - wording that implied a record-correctness guarantee stow does not make. A shipped PR with no backlog item, a queued umbrella whose phases had merged, and four decision holds left open after their answers shipped all survived repeated stows. Add a bounded pass that files record state from the same volatile input the rest of stow already uses: the open threads in context, minutes before the reset destroys them. It creates a record for an unfiled thread and corrects one the session knows is wrong, through the owning path, and states its boundary as part of the contract - it never enumerates the backlog, lists holds, or queries a forge, because it cannot be a reconciliation and must not be read as one. Correct the wording in AGENTS.md and the completion receipt so reset-safe means what it actually guarantees: nothing this session knew was lost. * no-mistakes(review): correct stow decision-hold inspection to read hold via tasks-axi * no-mistakes(document): note /stow open-record persistence in README command catalog * refactor(stow): state open-record persistence as principle, not procedure The first version enumerated triggers, named commands, and prescribed an ordered procedure. That is too rigid for an agent skill: it invites literal execution of a checklist instead of judgment, and every enumerated example is a way for the guidance to go stale. Reduce it to the intent - before a reset, the important open work you are holding in context must end up durably recorded rather than dying with the session, filing what is unfiled and correcting what is stale - and let the agent judge importance, the record, and the owning write path. Keep the scope bound, since it is a decided contract and not a mechanic: this covers the open work the session is holding, never a reconciliation of durable records against repository or forge reality. The wording corrections in AGENTS.md and the completion receipt are unchanged.
…eyed-answer path (kunchenguid#2490) * fix(decisions): close captain holds at answer time Firstmate had two "a decision is open" ledgers with asymmetric closing mechanics. The live status-log ledger closes atomically at answer time, because bin/fm-send.sh --resolve-key makes answering a decision be the act that closes it. The durable backlog hold ledger had no such coupling: answering and recording were two separate acts, and only the first was forced by the workflow. That asymmetry lost four real captain decisions. Their answers were captured durably to disk, keyed character for character by the hold decision keys, acknowledged, and even implemented and shipped, yet the holds stayed open for two days and the captain was asked to re-answer decisions already on his own disk. Give the hold ledger the same answer-time-closure property: - bin/fm-decision-hold.sh gains an `answer` subcommand, the hold ledger's counterpart to --resolve-key. It shares one unrouted close implementation with `decline`, so it carries every existing guard - the captain decision file, the active-hold requirement, retry identity, and the refusal to release still-routed work - and differs only in the resolution mode it records. `decline` keeps its stronger meaning that the answer routes no follow-up work at all. - bin/fm-procevent-lavish.sh wires the channel that actually carried the lost answers. `arm --decisions-origin` binds a deck to the origin whose holds it carries, `answers` reads the structured choices out of a captured poll result, `close-decisions` maps each key to its hold and closes it through the command above, and `autohandle` lets the runner apply that at capture time. Safety is preserved rather than traded away. Only rows tagged `choice` are read, so freeform captain prose cannot forge a decision key. Closure is confined to the one bound origin. The decision text is a pure function of the captured result, so a replayed capture is idempotent. A hold that is absent, already closed, or still blocking routed work is skipped and left for `resolve`, never forced. A deck armed without the binding touches no hold at all. And autohandle deliberately never reports full handling, because recording an answer is transcription while acting on it is firstmate's judgement - so the check wake still reaches the handler. fm-send --resolve-key is untouched. * no-mistakes(document): document state/lavish-decisions binding dir in AGENTS.md state inventory * refactor(decisions): make keyed-answer closure one general capability The previous pass gave holds answer-time closure but built it as bespoke Lavish wiring: the review adapter carried the source-to-origin binding, mapped keys to hold identities, wrote decision records, decided what to skip, and closed holds itself. That treated a review deck as a special decision source. It is not - it is an ephemeral discussion format that happens to carry answers. Collapse it into ONE general capability with one owner. bin/fm-decision-hold.sh now owns the whole of "a keyed answer closes its matching hold": - `answers <origin> --source <provenance>` is the channel-agnostic intake. It reads key/answer/label lines on stdin, maps each key to its hold, and closes it through the same `answer` path, so every guard applies identically whatever channel the answer came from. --source is provenance recorded in the decision, never a behavior switch; there is no per-channel branch and no knowledge of chat, decks, or transports. - `bind`/`unbind`/`binding` own the source-to-origin binding for any channel whose answers arrive detached from their origin. Every channel is now an ordinary caller that only turns what it received into keyed lines: - bin/fm-send.sh (chat) feeds the intake for a key that names an active hold. This also fixes a real gap: once `complete` transfers a decision to its hold it closes the live status copy, so --resolve-key alone could never answer a transferred decision. - bin/fm-procevent.sh feeds it generically. A bound source's captured result goes to `<adapter> answers <result-file>` and whatever that prints is piped into the intake. The runner names no adapter, parses no result, and carries no decision rule, so any future adapter with an `answers` command works with no change here. - bin/fm-procevent-lavish.sh keeps only `answers`, which reports the structured choices a review captured and stops. It maps nothing to a hold and closes nothing; it lost ~160 lines of decision logic. Feeding is independent of handling, so it never acknowledges a result and never suppresses a wake - recording an answer is transcription, acting on it stays firstmate's judgement. The regression that proves closure now drives a FIXTURE adapter that is not the review adapter, so what is proven is that any bound channel reaches the intake rather than that one channel is wired specially. A new regression drives the real fm-send over a stubbed transport for the chat side. Every prior guarantee still holds, and fm-send's status-log behavior is unchanged. * no-mistakes(review): test(decisions): drop source-content grep from hold-closure regression
…mlink (kunchenguid#2512) A Write aimed at CLAUDE.md followed the symlink and destroyed AGENTS.md. The installer now creates and migrates to a recoverable two-line pointer file.
* fix(lint): catch malformed GitHub workflows before merge A self-broken ci.yml cannot report its own breakage, so parse every workflow in the local lint path that no-mistakes already runs. * fix(lint): pin actionlint instead of Ruby for workflow lint A self-broken ci.yml still has to fail in the local lint path, and the named tool for that gate is actionlint, not a new Ruby runtime. * no-mistakes(document): Clarify pinned workflow lint documentation
…d#2546) * fix: install pinned shellcheck and actionlint on macOS and linux arm64 The installers were hardcoded to linux amd64 and sha256sum, so a Mac dev could not satisfy the refuse-on-mismatch lint gate. Select the official per-platform archive and checksum, and fall back to shasum -a 256. * no-mistakes(document): Document cross-platform pinned lint installers
…uid#2548) .no-mistakes.yaml has set test.evidence.store_in_repo: true since kunchenguid#2355, but CONTRIBUTING.md, docs/configuration.md, and docs/architecture.md still described the old policy of keeping evidence out of the repo in a temp directory. The current no-mistakes behavior for store_in_repo: true is to publish each run's test evidence to the orphan no-mistakes/evidence branch and link it from the PR body. That branch shares no history with code branches, so evidence never enters a pushed feature branch or the default branch, and CI's tracked personal fleet paths rule stays accurate. Docs only. No change to .no-mistakes.yaml or any workflow.
* docs: correct test evidence storage comment in .no-mistakes.yaml * no-mistakes: apply CI fixes
…nchenguid#2563) Make that a first-class option in always-loaded instructions so firstmate does not default to mediating and tearing the scout down between iteration rounds.
…chenguid#2570) * fix(bin): report remote secondmate delivery and state truthfully A steer to a remote secondmate crosses fm-on.sh to a host-local fm-send leg whose unconfirmed submit read-back (verdict=pending, typically a busy mate whose harness queues the steer) was flattened into exit 1, so the parent printed "error: text not submitted" / "error: text not sent" and discarded the pending-reply expectation for a steer that had actually landed. fm-send now carries the verdict across the ssh boundary as a documented delivered-unconfirmed exit 3: the parent reports the steer as delivered with confirmation pending, exits 0, keeps the expectation armed (awaiting_report), and closes --resolve-key decisions, while transport loss (ssh 255) and real remote failures keep failing loudly with the remote leg's stderr attached. A local unconfirmed submit now also exits 3 with an honest non-error message and still never closes a decision key. fm-crew-state.sh and fm-peek.sh no longer read a remote mate's endpoint through local probes (which misreported a healthy mate as "worktree gone" / "can't find session: remote"): both now use the true remote source over fm-on.sh, and an unreachable or unreadable remote reads as unknown-remote, never as gone or dead. * no-mistakes(document): Document remote delivery and state truth * no-mistakes: apply CI fixes
* Adopt quota-axi 0.1.29 spendPriority-primary array dispatch. quota-axi 0.1.29 publishes schema 5 with selection.spendPriority as the primary comparative signal and demotes derivation fields out of default --json. Rank comparable-fit candidates on that scalar, keep runway versus the completion horizon as a hard gate, and raise the compatibility floor so a pre-consolidation build cannot reach dispatch intake. * no-mistakes(review): Correct schema fixtures and remove prescriptive selection prompts * no-mistakes(document): Correct quota verification evidence chronology * Collapse quota-array-dispatch onto TOON-first spendPriority ranking. Decide from quota-axi's default TOON; keep --json as a rare defensive fallback. Rank by spendPriority after eligibility, reasoning-class, and runway-feasibility gates, and drop the hand-computed Pareto, pace, reserve, and window-id layers. * no-mistakes(review): Permit ambiguous JSON fallback and correct reset fixtures * no-mistakes(review): Correct runway semantics and escalate unresolved uncertainty * no-mistakes(document): Document TOON-first quota dispatch evidence
* docs: add GROK_BOT.md Grok Bot system prompt * docs: amend GROK_BOT.md with charter report-back and delegation marker * docs: classify GROK_BOT.md as public-product * docs: make GROK_BOT.md the plain Grok Bot system prompt
Refine language for clarity and consistency in instructions.
…#2595) * fix(bin): guarantee inactive-reconcile scan progress under second quantization The inactive-outcome scan computed its aggregate deadline in whole seconds, so a 1-second budget's effective value lands anywhere in (0,1]; a scan starting just before a wall-clock second boundary rounded its whole budget away mid-scan and exited having visited no child, while the durable cursor had already advanced past the never-examined child. This is the CI flake behind tests/fm-inactive-reconcile.test.sh's 'next bounded scan did not resume with the following child' (watcher-wake-lock family, portable serial 2, seen on the PR kunchenguid#2590 run). Every scan now visits at least its first due child with the per-child state-read bound floored at one second, so no invocation can be a zero-work no-op. The outer process-group kill moves to budget+1s: the scan's own deadline enforces the budget, and the kill is a backstop for a scan wedged in an unbounded wait instead of a racer that routinely preempts the clean bounded exit. The wake-lock-wait test bound tracks the backstop (3s -> 4s); the previously flaky assertion is unchanged. * no-mistakes(document): Document inactive-reconcile deadline backstop
Refactor the guidelines for Firstmate's role and delegation process, emphasizing the importance of crewmates and asynchronous work.
Clarified guidelines for handing off work to crewmates and managing secrets.
…tions deterministic (kunchenguid#2617) Three assertions in tests/fm-procevent.test.sh depended on a detached runner having finished work that the command starting it does not wait for. reconcile's replacement runner is started through detach_runner, which only forks: reconcile returns and counts the start before that runner has claimed its source or exec'd its child. Any assertion taken straight after reconcile therefore samples a race. - The publish-before-apply recovery section left its always-ready /bin/echo source registered across the recovery reconcile, so that reconcile launched a competing detached poll (observed: started=1) that then raced every later assertion for the source claim, the next capture sequence, and this home's applied record, and outlived the section holding a live claim. It is now retired before that reconcile - re-announcement is proven from the durable inbox alone and needs no registration - and started=0 is asserted so a competing poll cannot be reintroduced unnoticed. This is the same retire-before-reconcile discipline the self-announcing section already carries; that section acquired it after the identical race made its "not-autohandled: self-src" assertion read "already owned: self-src". - The crashed-leader replacement section snapshotted the replacement's claim file and execution log behind a fixed 0.5s settle window. On a loaded machine that window expires first, which is the CI flake behind "a replacement runner started without recording its own claim" and "reconcile did not start exactly one replacement source". Both effects are now waited for with the suite's bounded wait helpers; the exact one-replacement count is still asserted afterwards, unchanged. - The duplicate-start section slept 0.5s for reconcile's runner to record ownership before asserting that a second start loses to it. It now waits for that claim. Also tighten one assertion that could not fail as written: "autohandled: self-src" is a substring of "not-autohandled: self-src", so the applied path was accepted even when the runner reported the capture left for the handler. Evidence: on the unmodified suite, 128 full runs at 6-8x concurrency produced 6 failing runs, all in the crashed-leader section. On the fixed suite, 216 full runs under the same load produced none. Reverting the self-announcing section's retire-before-reconcile line reproduces "already owned: self-src" on the first iteration, confirming the shared mechanism.
) * fix(bin): keep pending-reply expectations honest on both send legs Two related asymmetries let the parent-owned secondmate reply guard drop or nag requests it should not have. Local delivered-unconfirmed dropped the expectation. A marked request whose submit read-back stayed unconfirmed (verdict=pending) is the same not-a-failure outcome the remote leg reports as delivered, but fm-send discarded the parent's pending-reply record for it, so a request that very likely landed stopped being tracked entirely. The record now stays armed on its unconfirmed-delivery marker: a correlated report still resolves it, and an unanswered one still surfaces through the library's own reconciliation. Exit 3 and the local rule that an unconfirmed answer never closes a decision key are unchanged. Remote replies were nagged for a repost they did not need. A remote mate's report reaches the parent's status log only through the asynchronous mirror in fm-procevent-remote-reply.sh, yet the guard read an absent correlated line as proof the mate never reported - even while the answer was still in flight, which is the common case because the mirror's poll window is comparable to the recovery grace. The mirror now publishes one caught-up watermark from a quiet window, and the guard admits a missing report as evidence only once that watermark passes the turn that should have produced it. A genuinely missed report still gets exactly one repost, and a channel that is behind, unarmed, or broken leaves the request durably open and un-nagged rather than nagging blind; the mirror escalates its own continuity failures as before. Tests: a local unconfirmed secondmate send keeps its expectation armed and resolvable; a mirrored correlated remote reply resolves with no repost; a stale or absent watermark withholds the repost while a fresh one still releases it; a quiet remote window publishes the watermark and retirement clears it. * no-mistakes(review): Distinguish preempted polls from quiet windows * no-mistakes(document): Clarify remote reply channel freshness * no-mistakes(lint): Annotate shared remote preemption exit constant
…d#2619) * fix(watch): honor a declared pause on a busy pane's completed-turn bound A worker that declares an external wait (`paused:`) and then blocks in one long foreground call - a review-hosting scout parked in a single blocking `lavish-axi poll`, a bounded watch loop, a rate-limit sleep - keeps its pane BUSY, so the stale path that already honors declared pauses never ran for it. The busy-pane completed-turn bound instead routed it straight into wedge_timer_check, which re-escalated "possible wedge, escalation N" (and, past the threshold, demand-deep-inspection) every FM_STALE_ESCALATE_SECS for as long as the review stayed open. busy_turn_bound_check now owns which absorber takes a crossed bound: a crew whose own last status line declares an external wait or a verified captain-held transfer takes the bounded FM_PAUSE_RESURFACE_SECS recheck, and everything else keeps the unchanged wedge timer. The discriminator is the declaration together with liveness (the caller has already confirmed the pane is busy), never a blanket silencing - a crew that declared nothing, or whose pane is not live, escalates exactly as before, and a declared pause still re-surfaces once per long cadence so a forgotten wait cannot rot invisibly. Away mode is untouched: the daemon owns pause triage there and already reads the same vocabulary. The two call sites also no longer clear pause bookkeeping in the same poll the pause cadence recorded it, which would have erased the re-surface throttle and turned the long cadence back into a per-poll re-surface. Tests: a three-phase regression fixture pins the absorbed pause, its long-cadence recheck, and the restored wedge escalation once the declaration is lifted on the same busy over-age pane. Also de-flakes tests/fm-watch-triage.test.sh, which failed spuriously on a loaded machine: fixed liveness budgets were reaping watchers mid-startup, so assertions on post-poll state passed vacuously or failed spuriously. Waits that describe a poll's outcome now wait for a completed poll cycle via the liveness beacon, the heartbeat test waits for the heartbeat it asserts on, and every wait_for_exit budget is the uniform 10s already used elsewhere in the file. * no-mistakes(review): Fail poll-cycle waits on timeout * no-mistakes(review): Prevent poll timeout test hangs * no-mistakes(document): Clarify paused busy-pane supervision
Added guidelines for decision communication to the captain.
Clarify communication protocols with crewmates regarding task delegation and reporting.
* fix(herdr): confirm local steers that native agent-state misses Herdr can leave agent_status idle for a landed Claude turn and can keep queued Enter text visible while busy, so fm-send was reporting false swallows. Confirm those cases through the shared queued-Enter verdict and a cleared composer, and keep a genuine idle pending composer as unconfirmed. * no-mistakes(review): Stop Herdr Enter retries on unreadable composers * no-mistakes(review): Reject queued delivery when all Herdr Enter sends fail * no-mistakes(review): Prevent confirmation after failed Herdr Enter * no-mistakes(review): Pace Herdr retries and clarify submit fallback * no-mistakes(review): Align Herdr submit docs with idle fallback * no-mistakes(document): Correct Herdr submit-confirmation documentation
* feat(bin): accept any-origin decision bindings with full-identity keys An aggregation surface (the bearings board) carries captain answers for holds across origins, but a binding was one-origin-per-source and the Lavish adapter capped question keys at 64 chars while real full hold identities measure 69-81. - fm-decision-hold.sh: bind <source-id> --any-origin records the (any) marker; binding prints it verbatim and answers accepts it, so the runner's feed seam carries an any-origin source with no runner change. In any-origin mode each key is a full hold identity <origin>-decision-<key>, split at its first -decision-; a key with no separator (merge/dispatch instructions) is skipped and feeds nothing, keeping non-decision answers out of the hold ledger by construction. Every existing close guard applies unchanged. - fm-procevent-lavish.sh: raise the question-key cap 64 -> 128 so a full hold identity fits; the slug-shape security property is unchanged. - tests: cross-origin closure through the real runner seam, an 81-char identity through the adapter, cap and shape refusals, routed-work skips, nonexistent-identity skips, and idempotent replay. * feat(bearings): add the /bearings lavish interactive fleet board /bearings lavish renders the bearings snapshot onto a shipped, reusable board template and arms it as a Lavish process-event source, so the captain answers Captain's Call items on the board and firstmate is woken by an ordinary check wake - no conversational turn ever blocks on a poll. - .agents/skills/bearings/assets/board-template.html: the shipped template (myfirstmate design system inlined, one fm-bearings-board.v1 JSON slot, fail-closed schema guard that renders an error card instead of an empty fleet). Per-invocation agent work is composing the payload only. - bin/fm-bearings-board.sh: build/refresh owner - fail-closed payload validation, slot injection with a round-trip check and \u003c escaping, stable board path, any-origin bind ALWAYS before arm, arm-if-absent. - bearings SKILL.md: the lavish invocation option, board composition rules, board-wake handling, and the captain-ruled merge-click authorization with its mandatory safeguards (PR resolved from the task's own meta record, wake-time green re-verification, never a red or changed PR, merges only through bin/fm-pr-merge.sh, chat echo with the full PR URL). - process-event-sources SKILL.md: one-line board-wake routing trigger. - tests: payload refusals, injection round-trip, bind-before-arm, idempotent re-arm, and template slot integrity. Fleet pickup: homes receive this after merge plus a firstmate self-update; landing timing is coordinated with the main firstmate. * no-mistakes(review): Harden bearings board validation and wake handling * no-mistakes(review): Require HTTPS for bearings board PR links * no-mistakes(review): Fail closed and bound bearings board answers * no-mistakes(review): Enforce UTF-8 byte limits for board answers * no-mistakes(review): Serve bearings board before arming and reject empty actions * no-mistakes(review): Prove bind-before-arm ordering through live answer consumption * no-mistakes(document): Document bearings board and cross-origin answers
…enguid#2707) * fix(bearings): always show decision options and a close/drop control Freeform-only Captain's Call cards hid the option buttons the board was designed around, and there was no way to drop a stale hold without inventing an answer. Require selectable options, keep freeform as a supplement, and route the reserved __drop__ answer through decline so the hold leaves Captain's Call. * no-mistakes(review): Fix drop closure and decision-only option validation * no-mistakes(review): Preserve answerability for non-decision cards * no-mistakes(document): Clarify decision drop documentation
Signature-only PRs can hide skipped review, test, or document steps. Fail unless no-mistakes >= 1.46.0 attests those three steps completed.
* fix(tests): make the changed-file map select per script and stabilize a budget flake
The changed-file map's bin/ fallback resolved a direct test reference to that
test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one
real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated
scripts, including a 341s presentation E2E with no dependency on it.
Resolve direct test references per script, and keep resolving consumer bin/
scripts through the curated map so recorded family-level coupling survives.
Also fix a load-sensitive flake: the tool-update budget deadline is whole-second
granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the
first budget check could already read as exhausted.
* feat(bin): make suite wall clock a result and let a family's concurrency be proven
--max-wall-ms fails a run whose wall clock exceeds the caller's budget, after
reporting the per-script results. A suite that stays green while outgrowing its
caller's invocation budget is the regression that got an agent killed mid-run
and retried invisibly, so duration has to be a result rather than a log note.
--pool on the isolation-proof harness runs the same concurrent proof over a
whole family, so 'is this family safe to parallelize?' is answered by a command
instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts
fail under concurrency on wall-clock assertions about reaching the next poll.
* perf(bin): schedule the changed suite concurrently, longest first
The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18
candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so
--changed now schedules its proven-concurrent scripts with bounded parallelism
and runs any unproven remainder serially afterwards, never beside them.
Concurrent runs are ordered longest-hint-first. Workers are handed scripts in
order, so alphabetical order started the 193s fm-watch-triage last and stranded
it running alone: 395s wall against a 205s balanced four-worker sum.
An explicit --jobs keeps its strict refusal, so every CI lane is unchanged.
* fix(bin): bound a hung test instead of letting it hang the suite
tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a
464ms recorded hint, and the suite had no per-script bound to stop it. An
unbounded suite is precisely what silently outruns a caller's invocation budget,
and --max-wall-ms is evaluated after the run so it cannot end one that never
finishes.
--per-script-timeout-secs terminates a script that outruns it and records exit
124, so the run still completes, accounts for the script, and fails. The
auto-concurrent --changed path applies 900s, far above the slowest real script
(the 341s Herdr presentation E2E), so it only ever converts a hang.
* no-mistakes(review): Enforce safe concurrency and descendant timeouts
* no-mistakes(review): Validate empty runs and isolation proof pools
* no-mistakes(review): Measure selection time in wall budget
* no-mistakes(review): Reap interrupted workers and bound finalization
* no-mistakes(review): Contain shutdown descendants and watchdog finalization
* no-mistakes(review): Honor remaining budget and close launch races
* no-mistakes(review): Restore timeout helper and simplify runner cleanup
* no-mistakes(review): Record isolation pool admission metadata
* no-mistakes(review): Bound Chrome reap and scope proof admission
* no-mistakes(review): Align proof scheduling and preserve budget summaries
* no-mistakes(review): Remove unreliable finalization watchdog
* no-mistakes(review): Freeze budget duration and enforce admission caps
* no-mistakes(document): Refresh test runner concurrency documentation
* no-mistakes(lint): Fix ShellCheck findings in test runner scripts
* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check`
* no-mistakes(review): Restore automatic changed-suite concurrency and timeout
* no-mistakes(review): Correct changed-suite contributor guidance
* no-mistakes(review): Reject gate-skipped isolation proofs
* no-mistakes(review): Correct automatic concurrency evidence
* no-mistakes(review): Isolate nested runner process groups
* no-mistakes(review): Remove unreliable signal cleanup machinery
* no-mistakes(test): Narrow changed-suite selection to executable contract owners
* no-mistakes(document): Document isolation proof skip and artifact semantics
* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed
* no-mistakes(review): Restore plain changed-suite automatic concurrency
* no-mistakes(review): Record resolved changed-suite worker count
* fix(bin): keep a runner change selecting its whole curated family
A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh
and bin/fm-test-isolation-proof.sh selected only their own two contract tests,
and the documentation surfaces only the audience test. That cut this branch's
own changed selection from 33 scripts to 5.
The runner executes every pure-contract-unit script, so its contract test
passing proves its logic is right, not that the suite it drives still runs.
Narrowing it also makes any wall-clock claim about the changed suite trivially
true by not running the work.
Only the unmapped bin/* grep fallback resolves per script; curated mappings keep
their recorded family coupling.
* perf(bin): admit the pure-contract-unit family to bounded concurrency
A runner-file change selects pure-contract-unit, so that family decides the
changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33
selected scripts fell to the serial tail and the selection measured 327.3s
against a 300s budget: the concurrent group was 19 scripts totalling 273.4s
while the tail alone was 215.7s.
bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice,
32 candidates, 0 failures, so the family is admitted on recorded evidence.
Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures,
inside a 300000ms budget. Also states the per-script guard's derivation.
* no-mistakes(review): Align contract-unit concurrency cap with recorded proof
* no-mistakes(document): Record final changed-suite performance evidence
* fix(bin): keep an empty changed selection clean on stock macOS Bash
Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an
unbound-variable error, while bash 4.4+ makes it a harmless no-op. The
concurrency work removed the early exit for an empty selection, so execution
fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a
contributor who changes only documentation and runs --changed got
bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable
with exit 1 and no summary, instead of a clean total=0 pass.
Restore the early exit, and guard every remaining array expansion reachable
with an empty selection. The reported duration is real elapsed invocation
time rather than a hardcoded zero, so a selection phase that outran
--max-wall-ms still fails.
Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable
error before, exit 0 with FM_TEST_SUMMARY total=0 after.
* no-mistakes(document): Document shell-bound changed-suite performance
---------
Co-authored-by: Kun Chen <kun-1@kunchenguid.com>
* feat(bin): publish per-home summary ledger * no-mistakes(review): Bound and schedule home summary publication * no-mistakes(review): Prove recurring watcher summary refresh cadence * no-mistakes(review): Bound refresh workers and publish durable spawns * no-mistakes(review): Fix atomic kill process-group coverage * no-mistakes(review): Bound state initialization within refresh timeout * no-mistakes(document): Document recurring bounded home-summary publication * no-mistakes(review): Bound and log all best-effort refresh failures * no-mistakes(review): Harden cadence and timeout regression coverage * no-mistakes(document): Document home-summary runtime tuning * no-mistakes(lint): Fix direct exit-code check in refresh test * no-mistakes(ci): Fixed remote secondmate retirement recreating the deleted home: teardown now skips side-band summary refresh when its overridden state directory was removed. Verified with remote lifecycle E2E, teardown tests, home-summary tests, ShellCheck, and git diff checks * no-mistakes(document): Clarify atomic home-summary publication guarantee
* fix(pi): gate first call on startup context * no-mistakes(document): Correct Pi startup prerequisite verification date * no-mistakes(review): Captain, fix startup process-group retirement after leader exit * no-mistakes(review): Captain, release reload exit listeners on shutdown * no-mistakes(review): Captain, complete startup exit lifecycle ownership * no-mistakes(review): Captain, release empty startup process-group ownership promptly * no-mistakes(review): Captain, supervise startup ownership and restore failure fallback * no-mistakes(review): Captain, restore live Pi supervisor execution * no-mistakes(document): docs: clarify Pi startup prerequisite delivery
* fix(pi): restore 0.84.4 adapter compatibility * no-mistakes(review): Restore Pi collapsed and expanded outcome parity * no-mistakes(review): Preserve Pi stock previews through capability probing * no-mistakes(document): Document Pi 0.84.4 renderer compatibility
…nchenguid#3273) * fix(bin): keep home-summary publication bounded and off the watcher beat A home whose tasks had accumulated ordinary status history could not publish state/home-summary.json at all, and every attempt starved the watcher's liveness beacon while it failed silently. The producer's per-task open-decision fold spent tens of milliseconds per status line on a bash 3.2 global bracket-class substitution used only as a blank-line guard. On a real home that made the whole ledger producer take minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every attempt and never completed. Replace that guard with an equivalent case glob in the one fold owner, which both the whole-file and cursor-backed folds use. Bound each per-task current-state read in the snapshot with FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh, whose dead-peer detection deliberately never kills a slow-but-alive remote command, so nothing else bounded it. Detach the watcher's two publication triggers from the poll loop. The loop owns the beacon that fm-guard.sh reads as proof supervision is alive, and an inline publication put up to a full publication deadline between two beacon touches. A single in-flight publication is tracked so a slow one cannot accumulate clones. Report a repeatedly failing publication at session start. Publication stays deliberately non-fatal to its caller, so the existing bounded home-local failure record is now surfaced as a HOME_SUMMARY bootstrap line once the ledger is absent or stale and failures have been recorded since. * no-mistakes(review): Preserve home-summary failure attempt ordering * no-mistakes(review): Enforce durable home-summary single-flight and ordering * no-mistakes(review): Derive failure ordering from publication boundaries * no-mistakes(review): Restore best-effort failure logging and publication scoping * no-mistakes(review): Make ordering regression sensitive to one failure * no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance
…henguid#3268) * fix(supervision): classify the appended status span, not the last line An actionable project update could be classified as routine and absorbed, so a worker that raised a decision, hit a blocker, failed, or finished stalled silently with the captain never told. Trigger, mask, symptom. A worker appends a captain-relevant event (`needs-decision`, `blocked`, `failed`, `done`). Any later routine append - a `working:` progress note - lands before the supervisor classifies the batch; the watcher's 30s signal-grace linger exists precisely to coalesce a status write with the same turn's turn-end, so this window is ordinary rather than rare. Both supervisors then asked "is the LAST line captain-relevant?", read the routine line, and absorbed the wake. The `.seen-*` suppressor advanced either way, so nothing ever re-read the event. When the crew was also provably working, the no-verb fallback absorbed it too, which is why the event disappeared completely instead of surfacing late. Reproduced end to end against a real watcher before any change: with the trailing `working:` append the watcher never exits and the wake queue stays empty; with that one line removed - the smallest counterfactual - the same `needs-decision` surfaces and queues. The away-mode daemon's `classify_signal` returns `self|routine signal` for a `blocked:` event under the same mask, which is the worse case because no captain is present to notice. The proven path was already in the tree: `status_open_decisions` fixed this exact masking for the durable decision fold, and its header states the rule - reading an append-only event log last-event-wins cannot represent an earlier event that a later unrelated line moved past. The classification path was never migrated to that read model. That is the earliest divergence, and the fix is to migrate it rather than to special-case the symptom. `status_span_first_actionable` in bin/fm-classify-lib.sh is the new single owner: it reads the bytes at or after a caller-supplied position and returns the first still-live captain-relevant event. Each supervisor supplies its own position, because the always-on watcher and the away-mode daemon classify the same stream independently and must not share one cursor: the watcher reads the size already recorded in its `.seen-*` signature (no new state) and its `.hb-surfaced-<task>` backstop marker, and the daemon its `.subsuper-seen-status-<task>` marker. Those two markers held the escalated line and now hold the escalated-through byte offset, which also removes a second defect in the same code - content dedup silently swallowed a genuinely new event whose text repeated an older one. An absent, malformed, or past-the-end position reads the whole log, so uncertainty surfaces events rather than losing them, and a marker an older build wrote as a status line reads that way too. Status logs are only ever appended to, including across a reused task id, so a recorded position keeps its meaning. A `needs-decision`/`blocked` event in the span is retired only when the whole-file fold proves its key closed; `status_open_decisions` stays the sole owner of that rule, so same-key reopening and reserved-key namespaces need no second implementation here. Every other captain-relevant event is terminal and always actionable. Both backstops now walk every status log instead of only those whose last line looks captain-relevant, because the event a backstop most needs to catch is exactly one a later append has moved past. That leaves `scan_captain_relevant_statuses` with no callers, and it is removed rather than left as a working copy of the defective read model. Regression coverage exercises the classifier and both supervisors through their own interfaces: the masked decision, the captain-reported release/install completion followed by cleanup chatter, and the away-mode blocker all surface; a routine append after an already-classified event stays absorbed, so the fix does not convert ordinary progress into wakes; and the heartbeat backstop catches a masked event the per-wake path missed. The end-to-end watcher tests drive a real fm-watch.sh with the crew reported as provably working, which is the configuration that made the original stall silent. Two further claims in the supplied RCA are deliberately not patched here. "Repeated operational recoveries produced all-clear replies despite known actions" is downstream of this same cause, not an independent contributor: an all-clear reply is the documented response when the specific event needs no action, so a classification that wrongly reported "no action" produces it, and correcting the classification removes it. "The project was subjected to validation requirements outside its accepted path" is delivery-mode selection, which AGENTS.md section 7 owns; no code changed here touches it, so it is out of scope. Harness and backend axes were inspected rather than assumed: nothing in this path reads a vendor-emitted signal. The status log's format and append protocol are Firstmate's own and identical for every harness, and no runtime backend reads or writes `.status` files (`bin/backends/*` contain no reference to them). The surrounding triage's only backend touchpoints - pane capture and the authoritative crew-state read - are unchanged. No live-harness guard applies and no per-harness verification record changes. Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `bin/fm-test-run.sh --changed --base origin/main`. * no-mistakes(review): Prevent status races and surface classification failures * no-mistakes(review): Surface unreadable signals and preserve AFK endpoints * no-mistakes(review): Route stale wakes through captured span verdicts * no-mistakes(review): Retire supervision offsets with reused task state * no-mistakes(review): Bind status offsets and preserve live decision origins * no-mistakes(review): Strengthen status identity with verified birth time * no-mistakes(review): Skip turn-end markers during status classification * no-mistakes(review): Preserve status presentation with platform-strength identities * no-mistakes(review): Retain failed wakes and advance routine checkpoints * no-mistakes(review): Surface all events and retain unreadable wakes * no-mistakes(review): Treat absent status logs as successful empty spans * no-mistakes(review): Bound repeated classification failures with durable receipts * revert(supervision): drop the failure-receipt and durable-retry machinery Captain-authorized revert to the minimal fix. Review rounds added a durable failure-receipt store and wake-retention-on-failure to bound repeated classification failures. That machinery grew larger than the fix it protected and kept producing its own defects: an unreadable log still looped forever because the always-on watcher never consulted the receipt, and the receipt was persisted before its diagnostic was durably queued, so a crash in between swallowed the alarm outright. Those two defects go away with the code that contained them rather than being repaired. Removed: the failure-receipt path, fingerprint, record and clear helpers and their retirement bookkeeping; the retention of a durable wake when classification fails; and the error-propagation plumbing in both supervisors that existed only to drive them. Kept, because it is the accepted fix rather than the declined machinery: span classification of the events appended since a supervisor last looked, in both supervisors and both backstops; reporting every actionable event in a span and committing a position only through what was reported; naming the live opening of a reopened decision; treating an absent log as ordinary and an unreadable one as worth reporting; the non-.status filter; and the platform-strength identity that guards a position commit without failing a read. Replacement behavior for a log that cannot be classified: report it once, do NOT advance the classification position so the content is classified from where it stopped once readable, and DO advance the wake signature so the report is bounded to one per distinct file state. Reporting and reading are different acts: telling the captain about a log is not the same as having read it, and only the latter may move a classification position. The residual risk is explicit and accepted: there is no guaranteed automatic retry inside a crash-mid-read window, and the locked session-start replay of the durable queue covers it. That rationale is recorded at mark_escalated_seen so a future reader does not reintroduce the retry as a "missing" guarantee. Also fixes lint failures that arrived with the review-fix commits and were never caught because the run never reached its lint step: an unfollowable conditional source directive, a second unquoted-expansion site left after a call was split across lines, cleanup of the file being read inside its own read loop (restructured to one post-loop teardown rather than three in-loop copies), stub functions in tests that are invoked indirectly, and a test local left unused when its assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch, so these were introduced here. Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue, wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures). * no-mistakes(review): Correct classification failure contract documentation * no-mistakes(review): Bound unreadable status reports without skipping classification * no-mistakes(review): Preserve escalation markers when buffering fails * no-mistakes(review): Detect permission recovery without advancing classification * no-mistakes(document): Document status span classification contract * no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed * no-mistakes(review): Escalate blockers while preserving declared-wait cadence * no-mistakes(review): Clarify actionable events override wait self-handling * no-mistakes(review): Surface rejected decisions and dangling status links * no-mistakes(document): Document reserved-key reconciliation classification * no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check * no-mistakes(document): Correct away-mode classification documentation
…#3289) * docs: split harness adapter operations reference * no-mistakes(review): Fix harness adapter routing and ownership contracts * no-mistakes(review): Prune duplicate harness adapter ownership prose * no-mistakes(review): Fix default effort routing and Grok max semantics * no-mistakes(review): Remove source-only routing test and duplicate semantics * no-mistakes(review): Add local harness adapter instruction evaluation * no-mistakes(review): Fix harness evaluation gating and change mapping * no-mistakes(test): Captain, require explicit harness instruction evaluator model * no-mistakes(document): Fix harness adapter documentation references
* test(fixtures): share fake-toolchain and spawn-world builders Future tests can start from tests/fixtures.sh instead of copying stubs, and a no-mistakes version-floor bump is one constant rather than a multi-file edit. Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen, fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile. Left for opportunistic migration: remaining make_spawn_fakebin copies (trace-context, kimi, muse, backend), the make_stubs send cluster, and the fake no-mistakes version banners in bootstrap/session-start/secondmate suites. Did not touch tests/fm-pr-check-security.test.sh. * no-mistakes(review): Prevent fake SSH test from blocking on stdin * no-mistakes(document): Clarify shared fixture documentation * no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks * no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head * no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head
* feat(bin): retire completed PR-check migration machinery Every registered home already carried both completion markers, and no installer still creates pre-migration checks. Remove the one-time migrate script, its bootstrap/watch/teardown/docs surface, and migration-path tests without weakening live check-trust or PR-poll authentication. * no-mistakes(review): Restore live PR-check security coverage * no-mistakes(document): Refresh retired PR-check documentation * no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks
…3247) * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * no-mistakes(review): Split extension conformance and forward remote transfer input * no-mistakes(review): Forward malformed remote payloads through fm-on * no-mistakes(review): Bound extension coordinator failure cleanup * no-mistakes(test): Skip repeated orphan sweep in coordinator children * no-mistakes(test): Queue isolated extension sections through bounded workers * no-mistakes(test): Bound extension coordinator lane cleanup * no-mistakes(test): Split remote lifecycle coordinator sections * no-mistakes(test): Coordinator probes pass; aggregate deadline remains * no-mistakes(test): Launch extension sections concurrently * no-mistakes(test): Fix coordinator marker publication * no-mistakes(test): Stabilize extension binding coordinator timing * no-mistakes(lint): Fix extension binding ShellCheck warnings * fix(extensions): prove invocation cleanup before retirement * no-mistakes(review): Harden process-event inbox confinement * no-mistakes(review): Preserve legacy capture parity * no-mistakes(review): Protect external registry staging * no-mistakes(test): Stabilize bounded extension conformance aggregate * no-mistakes(document): Document external evidence confinement * no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh` * no-mistakes(review): Harden extension staging and lifecycle reservation * no-mistakes(review): Harden external staging and lifecycle reservations * no-mistakes(review): Wire capture helper into remote conformance * no-mistakes(review): Pin external capture handoff and signal failures * no-mistakes(review): Bind pinned capture authority to inherited descriptor * no-mistakes(review): Harden descriptor-bound capture authority * no-mistakes(review): Harden core capture reservation authority * no-mistakes(review): Harden capture reservation boundaries * no-mistakes(review): Harden capture reservations and cleanup * no-mistakes(review): Harden capture handoff and reservation cleanup * no-mistakes(review): Bind capture handoff to claim descriptors * no-mistakes(review): Release lifecycle locks after host crashes * no-mistakes(review): Pin reservation recovery to recorded state roots * no-mistakes(review): Reject control bytes in claim state roots * no-mistakes(test): Stabilize extension capture descriptor handoff * no-mistakes(document): Document extension capture authority boundary * no-mistakes(lint): Fix ShellCheck extension binding warnings * no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks * no-mistakes(document): Correct extension namespace creation timing * no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes A promoted scout used to receive a free-form placeholder instead of the mode-specific Definition of done a briefed ship worker gets, so it never saw the ask-user escalation rule or the --yes prohibition. That gap is the concrete reason one incident's worker drove validation with --yes and answered its own ask-user findings. - Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh so the two contracts cannot drift. - bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the scratch inventory, clean base, ship branch, and that Definition of done, and prints the fm-send.sh command that delivers it. - State the --yes ban as a prohibition rather than a preference, without claiming an enforcement the tool does not provide. - Cover both through the real promotion and brief paths in tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh. * no-mistakes(review): Publish promotion instructions before committing task state * no-mistakes(review): Supersede conflicting scout delivery rules after promotion * no-mistakes(review): Reject invalid promotion instruction destinations * no-mistakes(document): Align documentation with promotion delivery contracts * no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check * no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean * no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head * no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
) * fix(bin): present complete Lavish board feedback as structured output Give the Lavish adapter a read-only presentation so a handler sees every annotation and the session-ending tag=message as its own field, instead of grepping a truncated raw capture. * no-mistakes(review): Preserve unquoted messages and prioritize captain prose * no-mistakes(document): Document structured Lavish result reads * no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves Dispatch and completion each moved a task's physical record and its backlog row as two independently timed steps, so a crash or a forgotten follow-up could leave the two disagreeing: a record with no in-flight row, an in-flight row with no owner, or a finished task still shown in flight. Fold each backlog transition into the script that performs the physical change, under the per-task lock it already holds and before it reports success. Dispatch moves the item to In flight after publishing the task record and fails loudly, removing its provisional record, when that transition cannot land. Completion records an authoritative close and performs it before removing the record, so an interrupted cleanup can be finished later, and its closing message now confirms what already happened rather than instructing a future step. Add a same-home reconciliation sweep to session start so a home that was interrupted mid-transition settles its own books on restart, replaying a recorded close and restoring an in-flight row it already owns a worker for. It never reads or writes another home; the fleet snapshot and the cross-home nudge stay as backstops. Close records are validated before they are trusted: the file is read as raw bytes and rejected outright when it carries a NUL or other control byte, every field must be well formed and non-duplicated, the id must match the record it was found under, the data location must resolve inside this home, and each close argument must carry a permitted, well-formed value. Writer and reader share one validator so a record this home publishes always remains replayable, independent of locale. Homes configured for a manual backlog, and homes with no backlog at all, stay exempt and are unaffected. * no-mistakes(review): Remove stale bootstrap migration helper invocation * no-mistakes(review): Preserve pending closes and narrow signal deferral * no-mistakes(review): Record close before destructive teardown * no-mistakes(review): Refuse pending closes before creating resources * no-mistakes(review): Guard relaunches and preserve cleanup warnings * no-mistakes(review): Reject symlinked records and clarify cleanup guidance * no-mistakes(review): Align dispatch eligibility and protect close replay * no-mistakes(review): Unify exact task incarnation parsing * no-mistakes(review): Render resolved configured backlog path * no-mistakes(review): Harden transition path boundaries against symlinks * no-mistakes(review): Validate lifecycle state before resource actions * no-mistakes(review): Enforce transition tooling and continuous state locks * no-mistakes(review): Consolidate same-home lifecycle file boundaries * no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts * no-mistakes(review): Reject final-component lifecycle record symlinks * no-mistakes(document): Document lifecycle record path boundaries * no-mistakes(lint): Quote literal done tokens in atomicity tests * no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched * no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check` * no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks * fix(records): validate record bytes without an uncurated tool The byte validation added for close records and directory paths shelled out to od. The spawn and teardown lifecycle runs under a curated command set that deliberately excludes it, so on any restricted PATH the check could not run, the data directory read as unresolvable, and dispatch and cleanup refused - wedging the lifecycle rather than protecting it. An earlier attempt made the failing test pass by adding od to that curated set. That fixed the test to agree with the defect and quietly widened the contract the fixture exists to pin, so it is reverted here. Inspect the bytes with perl instead, which is already in the curated set and already used in this repo for the same portability reason. The emitted values are identical to od's, so the rejection semantics are unchanged: NUL and other control bytes are still refused, legitimate paths containing spaces or non-ASCII characters still round-trip, and the check stays independent of the process locale. The restricted-PATH teardown case now passes because the validator no longer needs od, not because the fixture was loosened. * no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication * no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342) * fix: publish promote and Relay meta rewrites through contained replace Bare mv still rewrote live task records in place, so a symlink meta could be followed to a target outside state/. Route those field rewrites through the shared publisher and drop the unused library aliases. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse dangling symlinks during X metadata clear * no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects * no-mistakes(review): Exercise dangling symlink refusal through clear helper --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877) * fix(watch): absorb a turn-end whose pane churned since the previous poll The watcher's "absorb a benign turn-end when the crew is provably working" triage was structurally unreachable for any harness whose semantic busy state has no verified source. crew_absorb_class only reports working for an actively running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can only answer unknown for such an adapter, so codex crewmates surfaced a signal wake at every turn boundary with nothing to act on - a full supervisor drain, inspect and acknowledge turn per worker turn, scaling with the number of workers in flight and drowning the wakes that matter in identical noise. Widen the proof rather than bound the wake rate. A wake carrying only bare turn-ended markers is now also benign when the task's pane content changed since the previous poll, compared against the same state/.hash-* marker the staleness backbone already records and already trusts as liveness. That evidence claims no harness semantics, so it fabricates no busy verdict an adapter has not earned, and it needs no adapter cooperation. Absorb stays evidence-driven in both directions. A wake naming any status file keeps the strict proof, every captain-relevant verb still surfaces immediately, and an unresolvable task, a missing prior hash, a failed or empty capture, or an unchanged pane all surface exactly as before. The absorb defers rather than swallows: a crew that has stopped renders nothing further, so its now-static pane surfaces through the staleness backbone within a poll or two. Bounding the surfacing rate instead would have suppressed genuinely stopped workers. The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which owns it, and costs one bounded capture reached only for a no-verb turn-end whose crew is not already provably working. * no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Captain, reject malformed pane-churn hashes * no-mistakes(document): Document pane-churn turn-end evidence * no-mistakes: apply CI fixes * fix(watch): gate and bound the pane-churn turn-end absorb Make the pane-churn form of positive work evidence opt-in per home and bound how long it may defer one endpoint's bare turn-ends. Absorbing a bare turn-end on pane churn is now reached only when the home creates config/turnend-churn-absorb. The other two proofs read a verdict the harness itself vouches for, while this one infers execution from rendered bytes, so widening the absorb is a home's choice rather than a default every fleet inherits. With the flag absent the predicate returns on its first line and triage is unchanged. Churn and pane staleness read the same pane, so neither can be the other's only backstop. A pane that renders continuously never presents the two consecutive identical hashes the staleness backbone needs, so an unbounded churn absorb left a worker that had genuinely stopped behind such a renderer with no path to surface at all. One endpoint's turn-ends may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS, tracked in state/.churn-since-*, after which the wake surfaces and the window restarts. The bound is evaluated before any .stale- state is touched, so a wake that surfaces there leaves the staleness backbone's own classification alone. Covers both with behavioral tests: the same churning fixture that absorbs with the flag surfaces and queues without it, and a spent deferral window surfaces and restarts. The four existing safety guards now run with the flag enabled so they keep proving their specific guard. * no-mistakes(review): Fail closed on invalid churn deferral state * no-mistakes(review): Validate persisted churn deadlines before arithmetic * no-mistakes(review): Make churn deadlines transactional and bounds safe * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Clarify pane-churn supervision documentation * no-mistakes(lint): Fix watcher arithmetic lint issues * no-mistakes: apply CI fixes * no-mistakes(document): Clarify pane-churn fail-closed documentation * fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194) * fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Align pane-churn watcher documentation * no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement Agents were improvising rm of check files with unset STATE/ID, which wedges headless panes. Unregister validates the id and state directory first. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse explicitly empty custom-check state overrides * no-mistakes(document): Document custom-check retirement safety contract --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221) * Add quota exhaustion detection and safe fallback helpers - bin/fm-procevent-quota.sh: generic procevent adapter that arms a recurring quota-axi --json poll and wakes firstmate when a tracked provider's effectivePercentRemaining drops below a threshold or its runway.status becomes exhausted_now. - bin/fm-quota-choose.sh: worker-side helper that picks the first ranked harness:model candidate with positive effectivePercentRemaining. - AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document the new helpers and the mid-task quota-exhaustion wake path. - tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON source. * no-mistakes(review): Fix quota polling and scope bounds * no-mistakes(review): Enforce safe default quota selection * no-mistakes(review): Handle decimal quota values safely * no-mistakes(review): Fail closed on invalid quota inputs * no-mistakes(review): Reject empty quota candidate segments * no-mistakes(review): Harden quota parsing and timeout ownership * no-mistakes(review): Reuse captured quota snapshots consistently * no-mistakes(review): Match quota using explicit candidate providers * no-mistakes(review): Centralize fail-closed quota schema validation * no-mistakes(review): Reject out-of-range quota percentages * no-mistakes(review): Validate quota runway status enum * no-mistakes(review): Tighten quota scope and status contracts * no-mistakes(review): Preserve unknown quota and exact product bounds * no-mistakes(review): Preserve provider-level unknown quota * no-mistakes(review): Reuse canonical verified harness validation * no-mistakes(document): Document mid-task quota handling * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(docs): restore default routing contract, keep quota helper optional Restore the AGENTS.md section 4 always-loaded routing paragraph the PR had deleted, so the standing TOON-first intake, spendPriority ranker, every-candidate accounting, and load-trigger contract stay exactly as before this PR. The mid-task quota wake is optional and must not alter default routing. Restore the quota-array-dispatch skill ownership line to section 4 as the always-loaded intake boundary owner; keep the worker-side helper section as an addition only, without rewiring ownership or load triggers to section 13. * fix(bin): use harness-keyed quota matching in optional helper Revert fm-quota-choose.sh from harness:provider:model tuples back to harness:model candidates with harness-keyed provider matching, per the resolved ask-user finding. The helper is optional; authoritative multi-provider routing (provider discovery from the harness catalog and quota matching by that explicit provider) stays owned by AGENTS.md section 4 and the quota-array-dispatch skill intake procedure, not the helper. Document the multi-provider limitation in the helper header and the quota-array-dispatch skill: the helper maps each harness to one primary provider family only, so a candidate whose established provider differs from that primary family is checked against the wrong quota row. Use it only when the brief fixed the candidate order and every candidate's provider is the harness's primary family. The helper still consumes one already-captured default-TOON or JSON snapshot via stdin or --snapshot and never calls quota-axi itself, so it selects from the same quota state as the intake. * no-mistakes(review): Fix Muse quota mapping and helper contract docs * no-mistakes(review): Reject known-empty quotas and map quota tests explicitly * no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse * no-mistakes(review): Fix quota retirement and dependent regression coverage * no-mistakes(review): Accept zero-row quota TOON snapshots * no-mistakes(review): Enforce quota semantics status consistency * no-mistakes(review): Veto dispatch on any exhausted applicable scope * no-mistakes(review): Record exhausted quota scope in wake details * no-mistakes(review): Fix quota help and control dependency coverage * no-mistakes(review): Decode quoted TOON fields and document quota wakes * no-mistakes(review): Validate zero-row TOON and map timeout coverage * no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes * no-mistakes(review): Validate complete nonzero TOON envelopes * no-mistakes(review): Accept producer-shaped quota TOON envelopes * no-mistakes(review): Support empty quota arrays and validate counted rows * no-mistakes(review): Harden TOON completion, scopes, and quoted fields * no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields * no-mistakes(review): Allow unknown headroom under known semantics * no-mistakes(review): Reject noncanonical quota identities * no-mistakes(review): Preserve empty quota polling and validate attention identities * no-mistakes(review): Reject noncanonical provider watches * no-mistakes(review): Validate all candidates before quota selection * no-mistakes(document): Correct quota helper safety documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated read preferred element text over prompt, so an annotate-and-comment item dropped the captain's words. Surface prompt as its own field. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Filter non-comment prompts from Lavish reader output * no-mistakes(document): Clarify Lavish comment presentation contract * no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure * no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure * fix(bin): always emit Lavish comments and use real annotation fixtures Stop inferring comment provenance from prompt==text. Real pure annotations have no prompt, so always-emit does not duplicate. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420) * Fix public-followup register crashing on empty lock arrays under bash 3.2. bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock. The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there. * no-mistakes(document): Document stock Bash registration coverage * no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5 * no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment * no-mistakes(review): Clear inherited supervision model from Herdr launches * no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent A Discord support thread's screenshots were never seen by the agent handling the mention. The relay delivered them and the poll stashed them: the reporter's images arrived on the `thread_starter` entry of `in_reply_to_chain` while the mention's own media list was empty. The gap was in the responder's playbook, which enumerated a fixed field list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so made every other field, attachments included, invisible. Fix it where the gap is, in prose: - Read the complete payload object rather than a fixed field list, so media and later relay fields are never skipped again. - Fetch and view attached media with the agent's own tools, on the mention and on every chain entry, and call out the common shape where only the thread starter carries the screenshots. - Restrict those fetches to known-good platform media hosts over https (Discord: cdn.discordapp.com, media.discordapp.net, images-ext-1.discordapp.net, images-ext-2.discordapp.net; X: pbs.twimg.com, video.twimg.com), report a blocked host instead of working around it, and treat everything fetched as untrusted public input on the same terms as the surrounding thread text. The poll stays out of it and downloads nothing, so no third-party bytes are pulled on the polling path. The new test pins the contract the playbook depends on: a mention in the incident's shape, with an empty top-level media list and screenshots on the thread starter, must reach the inbox with the payload intact and its media URLs unfetched. * no-mistakes(review): Preserve media authority and enforce poll-only fetching * no-mistakes(document): Clarify Relay attachment safety prose
) * Defer inactive startup reconciliation * no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably * no-mistakes(review): Require worker phases to cover startup requests * no-mistakes(review): Make diagnostic wakes safely acknowledgeable * no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits * no-mistakes(review): Distinguish malformed presentation locks from live contention * no-mistakes(review): Bound no-ack drain queue lock acquisition * no-mistakes(document): Document bounded presentation-lock drain behavior * no-mistakes(lint): Annotate bounded lock output global * no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
Resolves ~75 conflicted paths spanning AGENTS.md, bin/, .agents/skills/, docs/, tests/, and .github/workflows/. Fork-only features (herdr fixes, decision-key fold, arm verification, afk handback, captain-messaging/ webhook receivers, fork-only delivery) are preserved; upstream's PR-check migration retirement, trusted process-event extensions, atomic backlog transitions, and Pi supervision-branch wake-claim model are integrated. Claude-Session: https://claude.ai/code/session_016PLevkhoYcz3TArb1x5NFe
fm-fork-deliver.sh execed tests/*.test.sh directly, relying on the shebang plus executable bit. Upstream ships some new test files as non-executable (confirmed on upstream/main itself), which the CI tests job already tolerates by invoking /bin/bash tests/foo.test.sh. Match that so a future upstream addition doesn't recurringly break fork-only delivery. Claude-Session: https://claude.ai/code/session_016PLevkhoYcz3TArb1x5NFe
…lake CI red on PR #72: portable serial 3 (job timeout) and 4 (real test failure). Both trace to the same root cause, neither a merge-resolution mistake: - tests/fm-remote-transport-lanes.test.sh (a brand-new upstream test, content untouched by any conflict resolution) polls for an abandoned stage directory to be reaped within 5s, but the worker's periodic stale-stage sweep defaults to a 60s scan interval (FM_REMOTE_JOB_REAP_SCAN_SECONDS). Verified: the test fails reliably without an override and passes reliably with one. Set it alongside the existing FM_REMOTE_JOB_STAGE_REAP_SECONDS=1 override. - 21 new test files landed with the upstream catch-up merge; 15 of them had no entry in bin/fm-test-run.sh's static portable_serial_weight_hints table and fell back to the 20s default weight, understating several by 3x or more. That skewed the longest-processing-time shard assignment enough to push two of four shards toward the 20-minute job timeout. Refreshed the table from a real local `--lane portable-serial --json` run per the documented procedure in docs/fm-test-portable-shards.md (147 scripts now measured, up from ~122); shard balance is now 733121/733121/733124/733125 ms (was reaching the 20-minute cap on two shards). Updated that doc's reference tables to match. Claude-Session: https://claude.ai/code/session_016PLevkhoYcz3TArb1x5NFe
Bre77
added a commit
that referenced
this pull request
Sep 2, 2026
PR #72 burned large GitHub-runner time on behavior-test shards and a timing-hint regeneration cycle for failures a local run had already triaged. Guard every ci.yml job to the upstream repo, matching the pattern no-mistakes-required.yml already uses, and update the fork-only delivery docs/tooling to reflect that the existing fm-fork-deliver.sh local gate is now the sole testing enforcement. Claude-Session: https://claude.ai/code/session_01Gf2V4SheNp3ucXtZYRCFjM
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
upstream/main(kunchenguid/firstmate): 107 upstream commits since the last catch-up (Merge upstream/main into fork integration line (catch-up 3, 42 commits) #70), from the wake-drain presentation-lock bound fix back through the retirement of the legacy PR-check migration machinery. Compare: kunchenguid/firstmate@6789876...f42a629references/*.md, trusted external process-event extensions, atomic task-record/backlog transitions, and the Pi supervision-branch wake-claim model.Conflict resolution, defaulting to upstream and keeping ours only where upstream has no equivalent or evolved a fork-only feature differently:
AGENTS.md- union throughout: kept every fork-only line (commit-identity/PR-label metadata,fm-notify-captain.shpaging, the three fork-only alert-response skills,--fork-onlybrief scaffolding) while taking upstream's new config/state rows, the deferred-startup-stage rewording, the durable steering-inbox description forfm-send, and the PR-check-migration references dropped frombootstrap-diagnostics' trigger list (matches the retirement below).bin/fm-pr-check-migrate.shdeleted, touchesfm-bootstrap.sh,fm-pr-lib.sh,fm-session-start.sh,fm-watch.sh, docs, tests) - took upstream's retirement. Verified this fork's own home already carries both completion markers (.pr-check-migration-scan-v1,.pr-check-migration-v1) and a quarantine artifact proving the fork's ownrecord_ambiguous_metadata_reasonaddition to that script had already fired, so nothing live depends on the removed machinery.bin/fm-watch.sh(8 hunks) - kept the fork's Herdr stale-dedup (stale_signal/fm_backend_busy_state) and its pane-signal-reset fix (clearing$sfon a changed hash, not just the wedge timer); took upstream's new away-mode daemon-handoff forbusy_turn_bound_check, theturnend-churn-absorbpane-churn absorb feature, and the steering-inbox loss detector (inbox_steer_check) wholesale, since the fork had no competing logic there.bin/fm-pr-lib.sh- union of two independent additions at the same insertion point: fork's PR-label functions (fm_pr_label_name/fm_pr_apply_label) plus upstream's new merge-notification canonical-identity marker.bin/fm-send.sh(initially looked like a fork-vs-upstream architecture split: the fork's--resolve-keydecision-closing vs. upstream's inbox+doorbell delivery rewrite) - on close inspection the decision-closing functions were already shared, unconflicted code fully wired into upstream's new architecture; the actual conflicts were header docs (took upstream's fuller version) and four now-dead inline remote-pane-typing hunks the surrounding unconflicted code already documents as unreachable, which upstream's removal made explicit.tests/fm-send-remote-delivery.test.shis upstream's parallel rewrite for the surviving mechanism (16/16 passing).bin/fm-spawn.sh,bin/fm-brief.sh,bin/fm-promote.sh- union of independent fork/upstream additions (spawn's--helpearly-exit alongside upstream's new backlog-directory guard; brief's--fork-onlymode-override kept alongside upstream'sfm-dod-lib.shsingle-owner DOD renderer and new steering-inbox section; promote's ship-instructions-file wording, which is what upstream's file-based flow actually produces).docs/scripts.md,docs/architecture.md,docs/configuration.md,docs/documentation-audiences.json,CONTRIBUTING.md- union; cross-checked several disputed rows against the actual resolved scripts' own header comments and--helpoutput rather than guessing (e.g.fm-test-run.sh's--proven-isolatedterminology,fm-wake-drain.sh's new per-actor claim model).tests/fm-pr-check-security.test.sh- the PR-check-migration retirement removed ~1,700 lines of now-unreachable migration-fault-injection tests and their private helpers; verified every dropped helper's only call sites were inside the removed blocks before deleting, so nothing else in the file references them.Fork-only surface
Spot-checked against
upstream/mainafter the merge, all present: ClickStack/BetterStack webhook receivers and their alert-response skills, captain messaging over Twilio RCS,fm-notify-captain.sh,fm-pr-activity-poll.sh,config/pr-label,fm-harbor.*,fm-fork-deliver.sh+--fork-onlybriefs,docs/crew-memory-cap.md, the Herdr stale-dedup and pane-signal-reset fixes infm-watch.sh.Fixed along the way
Two small bugs surfaced only once the merge tree was assembled (neither is a conflict-resolution mistake in isolation - both are interactions between independently-correct hunks):
bin/fm-brief.sh: the--fork-onlycase set$DODvia its own heredoc, but an unconditionalfm_dod_block "$MODE" "$ID"call three lines later (upstream's new single-owner DOD renderer) clobbered it for every mode, and that renderer has nofork-onlycase -fm-brief.sh id proj --fork-onlyexited 1. Guarded the call to skipfork-only.tests/lib.sh: fixture homes aremkdir -p'd with no explicit mode, so a permissive ambient umask (this host defaults to 0002) leavesstate/group-writable, which upstream's newfm_procevent_private_directory_validcheck correctly refuses. Pinnedumask 022once in the shared test library rather than patching every affected fixture.bin/fm-fork-deliver.shexecedtests/*.test.shdirectly by shebang; upstream ships some new test files non-executable (confirmed onupstream/mainitself), and the CI tests job already tolerates that by invokingbash tests/foo.test.sh. Matched that so a future upstream addition doesn't recurringly break fork-only delivery.Validation
bin/fm-lint.sh: clean (ShellCheck 0.11.0, actionlint 1.7.12, both pinned).bin/fm-test-run.sh --all: 184 scripts, 14 failed - all confirmed pre-existing and environment-specific in this host, not merge regressions, the same category Merge upstream/main into fork integration line (catch-up 3, 42 commits) #70 hit last round:.tsextensions via bareimport()with no repo-levelpackage.jsondeclaring"type": "module"(fm-busy-adapter-wiring,fm-pi-watch-extension,fm-pi-branch-extension, and the Pi-guard/Pi-continuation subtests insidefm-turnend-guard,fm-sessionstart-nudge,fm-watch-recovery-loop).fm-test-run.test.sh) needsruby, not installed here - same gap Merge upstream/main into fork integration line (catch-up 3, 42 commits) #70 hit.fm-backend-herdr-focus-flash-e2e.test.sh) and 1 file (fm-watcher-lock.test.sh) are timing-sensitive live Herdr/signal-delivery e2e tests; neither script was touched by any conflict resolution.fm-remote-transport-lanes.test.sh) failed once on a reap-age timing assertion during this ~65-minute full run; not reproduced in isolation.ruby, actionlint, and (presumably) a Node version this repo's own CI has already verified against; ran locally with actionlint installed viabin/fm-install-actionlint.shto confirm the two lint suites are otherwise clean.--skip-validatesince the above are host-environment gaps the gate can't get past locally, not because validation was skipped - the full suite ran to completion first.Nothing under
projects/was touched and no running firstmate home was self-updated;/updatefirstmateremains the captain's step after this lands.