chore: sync fork main with upstream a5f3cbe (25 commits) - #5
Merged
Merged
Conversation
* fix(bin): make Claude auto-arm continuity self-heal past a hung claim On a Claude primary, a Stop-hook auto-arm process that hung mid-arm held the single-flight owner lock with its epoch ledger frozen at outcome=arming, and the abandonment proof read any live lock holder in arming as legitimately deciding forever. Every later Stop firing exited 0 at the lock, the turn-end guard kept deferring to the hung owner as recovery under way, and the watcher was never auto-re-armed again for the rest of the session - supervision survived only on manual arms and lapsed between them (the 2026-08-26 watcher flap). Corrections layered onto the lock-held-across-arm shape each reopened the same concurrency class one level down, so this replaces the claim machinery wholesale with a generation-based optimistic design: - The epoch ledger's monotonic sequence IS the claim generation; the two-line entry (classic epoch record plus the claimant's MANDATORY pid-identity) is the claim. Every firing defers to a live OPEN claim: outcome arming, owner alive, identity recomputes and matches, and not stuck (entry and watcher beacon both older than the guard grace). - A finished, dead, identity-mismatched, identityless, or stuck claim is superseded by simply taking the next generation - no signalling or revocation of a steady-state predecessor. - No mutex is held across arming or output; the owner lock survives only as a micro-mutex around individual ledger writes. A superseded owner goes completely silent: ownership is re-verified before every arm invocation, episode-state mutation, ledger write, and continuation. - The irrevocable commit point of a translation is the exit status (the harness delivers the collected stderr only on exit 2), so the owned terminal ledger write is the atomic commit: the winning generation exits 2 unconditionally after it, a refused one exits 0 silently even after printing, and the once-per-episode failure notice commits in the same owned critical section as the winning failed write. Two bounded residuals are documented accepted intent: an owner dying between its owned write and its own exit, and a hung old-build owner resuming during the one legacy upgrade window. - The pre-generation lock-holding claim shape keeps defer-or-reclaim behavior through a legacy shim: a live identity-verified stuck owner is retired via TERM (with a queued TERM sufficient when the owner is stopped) before its lock is removed, an unverified or identityless pid is never signalled but never blocks a proven-abandoned reclaim, and the lock's identity evidence is grafted into the ledger (mtime-preserving) so pid-reuse protection survives the lock. - The guard reads the same predicates for recovery ownership and its terminal fail-open (which re-checks for a live open claim under the held locks before committing the attended alarm), with ledger reads anchored to line 1 so the identity line can never confuse them. Behavioral regression coverage exercises all three edge classes through the real hook and guard - a live open claim defers with no lock held, a stuck claim is superseded and the home re-arms, and an end-to-end run with a genuinely hung owner shows a concurrent firing deferring promptly mid-arm, a later firing superseding the stuck owner, and the superseded owner exiting silently without a second translation - plus the identityless/reused-pid loopholes, the superseded-owner arm boundary, and the legacy TERM, SIGSTOP, and signal-free reclaim paths. * no-mistakes(review): Refuse auto-arm commits when notice marker creation fails * no-mistakes(review): Make episode reset atomic with generation ownership * no-mistakes(document): Update auto-arm generation and commit documentation
…n unproved merge (kunchenguid#3064) * fix(pr): verify GitHub merge outcome * no-mistakes(review): Captain, fixed forge-only merge verification, queue guidance, metadata propagation * no-mistakes(document): Correct forge-specific merge documentation * no-mistakes(review): Captain: forge-only queue fix, focused tests pass * no-mistakes(review): Captain: suppress closed-state guidance and prove parent regression * no-mistakes(review): Captain: remove history proof; retain executable regressions * no-mistakes(document): Clarify GitHub recording timing in architecture docs * no-mistakes(document): Clarify outcome-aware PR merge recording documentation * no-mistakes: apply CI fixes * Revert "no-mistakes: apply CI fixes" This reverts commit c326cfa. The automatic CI repair round removed the up-front `gh` prerequisite check while keeping the `gh` dependency: `bin/fm-pr-merge.sh` still calls `gh api graphql` for the outcome read and `gh api` for the branch-rules read. That left the same hard requirement without the clear named error, and review immediately raised a new finding for exactly the failure the check prevents - `gh-axi pr merge` landing the merge while the follow-up read fails, so the PR metadata is never recorded. The check is also symmetric with the GitLab arm directly above it, which already refuses up front when `glab` or `jq` is missing, on the stated principle that a missing tool should be a named prerequisite rather than a merge that is armed and then refused for an unexplained reason. The workflows this round was chasing sit at `action_required` because this is a fork pull request; no code change can turn them green. * fix(pr): keep PR bookkeeping when a merge outcome read fails On the GitHub path a merge call that returned success was followed by `github_read_outcome || exit 1`, so a transient API failure, rate limit, or network blip during the read dropped out of the script before `record_pr_metadata` ever ran. The merge could have landed while `pr=` went unrecorded and the merge poll was never armed - bookkeeping lost on a real merge. The failure path just above already recorded metadata before exiting, so the error path was more careful than the success one. Record the PR before that refusal. Recording arms the later merge poll and is not a success claim, which is the same reasoning that keeps `record_pr_metadata` on the gh-axi failure path. The refusal itself is unchanged: exit stays non-zero and the message still names the concrete observed state. Metadata is withheld only when the read succeeds and proves the pull request neither merged nor queued. Pin it with a case that stubs `gh api graphql` into failure after a successful `gh-axi pr merge`, asserting both the non-zero exit and the recorded metadata. * no-mistakes(review): Aggregate queue rules and report conflicts explicitly * fix(pr): keep the merge abstraction reachable and its bookkeeping intact Two holes remained in the outcome-verified GitHub merge path, both on installations where gh-axi is present but gh is not. The verification preflight refused before bin/fm-pr-merge.sh ever reached the configured gh-axi merge abstraction, so an installation without gh could no longer merge at all. gh-axi now performs the merge unconditionally and the queue-aware gh read became an optional enrichment: with gh on PATH its GraphQL view still separates merged from queued, and without gh the gh-axi view still proves a landed merge while every outcome it cannot prove refuses. The PR metadata recording sat behind the outcome read, so a merge that landed before that read failed lost pr= and its merge poll. Recording now happens once, before either forge call, which arms the poll without claiming a landed outcome and leaves teardown a PR identity to verify against no matter how the read ends. Rebasing onto main also restored the durable merge-outcome reporting and the GitLab landed-state confirmation that the conflict resolution dropped. Tests pin each fix through the executable interface: the merge abstraction is reached and verified with gh absent, a failed fallback read keeps its bookkeeping, and a mock that snapshots the task meta during the forge call proves pr= is recorded before the merge can land. * no-mistakes(review): fix(pr): de-dup queue methods, fall back on failed gh read, refresh contracts * no-mistakes(review): fix(pr): quote forge output and explain armed auto-merge on refusal * no-mistakes(review): fix(pr): claim auto-merge armed only when the forge accepted it * no-mistakes(review): fix(pr): tell the operator what each GitHub refusal could not observe * no-mistakes(review): fix(pr): gate every forge-acceptance claim on a successful merge * no-mistakes(document): align merge docs with verified GitHub outcome contract
* fix(pi): stop reporting one merge to the captain twice The supervision branch's captain-outcome note told main, unconditionally, that the note "is not your own earlier output" and to relay it now. When main had already reported the same event, that assertion was false and the order turned the correct response - saying nothing new - into a mechanical re-report, so the captain saw one merge reported twice in 16 seconds. Two independent changes, both needed: - The relay instruction is now conditional. It still names itself as a supervision outcome so main cannot mistake it for its own earlier answer (the silent loss that instruction exists to prevent), and it now lets main stay quiet about an outcome it has already given the captain. - The merge case is closed at its source rather than left to that judgment. One merge reaches a home on two independent paths by design - main's own permanently main-owned merge poll, and the branch's task-local status wake - and main's captain-facing text only reaches the branch's mirror at main's turn end, so the branch can escalate before it could possibly see the captain was already told. bin/fm-pr-merge-notified.sh answers that question from bin/fm-pr-lib.sh's canonical merge-notification marker, so the answer holds regardless of mirror timing. A captain outcome naming an already-published merge is delivered as the ordinary rendered note instead of opening a follow-up turn: still appended, still visible, still recorded with the verdict the branch decided, minus the wasted turn. Any error, timeout, or unreadable state relays the outcome. A duplicate announces itself; a lost outcome does not. Regression coverage drives the real delivery path in both directions: a new outcome must still reach the captain in exactly one follow-up turn even beside an unrelated published merge, and an already-published merge must open no second turn while a different PR in the same task still does. The merge path's real producer and this new consumer are exercised end to end in tests/fm-pr-merge.test.sh. Pi-only by construction: the delivery path lives in .pi/extensions, so no other harness loads it, and the new script only reads existing markers. * no-mistakes(review): Document accepted latest-marker suppression residual * no-mistakes(review): Recheck ownership before merge outcome delivery * no-mistakes(document): Document merge-outcome suppression exception * refactor(pi): drop the source-level merge suppression, keep the envelope fix The captain reviewed this branch and judged the source-level duplicate suppression overly complicated for the problem it solved, and asked for the change to be reduced to the envelope wording alone. Remove the mergeIntoMain downgrade path, bin/fm-pr-merge-notified.sh, and every test and document that existed only for it. What remains is the conditional captain-outcome instruction: main is told to stay quiet about an outcome it has already reported and to relay anything else, which covers the duplicate without a second mechanism. The silent-loss protection is untouched - the note is still typed, self-describing, and delivered as one invisible follow-up turn - and the behavioral tests still assert that, now requiring both halves of the conditional instruction. * no-mistakes(ci): Clarified in code comments and owned documentation that this is intentionally an M1-only, model-facing conditional relay fix—not source-level suppression—addressing Greptile’s mistaken scope expectation without changing runtime behavior. Net diff remains 3 files and 27 insertions. Verified with fm-pi-branch-extension tests, fm-lint, doc audience check, and git diff --check; all passed * no-mistakes(ci): Strengthened the runtime delivery test to verify the captain outcome retains its required self-description and outcome text. Verified with `bash tests/fm-pi-branch-extension.test.sh`, `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `git diff --check`; all passed. The outer pipeline can now commit and attest the new head
* fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership
…unchenguid#3211) * fix(pi): surface requested supervision outcomes * no-mistakes(review): Mirror in-flight captain requests before branch dispatch * no-mistakes(review): Exercise real branch ownership and main outcome access * no-mistakes(review): Preserve request tails and align verdict guidance * no-mistakes(review): Preserve complete current captain requests * no-mistakes(review): Require visible requested outcomes and realistic classification * no-mistakes(document): Align supervision outcome documentation * no-mistakes(ci): Fixed Greptile’s runtime-ordering finding. The extension now stages Pi’s authoritative `before_agent_start` prompt before SessionManager persistence and suppresses the later duplicate entry. Updated docs and behavioral regression to reproduce real Pi ordering and verify each prompt is mirrored exactly once. Passed branch-extension tests, supervision tests, strict Pi typecheck, full lint, and diff checks * no-mistakes(review): Use canonical operational input classification * no-mistakes(review): Filter legacy operational inputs canonically * no-mistakes(document): Clarify captain request mirroring boundary * no-mistakes(ci): Fixed the CI time-boundary failure in tests/fm-public-followup.test.sh by pinning its clock, including context-registry setup. This prevents follow-up fixtures from expiring based on wall time. Verified the full regression suite passes, project-owned lint passes, and git diff checks are clean * no-mistakes(document): Clarify captain-visible supervision outcome documentation
…#3210) * feat(bin): per-home remote transport lanes with cancellation, bounded send, and closed stdin All remote commands for every home on one host used to serialize through one single-job-at-a-time worker on one shared queue: a timed-out caller abandoned a staged job that kept running, retries convoyed behind it, fm-send's remote leg had no time bound, and staging captured the caller's stdin to EOF so any fm-on.sh caller with an open stdin wedged staging indefinitely. - The worker now serves one lane per staged home: same-home jobs run strictly FIFO in a new staging-sequence order while different homes run concurrently, each lane as its own top-level worker process (a backgrounded subshell does not reliably reap dead children, so a zombie group leader kept a finished command's process group signalable). Long-poll preemption is lane-scoped. - A caller that disconnects or times out cancels its job: the entrypoint marks the record on any post-staging exit and probes its parent so a dead ssh channel cancels without a signal; the worker skips cancelled queued jobs, terminates a running cancelled job's process group, and reaps the record. - fm-send's remote leg is bounded by FM_SEND_REMOTE_BUDGET (default 30s) and a bound hit exits through the existing unconfirmed-delivery contract, which stays idempotent because the remote enqueue deduplicates. - fm-on.sh defaults the remote command's stdin to /dev/null; the three payload callers pass the new --stdin flag. Abandoned .stage.* litter is age-reaped. - The job execution deadline no longer loses up to a second to clock truncation. * no-mistakes(review): Protect live stages and validate send budgets early * no-mistakes(review): Preserve sequence lock ownership during stale recovery * no-mistakes(review): Allocate job sequences at publication boundary * no-mistakes(review): Bound remote keys and extend stale lock recovery * no-mistakes(document): Document bounded remote transport behavior * no-mistakes(lint): Suppress intentional deferred-expansion lint warning * no-mistakes(ci): Fixed stale sequence-lock recovery by reconciling the counter against published job records before allocating the next sequence, preventing duplicate sequences and same-home FIFO violations. Added a behavioral regression test reproducing displacement after publication and verifying execution order. Passed fm-remote-transport-lanes.test.sh, fm-remote-job.test.sh, fm-lint.sh, and git diff --check * no-mistakes(review): Use atomic sequence claims and lossless lane keys * no-mistakes(review): Recover regressed sequence hints and rate-limit claim reaping * no-mistakes(review): Restrict worker heartbeats to serving loop * no-mistakes(review): Verify supervisor identity before lane recovery signals * no-mistakes(review): Verify tracked lane and claim owner identities * no-mistakes(document): Clarify remote lane and transport contracts * no-mistakes(ci): Fixed the CI time-boundary failure by pinning fm-public-followup tests to a deterministic clock, including context-registry setup. Verified tests/fm-public-followup.test.sh, tests/fm-remote-transport-lanes.test.sh, shellcheck, and git diff --check * no-mistakes(review): Preserve assigned lane ownership of queued jobs * no-mistakes(review): Reserve homes owned by foreign queued lanes * no-mistakes(review): Preserve completed results during crash recovery * no-mistakes(review): Harden claim cleanup, expiry, and cancellation races * no-mistakes(review): Verify process groups and reap abandoned results * no-mistakes(review): Stop leaderless groups and reap cancelled publications * no-mistakes(document): Correct remote transport lifecycle documentation * no-mistakes(lint): Quote done state comparisons for ShellCheck
* fix(tests): make the changed-file map select per script and stabilize a budget flake
The changed-file map's bin/ fallback resolved a direct test reference to that
test's whole FAMILY. bin/fm-push-transition-lib.sh is named by exactly one
real-Herdr E2E, so a one-line change to it selected all 12 real-herdr-gated
scripts, including a 341s presentation E2E with no dependency on it.
Resolve direct test references per script, and keep resolving consumer bin/
scripts through the curated map so recorded family-level coupling survives.
Also fix a load-sensitive flake: the tool-update budget deadline is whole-second
granular, so a test budget of 1 left headroom anywhere in (0, 1] seconds and the
first budget check could already read as exhausted.
* feat(bin): make suite wall clock a result and let a family's concurrency be proven
--max-wall-ms fails a run whose wall clock exceeds the caller's budget, after
reporting the per-script results. A suite that stays green while outgrowing its
caller's invocation budget is the regression that got an agent killed mid-run
and retried invisibly, so duration has to be a result rather than a log note.
--pool on the isolation-proof harness runs the same concurrent proof over a
whole family, so 'is this family safe to parallelize?' is answered by a command
instead of a guess. Measured watcher-wake-lock and refused it: 3 of 18 scripts
fail under concurrency on wall-clock assertions about reaching the next poll.
* perf(bin): schedule the changed suite concurrently, longest first
The watcher-wake-lock family is proven concurrent-safe (two clean runs, 18
candidates, 0 failures at 4 workers; docs/fm-test-isolation-proof.md), so
--changed now schedules its proven-concurrent scripts with bounded parallelism
and runs any unproven remainder serially afterwards, never beside them.
Concurrent runs are ordered longest-hint-first. Workers are handed scripts in
order, so alphabetical order started the 193s fm-watch-triage last and stranded
it running alone: 395s wall against a 205s balanced four-worker sum.
An explicit --jobs keeps its strict refusal, so every CI lane is unchanged.
* fix(bin): bound a hung test instead of letting it hang the suite
tests/fm-calm-pi-extension.test.sh was observed running 17+ minutes against a
464ms recorded hint, and the suite had no per-script bound to stop it. An
unbounded suite is precisely what silently outruns a caller's invocation budget,
and --max-wall-ms is evaluated after the run so it cannot end one that never
finishes.
--per-script-timeout-secs terminates a script that outruns it and records exit
124, so the run still completes, accounts for the script, and fails. The
auto-concurrent --changed path applies 900s, far above the slowest real script
(the 341s Herdr presentation E2E), so it only ever converts a hang.
* no-mistakes(review): Enforce safe concurrency and descendant timeouts
* no-mistakes(review): Validate empty runs and isolation proof pools
* no-mistakes(review): Measure selection time in wall budget
* no-mistakes(review): Reap interrupted workers and bound finalization
* no-mistakes(review): Contain shutdown descendants and watchdog finalization
* no-mistakes(review): Honor remaining budget and close launch races
* no-mistakes(review): Restore timeout helper and simplify runner cleanup
* no-mistakes(review): Record isolation pool admission metadata
* no-mistakes(review): Bound Chrome reap and scope proof admission
* no-mistakes(review): Align proof scheduling and preserve budget summaries
* no-mistakes(review): Remove unreliable finalization watchdog
* no-mistakes(review): Freeze budget duration and enforce admission caps
* no-mistakes(document): Refresh test runner concurrency documentation
* no-mistakes(lint): Fix ShellCheck findings in test runner scripts
* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; `--changed --jobs auto` explicitly opts into bounded concurrency and the automatic hang timeout. Updated documentation and added behavioral coverage proving serial default behavior, explicit concurrent scheduling, and refusal of `--jobs auto` outside `--changed`. Verified with `bash tests/fm-test-run.test.sh`, `bin/fm-lint.sh`, and `git diff --check`
* no-mistakes(review): Restore automatic changed-suite concurrency and timeout
* no-mistakes(review): Correct changed-suite contributor guidance
* no-mistakes(review): Reject gate-skipped isolation proofs
* no-mistakes(review): Correct automatic concurrency evidence
* no-mistakes(review): Isolate nested runner process groups
* no-mistakes(review): Remove unreliable signal cleanup machinery
* no-mistakes(test): Narrow changed-suite selection to executable contract owners
* no-mistakes(document): Document isolation proof skip and artifact semantics
* no-mistakes(ci): Fixed Greptile’s concurrency-consent finding. `--changed` now remains serial by default; bounded concurrency requires explicit `--jobs auto`. Updated behavioral coverage, contributor guidance, and isolation-proof commands accordingly. Verified with `tests/fm-test-run.test.sh`, `bin/fm-doc-audience-check.sh`, `bin/fm-lint.sh`, Bash syntax checks, and `git diff --check`; all passed
* no-mistakes(review): Restore plain changed-suite automatic concurrency
* no-mistakes(review): Record resolved changed-suite worker count
* fix(bin): keep a runner change selecting its whole curated family
A pipeline fix round narrowed the curated changed-file map so bin/fm-test-run.sh
and bin/fm-test-isolation-proof.sh selected only their own two contract tests,
and the documentation surfaces only the audience test. That cut this branch's
own changed selection from 33 scripts to 5.
The runner executes every pure-contract-unit script, so its contract test
passing proves its logic is right, not that the suite it drives still runs.
Narrowing it also makes any wall-clock claim about the changed suite trivially
true by not running the work.
Only the unmapped bin/* grep fallback resolves per script; curated mappings keep
their recorded family coupling.
* perf(bin): admit the pure-contract-unit family to bounded concurrency
A runner-file change selects pure-contract-unit, so that family decides the
changed suite's wall clock. With only watcher-wake-lock admitted, 14 of its 33
selected scripts fell to the serial tail and the selection measured 327.3s
against a 300s budget: the concurrent group was 19 scripts totalling 273.4s
while the tail alone was 215.7s.
bin/fm-test-isolation-proof.sh --pool pure-contract-unit --jobs 4 passes twice,
32 candidates, 0 failures, so the family is admitted on recorded evidence.
Full 33-script plain --changed: 327.3s -> 181.8s / 178.5s / 172.7s, 0 failures,
inside a 300000ms budget. Also states the per-script guard's derivation.
* no-mistakes(review): Align contract-unit concurrency cap with recorded proof
* no-mistakes(document): Record final changed-suite performance evidence
* fix(bin): keep an empty changed selection clean on stock macOS Bash
Under set -u, bash 3.2 treats "${arr[@]}" on an EMPTY array as an
unbound-variable error, while bash 4.4+ makes it a harmless no-op. The
concurrency work removed the early exit for an empty selection, so execution
fell through to the unguarded existence loop: on stock /bin/bash 3.2.57 a
contributor who changes only documentation and runs --changed got
bin/fm-test-run.sh: line 1713: SCRIPTS[@]: unbound variable
with exit 1 and no summary, instead of a clean total=0 pass.
Restore the early exit, and guard every remaining array expansion reachable
with an empty selection. The reported duration is real elapsed invocation
time rather than a hardcoded zero, so a selection phase that outran
--max-wall-ms still fails.
Verified on this host with /bin/bash 3.2.57: exit 1 with the unbound-variable
error before, exit 0 with FM_TEST_SUMMARY total=0 after.
* no-mistakes(document): Document shell-bound changed-suite performance
---------
Co-authored-by: Kun Chen <kun-1@kunchenguid.com>
* feat(bin): publish per-home summary ledger * no-mistakes(review): Bound and schedule home summary publication * no-mistakes(review): Prove recurring watcher summary refresh cadence * no-mistakes(review): Bound refresh workers and publish durable spawns * no-mistakes(review): Fix atomic kill process-group coverage * no-mistakes(review): Bound state initialization within refresh timeout * no-mistakes(document): Document recurring bounded home-summary publication * no-mistakes(review): Bound and log all best-effort refresh failures * no-mistakes(review): Harden cadence and timeout regression coverage * no-mistakes(document): Document home-summary runtime tuning * no-mistakes(lint): Fix direct exit-code check in refresh test * no-mistakes(ci): Fixed remote secondmate retirement recreating the deleted home: teardown now skips side-band summary refresh when its overridden state directory was removed. Verified with remote lifecycle E2E, teardown tests, home-summary tests, ShellCheck, and git diff checks * no-mistakes(document): Clarify atomic home-summary publication guarantee
* fix(pi): gate first call on startup context * no-mistakes(document): Correct Pi startup prerequisite verification date * no-mistakes(review): Captain, fix startup process-group retirement after leader exit * no-mistakes(review): Captain, release reload exit listeners on shutdown * no-mistakes(review): Captain, complete startup exit lifecycle ownership * no-mistakes(review): Captain, release empty startup process-group ownership promptly * no-mistakes(review): Captain, supervise startup ownership and restore failure fallback * no-mistakes(review): Captain, restore live Pi supervisor execution * no-mistakes(document): docs: clarify Pi startup prerequisite delivery
* fix(pi): restore 0.84.4 adapter compatibility * no-mistakes(review): Restore Pi collapsed and expanded outcome parity * no-mistakes(review): Preserve Pi stock previews through capability probing * no-mistakes(document): Document Pi 0.84.4 renderer compatibility
…nchenguid#3273) * fix(bin): keep home-summary publication bounded and off the watcher beat A home whose tasks had accumulated ordinary status history could not publish state/home-summary.json at all, and every attempt starved the watcher's liveness beacon while it failed silently. The producer's per-task open-decision fold spent tens of milliseconds per status line on a bash 3.2 global bracket-class substitution used only as a blank-line guard. On a real home that made the whole ledger producer take minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every attempt and never completed. Replace that guard with an equivalent case glob in the one fold owner, which both the whole-file and cursor-backed folds use. Bound each per-task current-state read in the snapshot with FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh, whose dead-peer detection deliberately never kills a slow-but-alive remote command, so nothing else bounded it. Detach the watcher's two publication triggers from the poll loop. The loop owns the beacon that fm-guard.sh reads as proof supervision is alive, and an inline publication put up to a full publication deadline between two beacon touches. A single in-flight publication is tracked so a slow one cannot accumulate clones. Report a repeatedly failing publication at session start. Publication stays deliberately non-fatal to its caller, so the existing bounded home-local failure record is now surfaced as a HOME_SUMMARY bootstrap line once the ledger is absent or stale and failures have been recorded since. * no-mistakes(review): Preserve home-summary failure attempt ordering * no-mistakes(review): Enforce durable home-summary single-flight and ordering * no-mistakes(review): Derive failure ordering from publication boundaries * no-mistakes(review): Restore best-effort failure logging and publication scoping * no-mistakes(review): Make ordering regression sensitive to one failure * no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance
…henguid#3268) * fix(supervision): classify the appended status span, not the last line An actionable project update could be classified as routine and absorbed, so a worker that raised a decision, hit a blocker, failed, or finished stalled silently with the captain never told. Trigger, mask, symptom. A worker appends a captain-relevant event (`needs-decision`, `blocked`, `failed`, `done`). Any later routine append - a `working:` progress note - lands before the supervisor classifies the batch; the watcher's 30s signal-grace linger exists precisely to coalesce a status write with the same turn's turn-end, so this window is ordinary rather than rare. Both supervisors then asked "is the LAST line captain-relevant?", read the routine line, and absorbed the wake. The `.seen-*` suppressor advanced either way, so nothing ever re-read the event. When the crew was also provably working, the no-verb fallback absorbed it too, which is why the event disappeared completely instead of surfacing late. Reproduced end to end against a real watcher before any change: with the trailing `working:` append the watcher never exits and the wake queue stays empty; with that one line removed - the smallest counterfactual - the same `needs-decision` surfaces and queues. The away-mode daemon's `classify_signal` returns `self|routine signal` for a `blocked:` event under the same mask, which is the worse case because no captain is present to notice. The proven path was already in the tree: `status_open_decisions` fixed this exact masking for the durable decision fold, and its header states the rule - reading an append-only event log last-event-wins cannot represent an earlier event that a later unrelated line moved past. The classification path was never migrated to that read model. That is the earliest divergence, and the fix is to migrate it rather than to special-case the symptom. `status_span_first_actionable` in bin/fm-classify-lib.sh is the new single owner: it reads the bytes at or after a caller-supplied position and returns the first still-live captain-relevant event. Each supervisor supplies its own position, because the always-on watcher and the away-mode daemon classify the same stream independently and must not share one cursor: the watcher reads the size already recorded in its `.seen-*` signature (no new state) and its `.hb-surfaced-<task>` backstop marker, and the daemon its `.subsuper-seen-status-<task>` marker. Those two markers held the escalated line and now hold the escalated-through byte offset, which also removes a second defect in the same code - content dedup silently swallowed a genuinely new event whose text repeated an older one. An absent, malformed, or past-the-end position reads the whole log, so uncertainty surfaces events rather than losing them, and a marker an older build wrote as a status line reads that way too. Status logs are only ever appended to, including across a reused task id, so a recorded position keeps its meaning. A `needs-decision`/`blocked` event in the span is retired only when the whole-file fold proves its key closed; `status_open_decisions` stays the sole owner of that rule, so same-key reopening and reserved-key namespaces need no second implementation here. Every other captain-relevant event is terminal and always actionable. Both backstops now walk every status log instead of only those whose last line looks captain-relevant, because the event a backstop most needs to catch is exactly one a later append has moved past. That leaves `scan_captain_relevant_statuses` with no callers, and it is removed rather than left as a working copy of the defective read model. Regression coverage exercises the classifier and both supervisors through their own interfaces: the masked decision, the captain-reported release/install completion followed by cleanup chatter, and the away-mode blocker all surface; a routine append after an already-classified event stays absorbed, so the fix does not convert ordinary progress into wakes; and the heartbeat backstop catches a masked event the per-wake path missed. The end-to-end watcher tests drive a real fm-watch.sh with the crew reported as provably working, which is the configuration that made the original stall silent. Two further claims in the supplied RCA are deliberately not patched here. "Repeated operational recoveries produced all-clear replies despite known actions" is downstream of this same cause, not an independent contributor: an all-clear reply is the documented response when the specific event needs no action, so a classification that wrongly reported "no action" produces it, and correcting the classification removes it. "The project was subjected to validation requirements outside its accepted path" is delivery-mode selection, which AGENTS.md section 7 owns; no code changed here touches it, so it is out of scope. Harness and backend axes were inspected rather than assumed: nothing in this path reads a vendor-emitted signal. The status log's format and append protocol are Firstmate's own and identical for every harness, and no runtime backend reads or writes `.status` files (`bin/backends/*` contain no reference to them). The surrounding triage's only backend touchpoints - pane capture and the authoritative crew-state read - are unchanged. No live-harness guard applies and no per-harness verification record changes. Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and `bin/fm-test-run.sh --changed --base origin/main`. * no-mistakes(review): Prevent status races and surface classification failures * no-mistakes(review): Surface unreadable signals and preserve AFK endpoints * no-mistakes(review): Route stale wakes through captured span verdicts * no-mistakes(review): Retire supervision offsets with reused task state * no-mistakes(review): Bind status offsets and preserve live decision origins * no-mistakes(review): Strengthen status identity with verified birth time * no-mistakes(review): Skip turn-end markers during status classification * no-mistakes(review): Preserve status presentation with platform-strength identities * no-mistakes(review): Retain failed wakes and advance routine checkpoints * no-mistakes(review): Surface all events and retain unreadable wakes * no-mistakes(review): Treat absent status logs as successful empty spans * no-mistakes(review): Bound repeated classification failures with durable receipts * revert(supervision): drop the failure-receipt and durable-retry machinery Captain-authorized revert to the minimal fix. Review rounds added a durable failure-receipt store and wake-retention-on-failure to bound repeated classification failures. That machinery grew larger than the fix it protected and kept producing its own defects: an unreadable log still looped forever because the always-on watcher never consulted the receipt, and the receipt was persisted before its diagnostic was durably queued, so a crash in between swallowed the alarm outright. Those two defects go away with the code that contained them rather than being repaired. Removed: the failure-receipt path, fingerprint, record and clear helpers and their retirement bookkeeping; the retention of a durable wake when classification fails; and the error-propagation plumbing in both supervisors that existed only to drive them. Kept, because it is the accepted fix rather than the declined machinery: span classification of the events appended since a supervisor last looked, in both supervisors and both backstops; reporting every actionable event in a span and committing a position only through what was reported; naming the live opening of a reopened decision; treating an absent log as ordinary and an unreadable one as worth reporting; the non-.status filter; and the platform-strength identity that guards a position commit without failing a read. Replacement behavior for a log that cannot be classified: report it once, do NOT advance the classification position so the content is classified from where it stopped once readable, and DO advance the wake signature so the report is bounded to one per distinct file state. Reporting and reading are different acts: telling the captain about a log is not the same as having read it, and only the latter may move a classification position. The residual risk is explicit and accepted: there is no guaranteed automatic retry inside a crash-mid-read window, and the locked session-start replay of the durable queue covers it. That rationale is recorded at mark_escalated_seen so a future reader does not reintroduce the retry as a "missing" guarantee. Also fixes lint failures that arrived with the review-fix commits and were never caught because the run never reached its lint step: an unfollowable conditional source directive, a second unquoted-expansion site left after a call was split across lines, cleanup of the file being read inside its own read loop (restructured to one post-loop teardown rather than three in-loop copies), stub functions in tests that are invoked indirectly, and a test local left unused when its assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch, so these were introduced here. Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue, wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures). * no-mistakes(review): Correct classification failure contract documentation * no-mistakes(review): Bound unreadable status reports without skipping classification * no-mistakes(review): Preserve escalation markers when buffering fails * no-mistakes(review): Detect permission recovery without advancing classification * no-mistakes(document): Document status span classification contract * no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed * no-mistakes(review): Escalate blockers while preserving declared-wait cadence * no-mistakes(review): Clarify actionable events override wait self-handling * no-mistakes(review): Surface rejected decisions and dangling status links * no-mistakes(document): Document reserved-key reconciliation classification * no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check * no-mistakes(document): Correct away-mode classification documentation
…#3289) * docs: split harness adapter operations reference * no-mistakes(review): Fix harness adapter routing and ownership contracts * no-mistakes(review): Prune duplicate harness adapter ownership prose * no-mistakes(review): Fix default effort routing and Grok max semantics * no-mistakes(review): Remove source-only routing test and duplicate semantics * no-mistakes(review): Add local harness adapter instruction evaluation * no-mistakes(review): Fix harness evaluation gating and change mapping * no-mistakes(test): Captain, require explicit harness instruction evaluator model * no-mistakes(document): Fix harness adapter documentation references
* test(fixtures): share fake-toolchain and spawn-world builders Future tests can start from tests/fixtures.sh instead of copying stubs, and a no-mistakes version-floor bump is one constant rather than a multi-file edit. Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen, fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile. Left for opportunistic migration: remaining make_spawn_fakebin copies (trace-context, kimi, muse, backend), the make_stubs send cluster, and the fake no-mistakes version banners in bootstrap/session-start/secondmate suites. Did not touch tests/fm-pr-check-security.test.sh. * no-mistakes(review): Prevent fake SSH test from blocking on stdin * no-mistakes(document): Clarify shared fixture documentation * no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks * no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head * no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head
* feat(bin): retire completed PR-check migration machinery Every registered home already carried both completion markers, and no installer still creates pre-migration checks. Remove the one-time migrate script, its bootstrap/watch/teardown/docs surface, and migration-path tests without weakening live check-trust or PR-poll authentication. * no-mistakes(review): Restore live PR-check security coverage * no-mistakes(document): Refresh retired PR-check documentation * no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks
…3247) * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * feat(extensions): bind trusted external process-event adapters * no-mistakes(review): Enforce owner and remote-home conformance * no-mistakes(review): Enforce serialized remote extension package lifecycle * no-mistakes(review): Enforce identity-conditional extension retirement * no-mistakes(review): Serialize extension retirement and recover crash cuts * no-mistakes(review): Unify retirement worker and lifecycle lock ownership * no-mistakes(review): Harden extension lifecycle retirement serialization * no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries * no-mistakes(document): Clarify built-in-only captain answer routing * no-mistakes(lint): Captain: fix extension binding ShellCheck findings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Use isolated UID mapping for owner conformance * no-mistakes(review): Captain: remove forbidden CI ownership wrapper * no-mistakes(review): Serialize extension binding publication * no-mistakes(review): Document ordinary CI owner-fixture exclusion * no-mistakes(review): Quarantine orphaned handshake descendants * no-mistakes(test): Fix orphan attribution * no-mistakes(test): Harden process tracker baseline * no-mistakes(test): Harden detached descendant attribution * no-mistakes(test): Use exact invocation-group cleanup * no-mistakes(test): Bound remote conformance transport crossings * no-mistakes(test): Parallelize isolated extension conformance tests * no-mistakes(test): Lifecycle suite still exceeds deadline * no-mistakes(review): Split extension conformance and forward remote transfer input * no-mistakes(review): Forward malformed remote payloads through fm-on * no-mistakes(review): Bound extension coordinator failure cleanup * no-mistakes(test): Skip repeated orphan sweep in coordinator children * no-mistakes(test): Queue isolated extension sections through bounded workers * no-mistakes(test): Bound extension coordinator lane cleanup * no-mistakes(test): Split remote lifecycle coordinator sections * no-mistakes(test): Coordinator probes pass; aggregate deadline remains * no-mistakes(test): Launch extension sections concurrently * no-mistakes(test): Fix coordinator marker publication * no-mistakes(test): Stabilize extension binding coordinator timing * no-mistakes(lint): Fix extension binding ShellCheck warnings * fix(extensions): prove invocation cleanup before retirement * no-mistakes(review): Harden process-event inbox confinement * no-mistakes(review): Preserve legacy capture parity * no-mistakes(review): Protect external registry staging * no-mistakes(test): Stabilize bounded extension conformance aggregate * no-mistakes(document): Document external evidence confinement * no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh` * no-mistakes(review): Harden extension staging and lifecycle reservation * no-mistakes(review): Harden external staging and lifecycle reservations * no-mistakes(review): Wire capture helper into remote conformance * no-mistakes(review): Pin external capture handoff and signal failures * no-mistakes(review): Bind pinned capture authority to inherited descriptor * no-mistakes(review): Harden descriptor-bound capture authority * no-mistakes(review): Harden core capture reservation authority * no-mistakes(review): Harden capture reservation boundaries * no-mistakes(review): Harden capture reservations and cleanup * no-mistakes(review): Harden capture handoff and reservation cleanup * no-mistakes(review): Bind capture handoff to claim descriptors * no-mistakes(review): Release lifecycle locks after host crashes * no-mistakes(review): Pin reservation recovery to recorded state roots * no-mistakes(review): Reject control bytes in claim state roots * no-mistakes(test): Stabilize extension capture descriptor handoff * no-mistakes(document): Document extension capture authority boundary * no-mistakes(lint): Fix ShellCheck extension binding warnings * no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks * no-mistakes(document): Correct extension namespace creation timing * no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes A promoted scout used to receive a free-form placeholder instead of the mode-specific Definition of done a briefed ship worker gets, so it never saw the ask-user escalation rule or the --yes prohibition. That gap is the concrete reason one incident's worker drove validation with --yes and answered its own ask-user findings. - Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh so the two contracts cannot drift. - bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the scratch inventory, clean base, ship branch, and that Definition of done, and prints the fm-send.sh command that delivers it. - State the --yes ban as a prohibition rather than a preference, without claiming an enforcement the tool does not provide. - Cover both through the real promotion and brief paths in tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh. * no-mistakes(review): Publish promotion instructions before committing task state * no-mistakes(review): Supersede conflicting scout delivery rules after promotion * no-mistakes(review): Reject invalid promotion instruction destinations * no-mistakes(document): Align documentation with promotion delivery contracts * no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check * no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean * no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head * no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
) * fix(bin): present complete Lavish board feedback as structured output Give the Lavish adapter a read-only presentation so a handler sees every annotation and the session-ending tag=message as its own field, instead of grepping a truncated raw capture. * no-mistakes(review): Preserve unquoted messages and prioritize captain prose * no-mistakes(document): Document structured Lavish result reads * no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves Dispatch and completion each moved a task's physical record and its backlog row as two independently timed steps, so a crash or a forgotten follow-up could leave the two disagreeing: a record with no in-flight row, an in-flight row with no owner, or a finished task still shown in flight. Fold each backlog transition into the script that performs the physical change, under the per-task lock it already holds and before it reports success. Dispatch moves the item to In flight after publishing the task record and fails loudly, removing its provisional record, when that transition cannot land. Completion records an authoritative close and performs it before removing the record, so an interrupted cleanup can be finished later, and its closing message now confirms what already happened rather than instructing a future step. Add a same-home reconciliation sweep to session start so a home that was interrupted mid-transition settles its own books on restart, replaying a recorded close and restoring an in-flight row it already owns a worker for. It never reads or writes another home; the fleet snapshot and the cross-home nudge stay as backstops. Close records are validated before they are trusted: the file is read as raw bytes and rejected outright when it carries a NUL or other control byte, every field must be well formed and non-duplicated, the id must match the record it was found under, the data location must resolve inside this home, and each close argument must carry a permitted, well-formed value. Writer and reader share one validator so a record this home publishes always remains replayable, independent of locale. Homes configured for a manual backlog, and homes with no backlog at all, stay exempt and are unaffected. * no-mistakes(review): Remove stale bootstrap migration helper invocation * no-mistakes(review): Preserve pending closes and narrow signal deferral * no-mistakes(review): Record close before destructive teardown * no-mistakes(review): Refuse pending closes before creating resources * no-mistakes(review): Guard relaunches and preserve cleanup warnings * no-mistakes(review): Reject symlinked records and clarify cleanup guidance * no-mistakes(review): Align dispatch eligibility and protect close replay * no-mistakes(review): Unify exact task incarnation parsing * no-mistakes(review): Render resolved configured backlog path * no-mistakes(review): Harden transition path boundaries against symlinks * no-mistakes(review): Validate lifecycle state before resource actions * no-mistakes(review): Enforce transition tooling and continuous state locks * no-mistakes(review): Consolidate same-home lifecycle file boundaries * no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts * no-mistakes(review): Reject final-component lifecycle record symlinks * no-mistakes(document): Document lifecycle record path boundaries * no-mistakes(lint): Quote literal done tokens in atomicity tests * no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched * no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check` * no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks * fix(records): validate record bytes without an uncurated tool The byte validation added for close records and directory paths shelled out to od. The spawn and teardown lifecycle runs under a curated command set that deliberately excludes it, so on any restricted PATH the check could not run, the data directory read as unresolvable, and dispatch and cleanup refused - wedging the lifecycle rather than protecting it. An earlier attempt made the failing test pass by adding od to that curated set. That fixed the test to agree with the defect and quietly widened the contract the fixture exists to pin, so it is reverted here. Inspect the bytes with perl instead, which is already in the curated set and already used in this repo for the same portability reason. The emitted values are identical to od's, so the rejection semantics are unchanged: NUL and other control bytes are still refused, legitimate paths containing spaces or non-ASCII characters still round-trip, and the check stays independent of the process locale. The restricted-PATH teardown case now passes because the validator no longer needs od, not because the fixture was loosened. * no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication * no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342) * fix: publish promote and Relay meta rewrites through contained replace Bare mv still rewrote live task records in place, so a symlink meta could be followed to a target outside state/. Route those field rewrites through the shared publisher and drop the unused library aliases. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse dangling symlinks during X metadata clear * no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects * no-mistakes(review): Exercise dangling symlink refusal through clear helper --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877) * fix(watch): absorb a turn-end whose pane churned since the previous poll The watcher's "absorb a benign turn-end when the crew is provably working" triage was structurally unreachable for any harness whose semantic busy state has no verified source. crew_absorb_class only reports working for an actively running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can only answer unknown for such an adapter, so codex crewmates surfaced a signal wake at every turn boundary with nothing to act on - a full supervisor drain, inspect and acknowledge turn per worker turn, scaling with the number of workers in flight and drowning the wakes that matter in identical noise. Widen the proof rather than bound the wake rate. A wake carrying only bare turn-ended markers is now also benign when the task's pane content changed since the previous poll, compared against the same state/.hash-* marker the staleness backbone already records and already trusts as liveness. That evidence claims no harness semantics, so it fabricates no busy verdict an adapter has not earned, and it needs no adapter cooperation. Absorb stays evidence-driven in both directions. A wake naming any status file keeps the strict proof, every captain-relevant verb still surfaces immediately, and an unresolvable task, a missing prior hash, a failed or empty capture, or an unchanged pane all surface exactly as before. The absorb defers rather than swallows: a crew that has stopped renders nothing further, so its now-static pane surfaces through the staleness backbone within a poll or two. Bounding the surfacing rate instead would have suppressed genuinely stopped workers. The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which owns it, and costs one bounded capture reached only for a no-verb turn-end whose crew is not already provably working. * no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Captain, reject malformed pane-churn hashes * no-mistakes(document): Document pane-churn turn-end evidence * no-mistakes: apply CI fixes * fix(watch): gate and bound the pane-churn turn-end absorb Make the pane-churn form of positive work evidence opt-in per home and bound how long it may defer one endpoint's bare turn-ends. Absorbing a bare turn-end on pane churn is now reached only when the home creates config/turnend-churn-absorb. The other two proofs read a verdict the harness itself vouches for, while this one infers execution from rendered bytes, so widening the absorb is a home's choice rather than a default every fleet inherits. With the flag absent the predicate returns on its first line and triage is unchanged. Churn and pane staleness read the same pane, so neither can be the other's only backstop. A pane that renders continuously never presents the two consecutive identical hashes the staleness backbone needs, so an unbounded churn absorb left a worker that had genuinely stopped behind such a renderer with no path to surface at all. One endpoint's turn-ends may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS, tracked in state/.churn-since-*, after which the wake surfaces and the window restarts. The bound is evaluated before any .stale- state is touched, so a wake that surfaces there leaves the staleness backbone's own classification alone. Covers both with behavioral tests: the same churning fixture that absorbs with the flag surfaces and queues without it, and a spent deferral window surfaces and restarts. The four existing safety guards now run with the flag enabled so they keep proving their specific guard. * no-mistakes(review): Fail closed on invalid churn deferral state * no-mistakes(review): Validate persisted churn deadlines before arithmetic * no-mistakes(review): Make churn deadlines transactional and bounds safe * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Clarify pane-churn supervision documentation * no-mistakes(lint): Fix watcher arithmetic lint issues * no-mistakes: apply CI fixes * no-mistakes(document): Clarify pane-churn fail-closed documentation * fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194) * fix(bin): bind the live pipeline-owned run instead of a superseded failed row fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead of the LIVE replacement run: the live run's pipeline-owned lane head is not a git object in the task worktree, so head-equality attribution rejected it and the coarse runs-list fallback silently continued past the RUNNING row onto an older failed row whose head equalled the stale worktree HEAD. The home summary then flipped invalid and Bearings hid the home's live work (F10). Attribution precedence now follows the daemon's own identity: - An ACTIVE run for the task's branch binds without head equality while branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active); the pipeline owning the branch is itself the attribution. - A genuinely failed run with no later run on the branch still reports failed through the unchanged head-equality path - real failures are not hidden. - In the coarse runs scan, an unresolvable head is unknown attribution and stops the scan (fm_nm_head_resolvable) instead of falling through to an older row; a resolvable-but-mismatched head keeps the historical reused-branch skip. The exemption never applies to a terminal run and requires pipeline_owned specifically, both pinned by negative-control tests. Fixture shape verified against the live incident run's real axi status output. * no-mistakes(document): Updated run-attribution documentation ownership * no-mistakes(review): Captain, make watcher marker identities injective * no-mistakes(review): Captain, localize pane-churn collision guard * no-mistakes(review): Compose turn-end evidence per task from one snapshot * no-mistakes(review): Restore strict turn-end fallback guards * no-mistakes(document): Align pane-churn watcher documentation * no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement Agents were improvising rm of check files with unset STATE/ID, which wedges headless panes. Unregister validates the id and state directory first. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Refuse explicitly empty custom-check state overrides * no-mistakes(document): Document custom-check retirement safety contract --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221) * Add quota exhaustion detection and safe fallback helpers - bin/fm-procevent-quota.sh: generic procevent adapter that arms a recurring quota-axi --json poll and wakes firstmate when a tracked provider's effectivePercentRemaining drops below a threshold or its runway.status becomes exhausted_now. - bin/fm-quota-choose.sh: worker-side helper that picks the first ranked harness:model candidate with positive effectivePercentRemaining. - AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document the new helpers and the mid-task quota-exhaustion wake path. - tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON source. * no-mistakes(review): Fix quota polling and scope bounds * no-mistakes(review): Enforce safe default quota selection * no-mistakes(review): Handle decimal quota values safely * no-mistakes(review): Fail closed on invalid quota inputs * no-mistakes(review): Reject empty quota candidate segments * no-mistakes(review): Harden quota parsing and timeout ownership * no-mistakes(review): Reuse captured quota snapshots consistently * no-mistakes(review): Match quota using explicit candidate providers * no-mistakes(review): Centralize fail-closed quota schema validation * no-mistakes(review): Reject out-of-range quota percentages * no-mistakes(review): Validate quota runway status enum * no-mistakes(review): Tighten quota scope and status contracts * no-mistakes(review): Preserve unknown quota and exact product bounds * no-mistakes(review): Preserve provider-level unknown quota * no-mistakes(review): Reuse canonical verified harness validation * no-mistakes(document): Document mid-task quota handling * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(docs): restore default routing contract, keep quota helper optional Restore the AGENTS.md section 4 always-loaded routing paragraph the PR had deleted, so the standing TOON-first intake, spendPriority ranker, every-candidate accounting, and load-trigger contract stay exactly as before this PR. The mid-task quota wake is optional and must not alter default routing. Restore the quota-array-dispatch skill ownership line to section 4 as the always-loaded intake boundary owner; keep the worker-side helper section as an addition only, without rewiring ownership or load triggers to section 13. * fix(bin): use harness-keyed quota matching in optional helper Revert fm-quota-choose.sh from harness:provider:model tuples back to harness:model candidates with harness-keyed provider matching, per the resolved ask-user finding. The helper is optional; authoritative multi-provider routing (provider discovery from the harness catalog and quota matching by that explicit provider) stays owned by AGENTS.md section 4 and the quota-array-dispatch skill intake procedure, not the helper. Document the multi-provider limitation in the helper header and the quota-array-dispatch skill: the helper maps each harness to one primary provider family only, so a candidate whose established provider differs from that primary family is checked against the wrong quota row. Use it only when the brief fixed the candidate order and every candidate's provider is the harness's primary family. The helper still consumes one already-captured default-TOON or JSON snapshot via stdin or --snapshot and never calls quota-axi itself, so it selects from the same quota state as the intake. * no-mistakes(review): Fix Muse quota mapping and helper contract docs * no-mistakes(review): Reject known-empty quotas and map quota tests explicitly * no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse * no-mistakes(review): Fix quota retirement and dependent regression coverage * no-mistakes(review): Accept zero-row quota TOON snapshots * no-mistakes(review): Enforce quota semantics status consistency * no-mistakes(review): Veto dispatch on any exhausted applicable scope * no-mistakes(review): Record exhausted quota scope in wake details * no-mistakes(review): Fix quota help and control dependency coverage * no-mistakes(review): Decode quoted TOON fields and document quota wakes * no-mistakes(review): Validate zero-row TOON and map timeout coverage * no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes * no-mistakes(review): Validate complete nonzero TOON envelopes * no-mistakes(review): Accept producer-shaped quota TOON envelopes * no-mistakes(review): Support empty quota arrays and validate counted rows * no-mistakes(review): Harden TOON completion, scopes, and quoted fields * no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields * no-mistakes(review): Allow unknown headroom under known semantics * no-mistakes(review): Reject noncanonical quota identities * no-mistakes(review): Preserve empty quota polling and validate attention identities * no-mistakes(review): Reject noncanonical provider watches * no-mistakes(review): Validate all candidates before quota selection * no-mistakes(document): Correct quota helper safety documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated read preferred element text over prompt, so an annotate-and-comment item dropped the captain's words. Surface prompt as its own field. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Filter non-comment prompts from Lavish reader output * no-mistakes(document): Clarify Lavish comment presentation contract * no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure * no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure * fix(bin): always emit Lavish comments and use real annotation fixtures Stop inferring comment provenance from prompt==text. Real pure annotations have no prompt, so always-emit does not duplicate. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420) * Fix public-followup register crashing on empty lock arrays under bash 3.2. bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock. The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there. * no-mistakes(document): Document stock Bash registration coverage * no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5 * no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
Merge upstream/main (kunchenguid/firstmate) at a5f3cbe into this fork's main at b93d621, which was 25 commits behind and 3 ahead. A merge commit, not a rebase, so every home can fast-forward onto the result. Three files needed hand resolution. bin/fm-crew-state.sh, header: took upstream's wording, then corrected its ownership pointer. Upstream routed all of "branch, head, pipeline-custody, and newest-first" attribution to bin/fm-nm-run-lib.sh, but the newest-first selection order lives in this script's own nm_runs_status_for_branch. bin/fm-crew-state.sh, nm_runs_status_for_branch: kept both sides. Upstream's bca584a (kunchenguid#3194) and this fork's d82ae28 + 61ccb10 fix the same class of incident at different layers, so they are complementary. Upstream's pipeline-owned exemption on the rich `axi status` path is kept in full with its helpers and all five of its tests. The fork's coarse-scan rules are kept because kunchenguid#3194's coarse guard rests on a premise that does not hold here: it treats an unresolvable run head as unknown attribution, but firstmate task worktrees share one object store with the primary checkout and the pipeline publishes lane heads into it as refs/no-mistakes/sync/<run-id>, so a live run's advanced head IS resolvable (verified 2026-08-31 against the live fleet: 7 of 8 real run heads, the running row included). A resolvability guard therefore reads a live row as a proven mismatch and walks on to the older superseded row, which is the original false-failure incident. docs/architecture.md: took upstream's sentence naming bin/fm-nm-run-lib.sh as the attribution owner, and kept a corrected sentence for the coarse fallback, whose selection order is still owned in the script. bin/fm-nm-run-lib.sh carries a doc-only correction: fm_nm_head_resolvable is left in place but no longer has a caller, so its comment no longer instructs callers to stop the scan on an unresolvable head. Two new regressions pin why the coarse rules are required, and both fail against kunchenguid#3194's coarse shape: - test_coarse_live_row_with_resolvable_diverged_head_binds - test_coarse_active_row_binds_when_the_pane_is_idle One upstream test changed deliberately: test_coarse_unresolvable_active_row_never_falls_to_older_row -> test_coarse_active_row_binds_and_never_falls_to_older_row. Its safety property is unchanged and still asserted; only the attribution expectation moved from "stop and let the pane answer" to "the branch's newest row binds".
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merges
upstream/main(kunchenguid/firstmate) ata5f3cbeinto this fork'smainatb93d621, which was 25 commits behind and 3 ahead.This is a merge commit, not a rebase: first parent is fork
main(b93d621, including the two CI fixes from #3), second parent isupstream/main(a5f3cbe), so every home can fast-forward onto the result. After the merge the branch is 0 commits behind upstream.Upstream commits merged (25)
10b93b2fix(bin): recover Claude auto-arm from hung claims (fix(bin): recover Claude auto-arm from hung claims kunchenguid/firstmate#3156)7ee0c19fix(bin): verify the real GitHub merge outcome instead of reporting an unproved merge (fix(bin): verify the real GitHub merge outcome instead of reporting an unproved merge kunchenguid/firstmate#3064)4f89f5bfix(pi): prevent duplicate captain outcome reports (fix(pi): prevent duplicate captain outcome reports kunchenguid/firstmate#3184)bca584afix(bin): prioritize active pipeline-owned crew runs (fix(bin): prioritize active pipeline-owned crew runs kunchenguid/firstmate#3194)c651b59fix(pi): surface requested outcomes without replaying fleet events (fix(pi): surface requested outcomes without replaying fleet events kunchenguid/firstmate#3211)1fd7ea2feat(bin): add concurrent bounded remote transport lanes (feat(bin): add concurrent bounded remote transport lanes kunchenguid/firstmate#3210)4207214fix(bin): accelerate and bound changed test runs (fix(bin): accelerate and bound changed test runs kunchenguid/firstmate#3250)a390659feat(bin): publish per-home summary ledgers (feat(bin): publish per-home summary ledgers kunchenguid/firstmate#3222)52b59a1fix(pi): gate first provider call on startup context (fix(pi): gate first provider call on startup context kunchenguid/firstmate#3158)f66be0ffix(pi): restore Pi 0.84.4 renderer compatibility (fix(pi): restore Pi 0.84.4 renderer compatibility kunchenguid/firstmate#3261)5f31097fix(bin): keep home-summary publication from starving supervision (fix(bin): keep home-summary publication from starving supervision kunchenguid/firstmate#3273)4eb587dfix(bin): prevent routine updates from hiding actionable status (fix(bin): prevent routine updates from hiding actionable status kunchenguid/firstmate#3268)c731c36docs(skills): split harness adapter operations reference (docs(skills): split harness adapter operations reference kunchenguid/firstmate#3289)0ace60atest: centralize shared shell fixtures (test: centralize shared shell fixtures kunchenguid/firstmate#3296)9e3df47refactor: retire legacy PR-check migration machinery (refactor: retire legacy PR-check migration machinery kunchenguid/firstmate#3299)1fbc7bbfeat(bin): add trusted process-event extension bindings (feat(bin): add trusted process-event extension bindings kunchenguid/firstmate#3247)c7fdef9fix(bin): deliver safety rules to promoted workers (fix(bin): deliver safety rules to promoted workers kunchenguid/firstmate#3269)debe4bffix(bin): present Lavish feedback as structured output (fix(bin): present Lavish feedback as structured output kunchenguid/firstmate#3321)1260adcfix: keep task records and backlog transitions atomic (fix: keep task records and backlog transitions atomic kunchenguid/firstmate#3322)d71f4b9fix(bin): contain promote and Relay metadata publishing (fix(bin): contain promote and Relay metadata publishing kunchenguid/firstmate#3342)a56a78afix(bin): absorb turn-end wakes during bounded pane churn (fix(bin): absorb turn-end wakes during bounded pane churn kunchenguid/firstmate#2877)0866a77fix(bin): safely unregister custom checks (fix(bin): safely unregister custom checks kunchenguid/firstmate#3369)4ad8cbarefactor(quota): extract mid-task polling and candidate selection into dedicated scripts (refactor(quota): extract mid-task polling and candidate selection into dedicated scripts kunchenguid/firstmate#3221)6c1d2dbfix: surface comments on Lavish annotations (fix: surface comments on Lavish annotations kunchenguid/firstmate#3371)a5f3cbefix: support first public-followup registration on Bash 3.2 (fix: support first public-followup registration on Bash 3.2 kunchenguid/firstmate#3420)Conflicts and how they were resolved
Three files needed hand resolution; everything else auto-merged.
1.
bin/fm-crew-state.sh- header comment (block 1)Took upstream's wording, then corrected its ownership pointer.
Upstream routed all of "branch, head, pipeline-custody, and newest-first" attribution to
bin/fm-nm-run-lib.sh, but the newest-first selection order lives in this script's ownnm_runs_status_for_branch.The merged header now names each owner separately, keeping the one-owner rule honest.
2.
bin/fm-crew-state.sh-nm_runs_status_for_branch(block 2)The substantive conflict. Resolution and evidence are in the next section.
3.
docs/architecture.mdTook upstream's sentence naming
bin/fm-nm-run-lib.shas the attribution owner, and kept a corrected version of our sentence describing the coarse fallback, since that fallback's selection order is still owned here rather than in the lib.bin/fm-nm-run-lib.shalso has a non-conflicting doc correction, explained below.The attribution-commit decision
Our two local commits and upstream's
bca584a(kunchenguid#3194) address the same class of incident - a superseded run row being reported while a live run for the branch is healthy - but at different layers, so they are complementary rather than redundant. Both were kept.d82ae28+61ccb10bca584a(kunchenguid#3194)no-mistakes runsfallbackaxi statusprimary pathbranch_sync.state=pipeline_ownedUpstream's primary-path exemption is a genuine new capability we did not have, and it is kept in full, along with its
fm_nm_run_is_pipeline_owned_active/fm_nm_branch_sync_state/fm_nm_run_is_activehelpers and all five of its tests.Why our coarse-scan commits were not dropped
kunchenguid#3194's coarse-scan guard rests on one premise: that a live run's lane head is not a resolvable git object in the task worktree, so an unresolvable head can be read as "unknown attribution" and stop the scan.
That premise does not hold for firstmate's own worktrees. Task worktrees share a single object store with the primary checkout, and the pipeline publishes lane heads into it as
refs/no-mistakes/sync/<run-id>through this repo's configuredno-mistakesremote. Verified against the live fleet on 2026-08-31:7 of the 8 real run heads, including the currently running row, resolve inside a task worktree. So a resolvability test reads a live run's advanced head as a proven mismatch, walks past it, and reports the older superseded row - the exact incident it is meant to stop.
The coarse path is not a rare branch either:
axi statusanswers with one globally active-or-most-recent run, so in a multi-lane fleet a task's own live run reaches attribution only through the coarse list (runs_on_current_branch: 0with 8 runs listed, observed above).The proof
Reverting only the coarse scan to kunchenguid#3194's shape and running the suite reproduces the original defect:
Three tests pin this and each one fails against kunchenguid#3194's coarse shape (checked by reverting the hunk and re-running, so none is vacuous):
test_coarse_live_row_with_resolvable_diverged_head_binds- new. The firstmate-shaped case: live row, head resolvable but diverged, idle pane. Asserts the superseded failed row does not surface and thatcrew_absorb_classisworking. Against fix(bin): prioritize active pipeline-owned crew runs kunchenguid/firstmate#3194's shape:state: failed.test_coarse_active_row_binds_when_the_pane_is_idle- new. Why binding beats "stop and let the pane answer": a crew waiting on its pipeline is idle, so the pane fallback degrades to the crew's own possibly-stale status line. Against fix(bin): prioritize active pipeline-owned crew runs kunchenguid/firstmate#3194's shape:source: status-log.test_live_run_outranks_superseded_row_on_worktree_sha- fromd82ae28, unchanged.One upstream test changed deliberately
test_coarse_unresolvable_active_row_never_falls_to_older_row->test_coarse_active_row_binds_and_never_falls_to_older_row.Its safety property is unchanged and still asserted: the older failed row must not surface. What changed is the attribution expectation - the branch's newest row now binds (
source: run-step) instead of stopping so the pane can answer (source: pane). Both shapes reportworkingon that fixture because its pane is busy; the two new tests above pin why binding is the stronger answer once the pane is idle, which is the normal posture while a pipeline validates. The rename and the reasoning are recorded in the test's own comment.fm_nm_head_resolvableLeft in
bin/fm-nm-run-lib.shrather than deleted - removing an upstream primitive is beyond what this merge requires. Its call site is gone, so its comment was corrected: it previously instructed callers to stop the scan on an unresolvable head, which is no longer whatnm_runs_status_for_branchdoes and would be misleading guidance for a future caller in this repo. The comment now records the shared-object-store evidence to weigh before reintroducing it as an attribution gate.Note on upstream, not changed here
fm_nm_head_resolvable's "unresolvable head" premise appears to be environment-specific rather than universal: it holds where the gate mirror's objects never reach the task worktree, and not where the worktree shares an object store with a checkout that has theno-mistakesremote configured. Upstream may want to revisit the coarse-scan guard on that basis. No upstream behavior was changed in this PR beyond the merge resolution itself.Gates run on the merge result
bin/fm-lint.sh: ShellCheck 0.11.0 (pinned), full extended analysis; actionlint 1.7.12 (pinned), 3 workflow files valid.bin/fm-doc-audience-check.sh:ok surfaces=89 local_links=300.tests/fm-crew-state.test.sh: all pass, including the five upstream fix(bin): prioritize active pipeline-owned crew runs kunchenguid/firstmate#3194 cases and the fork's coarse-scan cases.Local behaviour-suite failures, and why they are not from this merge
Six scripts fail on the development machine (macOS, stock Bash 3.2, with other lanes live). Every one of them behaves identically on pristine
upstream/maina5f3cbe- same exit status, script for script - so none is introduced here (0= passes, listed as controls):upstream/maina5f3cbeis green on all twelve upstream CI lanes (Lint, Repo invariants, Behavior portable parallel 1-2, portable serial 1-4, Herdr, Test coverage guard, Behavior timing aggregate, Stock macOS Bash snapshot compatibility), so these are local-environment results rather than code defects. This PR's own CI result is the authority.Worth noting: this merge improves
tests/fm-public-followup.test.shlocally. Onb93d621it fails at the first assertion (could not register the public commitment- the stock Bash 3.2 registration bug); with upstream'sa5f3cbefix merged it gets 34 assertions further before hitting an unrelated pre-existing path-construction failure in the rechain case.The gated
real-herdr-gatedfamily was not run: it drives real Herdr lifecycle behaviour, which this task was not scaffolded to do.