Conversation
…spatch (#1) * feat(harness): add the cline crewmate/scout adapter with ClinePass dispatch Add Cline CLI 3.0.62 as a verified crewmate/scout harness following the agy pattern: ancestry detection on the native .cline process, a launch-then-send TUI launch, per-task .cline/hooks busy/turn-end wiring under a new cline-hook busy source, Escape interrupt, /exit, and control-plane tables. Wire the ClinePass open-weights pool into the crew-dispatch example and document the known composer-empty gap (placeholder luminance above the shared ghost ceiling) with a tmux live guard as the refresh command. * no-mistakes(review): fix(docs,quota): correct cline resume grouping and cline-pass family id * no-mistakes(document): docs: cover cline in tmux liveness/anchoring list and configuration.md secondmate-refusal note * no-mistakes(lint): {"summary": "lint: no code changes needed, fm-lint.sh passes with shellcheck on PATH"} * chore(gitignore): drop the stray .omc handoff artifact and ignore .omc/ The no-mistakes gate agent's Claude Code oh-my-claudecode plugin writes .omc/handoffs/last-session-end.md into the run worktree at session end, and a later pipeline step committed it into this branch. Remove the committed file and ignore .omc/ so a home-environment handoff artifact can never ride into a PR. --------- Co-authored-by: firstmate-worker <worker@local>
Two Claude subscriptions need to run concurrently across lanes without moving every claude spawn onto one account. --claude-config-dir picks the CLAUDE_CONFIG_DIR one claude spawn's pane resolves into: validated before any worktree or endpoint exists, recorded in the task's own meta, reused unchanged on --relaunch, and threaded through both the pre-launch trust registration and the launch's own environment so the two halves can never land in different stores. A spawn naming no seat is byte-identical to before.
…rop stray artifact
Count live crewmate, scout, and local secondmate lanes grouped by the billing provider a candidate actually draws on, expose the load for dispatch intake, and refuse a spawn that would push a provider past its configured cap. Provider identity comes from the resolved model string, not the harness name: a provider-qualified prefix or a model-id pattern decides the pool, and the harness table is only the fallback. That keeps two models on one pool counting together while a different pool stays separate. The mapping lives once in bin/fm-provider-lib.sh, whose fallback reuses the existing quota tables rather than restating them. A lane occupies a seat unless its recorded endpoint is provably dead or missing, so a cleared seat is visible before the next dispatch. The cap comes from providerCaps in config/crew-dispatch.json (per provider, else default), falling back to 4; bootstrap now rejects a malformed providerCaps instead of silently ignoring it. bin/fm-provider-load.sh prints the current per-provider used/cap for intake. Stranded-record detection stays with fm-lane-account-dead-records; this counter reads the current endpoint classifier.
Deliver bin/fm-hold-reverify.sh, an armed watcher check that re-checks each captain hold past an age threshold against shipped reality and reports it as dead, still_live, not_a_decision, or unestablishable - the reconciliation vocabulary captain-hold-lifecycle already owns. It reports only: it never calls answer and never closes or annotates a call, so only the captain's own words or an explicit evidence-backed reconciliation can resolve one. Dead is never inferred from absence or an unreadable source. Aged holds come from the canonical local backlog projection (fm-fleet-snapshot.sh --contribution-input); recorded pull requests are read through fm-pr-lib.sh. Each sweep writes a docket and prints one line only when the finding set changes, with a report record keyed on that set.
…rget
A firstmate home had no way to see that another home was already working a
shared external target, so the main home and a secondmate could both arm to
land the same PR with nothing to stop a double merge.
bin/fm-claim.sh records, releases, and inspects a work claim on a shared
external target - a PR, an issue id, or a declared file area. The store is a
machine-wide directory (FM_CLAIM_ROOT, default
${XDG_STATE_HOME:-$HOME/.local/state}/firstmate/claims), the sibling of the
existing process-event source claim root, because one owner per canonical
target cannot live inside a single home. Local homes share one filesystem; a
remote secondmate is a separate host and stays outside the mechanism.
Acquire is atomic and fails closed: a second home's live claim refuses rather
than racing. A claim is released on cleanup or reclaimed only when its holder
is provably gone (its home directory is absent, or its task record is absent
past FM_CLAIM_PENDING_GRACE). Any uncertainty keeps the claim.
bin/fm-spawn.sh --claim records the claim before any endpoint or task record
exists and refuses the spawn on conflict, recording the canonical keys on the
task as claims=; fm-teardown.sh releases them on cleanup. The flag is refused
on --secondmate, --relaunch, and a batch dispatch.
Tests: tests/fm-claim.test.sh drives the real CLI across two simulated homes
sharing one claim root, plus a real spawn that records its claim and a second
dispatch that is refused.
Merge authority was decided once at intake and never revisited, so a task dispatched yolo=on kept that authority even after firstmate held it for the captain, and the recorded authority and the merge path could disagree. Fold the captain-hold predicate into fm_merge_authority_resolve, the single owner of a task's standing merge authority, so a held task resolves to captain-hold (or hold-unreadable) whatever its yolo posture or away grants say. Remove bin/fm-pr-merge.sh's duplicate require_released_captain_hold and fold its refusal into the shared gate, and have bin/fm-merge-local.sh share the same predicate instead of repeating it.
A worker parked on a provider quota wall kept a live, painting harness while its turn could not advance, so every one of them read as working from its semantic busy record. The measured fleet incident had all of one provider's workers stalled at the same weekly limit while supervision saw a healthy fleet. Recognize the wall from the pane text the busy reader already inspects. The signal is built from two independent rendered families - a wall-shaped limit phrase and a scheduled retry/reset phrase - within the last few non-empty lines, so no single vendor string is load-bearing and ordinary worker prose does not match. A busy verdict over that wall reports `quota` instead of busy. fm-crew-state.sh surfaces it as its own `state: quota` rather than collapsing it into working or a declared pause, because a quota-killed worker cannot be relaunched in place; the recovery skill now states that preserve-and-replace under a new id is the path. The portable regression pins the logic and its divergence cases over synthetic transcripts. The live guard drives the real installed OpenCode TUI against a local 429 stub so its own retry modal renders with no model tokens spent, and proves the same task reads working before the wall and quota after it.
Teardown refused any record whose endpoint was already cleared, so a lane could never be retired once its window was gone and it kept occupying an in-flight row. Accept an explicit endpoint_cleared stamp as stronger agent-less evidence than a dead window, with no flag and no --force, while keeping the unlanded-work refusal unchanged. A projected Herdr teardown confirmed only the task pane was gone, so a workspace whose recorded pane vanished before its close survived for a restart to restore as a live agent in the primary clone. Remove the workspace's remaining panes through the same focus-preserving pane close and require the workspace gone before retiring the journal. A dead pane whose display redrew re-alarmed on every new hash, a supervision tax that grew with each dead lane. Absorb a redrawn dead display against the existing once-record, while a relaunched agent re-arms the incarnation and its own death still reports in full.
…human-read text
…e ship definition of done
…ot wedge a supervisor
… liveness sweep
…control exit works
…nnot race one target
… alongside the reliability batch # Conflicts: # bin/fm-bootstrap.sh # bin/fm-control-lib.sh # bin/fm-quota-choose.sh # docs/configuration.md # docs/examples/crew-dispatch.json
… fork merge Committed only to let a clean merge proceed without touching this work; not otherwise reviewed or authored by this session.
…s the preserved cline adapter # Conflicts: # AGENTS.md
Make the Fireworks DeepSeek dispatch lane usable: openhands is now a verified crewmate/scout harness with launch, readiness, busy, interrupt, and exit mechanics, so config/crew-dispatch.json no longer fails as an unverified adapter.
# Conflicts: # .agents/skills/harness-adapters/SKILL.md # AGENTS.md # bin/fm-agent-process-lib.sh # bin/fm-bootstrap.sh # bin/fm-busy-lib.sh # bin/fm-composer-lib.sh # bin/fm-control-lib.sh # bin/fm-harness.sh # bin/fm-spawn.sh # docs/agent-control.md # docs/architecture.md # docs/configuration.md # docs/trace-context.md # docs/verification/runtime-backends.md # tests/fm-control.test.sh
# Conflicts: # .gitignore # bin/fm-spawn.sh
* fix(bin): read aged holds via backlog-json, not contribution-input The hold re-verify sweep only needs canonical backlog rows; contribution-input also walks every task meta for merge-authority resolution, which took ~96s at this fleet size and always exceeded the five-second projection bound. Add fm-fleet-snapshot.sh --backlog-json for that narrower read and point the sweep at it so failures stay loud without raising the timeout. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): apply review decisions for aged hold re-verify sweep Sort aged captain holds oldest-first before the per-sweep cap, let FM_HOLD_REVERIFY_BUDGET_SECS govern the backlog projection bound, clamp forge probes to remaining budget, skip probes for predetermined not-a-decision rows, drop the classify subcommand, and document --backlog-json. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): use backlog_json output mode label for shellcheck Co-authored-by: Cursor <cursoragent@cursor.com> * docs: fold pipeline document-step hold-reverify pointers into fork tip Adds the toolbelt row and prose alignment from the failed run's document step without rebasing onto upstream main. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): correct hold-reverify docket contents description * chore: gitignore .omc session state and document in AGENTS.md Apply captain inbox 006 on the fork publication branch without rebasing onto upstream main. --------- Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: expect 19 snapshot/fleet-view tests under stock macOS Bash The fork's tests/fm-fleet-snapshot-view.test.sh carries test_large_payloads_compose_through_files, the regression for bin/fm-fleet-snapshot.sh composing payloads above 128KB through files instead of argv. It landed in d9356ca together with that snapshot change and passes under /bin/bash 3.2 on the macOS runner, which already counted 19 ok lines, so the hard-coded 18 in the macOS job was the only thing left behind. * docs: declare the cline-pass provider on the documented cline profiles The resolver refuses docs/examples/crew-dispatch.json with "profiles whose harness lacks one authoritative provider family require provider: cline", so the documented-example check in tests/fm-dispatch-resolve.test.sh has failed since the cline profiles were added to the example. The example is the wrong side. The resolver's provider is the quota-axi provider family it ranks candidates by, not the launch prefix cline reads from its model id, and the single-provider table in docs/configuration.md leaves cline out on purpose: the model prefix, not the harness, decides who bills the run, exactly as for pi and omp. The Pi default in the same example already declares provider: claude for that reason. Add provider: cline-pass to the four cline profiles, correct the one sentence in docs/configuration.md that claimed no field was needed, and give the test's canned Choice answer the fourth rule the example now has, since the resolver checks the answer against the full option set. * test: corrupt the claim record under test, not the first one find returns The corrupt-record check picked its victim with find | head -n 1 while three claims exist, so on a filesystem whose directory order differs from the author's it corrupted o/r#7 or repos/example#9 and then asked about owner/repo#11, whose intact record answered "held" with exit 0. CI's stdout showed exactly that line. Select the record by its documented key= line instead. The guard itself is intact: corrupting the right record returns exit 5 on the same inputs. * test: give the restart watchers a refresh bound a slow runner can meet In the watcher-restart section of tests/fm-home-summary-refresh.test.sh the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of reports 999999 for a missing ledger, so its detached refresh fires on every one-second poll. After the lock holder is killed that refresh steals the dead lock ahead of the test's idle-only refresh, which then returns without publishing, and with the section's FM_HOME_SUMMARY_TIMEOUT=2 a slow runner kills the watcher's attempt before it publishes. The next poll dies the same way and the ledger never appears, which is the "a dead publication lock wedged publication" failure in CI run 35683307194. Raise the three restart watchers' bound to 30 seconds, the deadline this file already uses for its accumulated-home publication. Under a 30% CPU quota the unchanged test fails on exactly that line and the changed one passes all 21 checks. The idle-only refresh, the dead-lock reclamation and the 10-second wait are untouched.
…wner (adopt kunchenguid#5993) (#7) * Fix reassigned teardown slot collisions (cherry picked from commit 115333b) * no-mistakes(document): Document claim-first pool-slot teardown behavior (cherry picked from commit dcbe51a) * no-mistakes(document): Confirm teardown documentation reflects slot ownership behavior (cherry picked from commit bdcedcf) * fix(teardown): avoid presentation-lock races after claim-first slot cleanup Reassigned-slot teardown with a dead Herdr husk no longer takes the shared presentation session lock, restoring the pre-5993 contention profile for stale records while keeping live-slot protection intact. Herdr presentation recovery spawns now wait up to 120s for that lock and release it on abort so concurrent cross-home recovery cannot fail the 5s try loop after a legitimate holder. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: QIanGua <15757826110@163.com> Co-authored-by: Cursor <cursoragent@cursor.com>
…in config (#9) * Fix OpenCode v2 worker launch to use --standalone and config model. OpenCode 2.x removed the interactive --model flag; carry the resolved model in OPENCODE_CONFIG_CONTENT and launch with --standalone so the model and permission block are honored off the shared service. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(bin): honor opencode v2 top-level model and gate --standalone Always write the resolved model as OPENCODE_CONFIG_CONTENT top-level model on 2.x, drop unverified agent.build variant JSON, and keep the 1.x --model launch shape when opencode --version reports major 1. Co-authored-by: Cursor <cursoragent@cursor.com> * test: stub opencode --version in shared spawn fakebin Spawn tests prepend a fakebin to PATH; fm-spawn now probes opencode --version for the v1/v2 launch gate, so every spawn fakebin must answer it. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)
* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it
A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.
The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.
Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.
The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.
Closes #3055
* no-mistakes(review): require an unanswered decision before deferring a parked gate
* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records
* test(watch): pass the pane hash wedge_timer_check now takes
Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.
* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate
* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling
* feat(watch): make the parked-gate wait deferral opt-in
The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.
config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.
It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.
The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.
* test(watch): pin that the away-silenced hold leaves the idle timer alone
The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.
* no-mistakes(review): document away-silence rationale, pin captured gate component
* no-mistakes(test): anchor gate row scan to the braced findings header
* no-mistakes(document): pin same-block gate row invariant in crew-state comment
* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)
* fix(control): let the owning seat reclaim a task whose endpoint is gone
A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.
`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.
- fm-spawn --relaunch accepts a positively proven `missing` and creates
one fresh endpoint in the recorded worktree; the record it already
republishes rebinds the task to it. A `dead` endpoint is still adopted
in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
relaunch transaction's stop step no longer dead-ends, and re-resolves
the endpoint from the record before verifying the replacement.
The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.
A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.
Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.
* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds
* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session
* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly
* no-mistakes(review): stop refusals and docs asserting unestablished causes
* no-mistakes(review): stop herdr fixture helper losing tmp-root registration
* no-mistakes(review): document workspace drift and absence-probe server residue
* no-mistakes(review): correct rebind limitation to its one reachable case
* no-mistakes(review): stop claiming reclaim leaves instructions untouched
* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage
* no-mistakes(rebase): read the staged launch file in the herdr fixture
Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.
Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(document): note reclaim placement in herdr and scripts inventories
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(bin): stamp status events with their emission time (#3764)
* test(status): reproduce missing event emission time
* wip(status): preserve optional event emission time
* test(status): document indirect clock stub invocation
* no-mistakes(review): Preserve historical status bytes during reply recovery
* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies
* no-mistakes(review): Preserve captain regex overrides for timestamped status events
* no-mistakes(document): Clarify status event timing and publication contracts
* no-mistakes(lint): Quote literal done to satisfy ShellCheck
* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed
* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags
* no-mistakes(test): Stamp Rovo spawn failures with emission time
* no-mistakes(document): Verify status event documentation
* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests
* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed
* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed
* no-mistakes(review): Stamp remote escalations at call sites, drop new flag
* no-mistakes(review): Accept stamped escalation and close lines in test assertions
* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes
* test(status): accept optional emission time in PR-provenance assertions
The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.
* no-mistakes(review): Accept stamped ready signal in PR fallback scrape
* no-mistakes(review): Drop relay flag, stamp parent events at call sites
* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture
* no-mistakes(review): Accept optional stamp in live cmux drift guard
* no-mistakes(review): Restore original test invocation order in two suites
* no-mistakes(review): Strip only well-formed numeric status time tags
* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc
* no-mistakes(review): Stamp agy spawn-failure status lines with event time
* fix(bin): normalize status event times in-shell and freeze the budget test clock
Two paths made a status event's emission time cost more than it should.
The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.
tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.
Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.
* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion
* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs
* test: fold emission-time snapshot coverage into the fixture case
Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.
* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch
* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape
* no-mistakes(review): strip undelimited at-tags; correct brief stamp header
* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness
* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles
* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb
* no-mistakes(review): read note and key past colon-bearing stamps
* test(status): keep inactive reconcile assertions stamp-tolerant
These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap
* no-mistakes(document): correct stale unstamped status-line spellings in docs
* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments
* no-mistakes(ci): rename subshell-local epoch in delivery-race stub
The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(bin): unify Lavish host and disconnect handling (#5060)
* fix: ship clean Lavish host fixes
* no-mistakes(review): Fix Lavish classifications and fail-closed host loading
* no-mistakes(review): Restore Lavish host state across retries and launches
* no-mistakes(review): Preserve destination Lavish host when configuration is absent
* no-mistakes(document): Document Lavish status and host guarantees
* feat: act on captain's away words during AFK supervision (#5076)
* feat(afk): make the captain's away words the whole mandate
Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.
The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.
Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.
* no-mistakes(review): carry the away read-back to the session verbatim
* no-mistakes(review): match the exact away-action marker in the return brief
* no-mistakes(review): refuse a words block truncated by a damaged line
* no-mistakes(document): Refresh away-role contract documentation
* fix(bin): render the remote charter's steering-inbox path host-local (#5049)
* fix(bin): render the remote charter's steering-inbox path host-local
A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.
Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.
Closes #5012
* no-mistakes(document): document remote charter's host-local steering inbox
* feat: route Lavish feedback directly to owning workers (#5099)
* feat(procevent): route worker-owned Lavish rounds
* no-mistakes(review): drop duplicate artifact field from task-owned registration
* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic
* no-mistakes(review): keep worker board owned until terminal round acknowledged
* no-mistakes(review): refuse every retirement of an open worker-owned round
* no-mistakes(review): use real lavish reply flag, isolate reply generations
* no-mistakes(review): drop .posted marker for best-effort reply posting
* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures
* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms
* no-mistakes(review): re-arm only to acknowledge an open round
* no-mistakes(review): conclude only a still-open terminal round
* no-mistakes(review): record the acknowledgement before retiring the board
* no-mistakes(review): retain the registration across a conclude, qualify terminal docs
* no-mistakes(document): Document worker-owned Lavish round lifecycle
* fix(bin): fit pull observation within the contribution poll budget (#5107)
* fix(bin): reserve contribution observation budget
* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)
* feat(bin): add idempotent inbox orders, receipts, replies, and readiness
Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.
* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict
* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json
Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.
* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor
* no-mistakes(document): Note read-only lock inspection in scripts inventory
* no-mistakes(lint): Pass missing id argument to malformed-reply test printf
---------
Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)
* fix(composer): stop a harness footer row from reading as a composer holding text
A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.
Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.
A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.
Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.
* no-mistakes(review): make composer footer-zone demotion shape-independent
* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan
* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100
---------
Co-authored-by: Koen Muller <koen@catapult.nl>
* feat(bin): append optional home-local include to briefs (#5115)
Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)
* fix(bin): stop misreading a no-run branch as an unreadable runs table
Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.
Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.
Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.
* fix: recovered same-branch inventory awk misreads empty result as unreadable
fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.
Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.
* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup
* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings
* no-mistakes(review): revert repo lookup to exact working_path match
* no-mistakes(document): note state-db inventory read under crew-state nm timeout
* fix(bin): require a non-draft pull request before a PR-based done report (#5141)
* fix(bin): require a non-draft pull request before a PR-based done report
A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.
The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.
Closes #4757
* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix: support quota-axi schema 6 snapshots (#4904)
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey
quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.
- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
accountKey required on every row and uniqueness on
provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
is the one join every consumer uses: schema 5 binds by provider alone,
schema 6 binds to the row keyed by the candidate's Pi lane, else the
provider's default row, else no row (unmeasured, never blocked, never
by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
the shared function; an expanded provider with no row for the
candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
with a schema 5 case on the same path; every new case fails on the
previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.
* no-mistakes(review): Fix native Codex quota and expanded provider watches
* no-mistakes(review): Align native Codex account matching across dispatch paths
* no-mistakes(document): Align quota documentation with account-aware snapshots
* no-mistakes(document): Align quota dispatch documentation with account matching
* fix(bin): keep CI lint and the quota watch test portable
- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
that source this library, so full-mode ShellCheck reported SC2034 on
the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
assertions used rg, which CI runners do not install, so the case
failed with 'rg: command not found' rather than on behavior; use grep
like the rest of the file.
* no-mistakes(document): Documented schema-version account-row compatibility
* test: fix Claude session-start drain live E2E (#5165)
* test: repair Claude live auto-arm regression
* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event
* no-mistakes(document): Consolidate Claude live verification references
* ci: pin the no-mistakes required check to v1.80.1 (#5195)
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
* fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174)
* fix: preserve Pi watcher ownership across session replacement
* no-mistakes(document): Scope Pi predecessor retention away from omp
* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
* fix(bin): allow cleanup of windowless legacy task records (#5236)
* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown
Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers
* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator
* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity
* no-mistakes(document): Clarify windowless teardown retry documentation
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: exempt kunchenguid from the no-mistakes required check (#5256)
* fix(bin): surface launches parked on an interactive prompt as not-started (#5250)
* fix: surface parked launch prompts as not started
* no-mistakes(document): docs: record launch-prompt busy backstop classification
* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* fix: record away posture immediately on /afk (#5260)
* feat(afk): make /afk itself the go with a same-turn record write
Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.
* no-mistakes(document): Refresh away-entry documentation evidence
* fix(bin): recognize passed-with-override as a passing outcome (#5294)
* fix(bin): map passed-with-override to done instead of unknown
no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.
Map passed-with-override to the same done/terminal handling as a
clean passed in both places.
* fix(document): Replace stale outcome mapping with authoritative pointer
* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: clean up workers after their pull requests land (#5317)
* fix: close landed workers from supervision in both postures and at return
During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.
- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
inactive-outcome, or heartbeat row on a done task with a merged PR, as the
moment to claim the lease and run bin/fm-teardown.sh with no flags; a
refusal is reported, never forced or worked around. Add teardown to the
handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
records only (a live task record whose recorded PR carries the
merge-notification marker), between could-not-fix and handled, without
holding the gate; the afk skill's return step closes each listed task
through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
in fm-afk-return through the real marker writer.
* no-mistakes(document): Document landed-task cleanup ownership
* fix: surface green no-mistakes PRs awaiting merge (#5327)
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring
A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.
Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.
Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.
Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.
* no-mistakes(review): read the full ci log when checking checks-green
* no-mistakes(review): correct stale ci log tail wording in docs
* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling route from board session (#5334)
* fix: derive Lavish polling server from its board session
* no-mistakes(document): Document session-derived Lavish polling
* no-mistakes(document): Correct Lavish routing verification claims
* fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900)
* fix(bin): ignore vanished state scratch files on secondmate relaunch
Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.
Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes #4765.
* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
* fix(bin): stop each keyed answer from re-waking this home (#4907)
* fix(bin): treat home-owned status closes as already read
Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.
* no-mistakes(review): Keep owned closes in unread status; lock ledger writes
* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake
* no-mistakes(review): Require real owned growth before ledger marks status seen
* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS
* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test
* no-mistakes(review): Align ledger docs and scope ledger to wake path only
* no-mistakes(review): Restore stranded historical-annotation test comment to its function
* no-mistakes(review): Retire the home-appends lock alongside its ledger
* no-mistakes(document): Note ledger's lock-helper dependency in classify library
* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions
* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger
* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger
* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
* fix: deliver failed public follow-ups with updated AXI floors (#5350)
* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest
Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.
tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.
Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.
* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test
* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures
* no-mistakes(document): Document failed public-followup delivery behavior
* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
* fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878)
* fix(bin): refuse ship done: when the named head lives only in the worker copy
A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.
* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff
Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.
* no-mistakes(review): Bind named-head gate to recorded PR and forge heads
* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries
* no-mistakes(review): Align worker done wording, test mapping, pending-retry test
* no-mistakes(test): Raise watcher test time limit to stop load flake
* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage
* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict
* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery
* no-mistakes(document): Name fm-crew-state among named-head gate callers
* fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204)
* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm
A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.
* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
* test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335)
The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue #3793).
The original 0.25s window after confirmation was widened to 80 polls in
#3837, which left the same race at a larger size.
Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.
A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix: stop watchers reliably during blocked polls (#5362)
* fix(bin): let one TERM always stop the watcher on bash 5.2
Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.
HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.
The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.
* no-mistakes(document): Clarify watcher stop-signal documentation
* fix: submit stuck inbox doorbells instead of skipping them (#5374)
* fix(bin): submit our own stuck doorbell instead of skipping every later ring
* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit
* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat: add opt-in fleet activity ledger (#5375)
* feat(bin): add the opt-in fleet activity ledger
Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.
* no-mistakes(review): Record task.status text verbatim after the first colon
* no-mistakes(document): Clarify fleet ledger status and setup documentation
* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
* fix: validate public follow-up deliverables and wake on rejection (#5352)
* fix(bin): format, validate, and surface public-followup deliverables
brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.
* no-mistakes(review): refuse emits missing a required deliverable in both destinations
* no-mistakes(review): require promised deliverables and keep rejections recoverable
* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule
* no-mistakes(review): keep a rejection wake whose line cannot be read
* no-mistakes(review): key emit-time rules on the promise, not the outcome
* no-mistakes(review): bound deliverable keys and values as tasks-axi does
* no-mistakes(review): state rejection wakes as at-least-once and pin it
* no-mistakes(review): enforce the promised contract tasks-axi holds at emit
* no-mistakes(review): stop inferring a staged promise from its outcome
* no-mistakes(document): Refresh public follow-up documentation
* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks
* Revert unrelated CI auto-fix edits to the watcher and bearings board test
The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.
* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing
* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
* feat: add Devin CLI crewmate and scout adapter (#5380)
* Add verified Devin CLI worker adapter
* no-mistakes(review): Drop Devin resolver refusal and launch marker
* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs
* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts
* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals
* fix(control): never pair Devin interrupt presses on an idle agent
A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.
* fix(devin): disable Claude hook import and commit attribution for workers
The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.
* test(devin): extend live guard and record Herdr and revert-picker evidence
The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.
* no-mistakes(document): Correct Devin documentation links and lifecycle guidance
---------
Co-authored-by: Denis Beliaev <battler73@yandex.ru>
* fix(bin): recognize passed-with-skips as a passing outcome (#5322)
fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.
Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.
* fix(bin): refuse unavailable backend adapters before sourcing (#5382)
* fix: refuse missing backend adapter before source
* no-mistakes(review): Gate backend precheck under stock Bash
* no-mistakes(document): Clarify adapter precheck docs
* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
* test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338)
* fix(test): repair tmux liveness and calm follow-up loaded_off regressions
Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.
tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).
tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.
Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.
These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.
Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded
* fix(test): give wake-queue observation checkpoints the alerting ceiling
tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:
not ok - a foreign queue with no progress did not alert: check: rearm-resurface
not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface
Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.
* no-mistakes(document): docs: correct export-DOM Chrome render root cause
* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup
* chore: re-trigger fork workflow approval for triage
---------
Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
* fix: keep watcher status classification bounded to new log spans (#5383)
* fix(bin): classify a status span without re-folding the whole log
A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.
Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.
Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.
Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.
A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).
* no-mistakes(document): Clarify span classification and watcher regression coverage
* test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381)
* test: fix watcher timing flakes in fm-pr-check-security
The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.
The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.
The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.
The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.
* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed
* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"
This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6.
* feat: record fleet status immediately and emit PR-ready events (#5385)
* feat: record task.pr_ready in the fleet ledger when a task PR is registered
* feat: record worker status lines in the fleet ledger as they are written
* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger
* no-mistakes(review): Resolve relative config override before embedding in worker command
* no-mistakes(document): Clarify fleet ledger status capture timing
* test: synchronize foreign queue stall checks with watcher progress (#5386)
* test: synchronize foreign secondmate stall legs on the watcher's recorded observation
Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.
Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.
* no-mistakes(review): Wait for full stall reset before stopping watcher leg
* test: isolate the bearings render fixture from the shared Lavish store (#5391)
The listener resolves its server from that store before it polls. Without a session for this bo…
…ession start (#12) * feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes #3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment * fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up #4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(bin): stamp status events with their emission time (#3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The #4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead of carrying a second copy of the rule in awk. The copies had already drifted: the shell side stripped tags from lines with no colon, which the awk rule left whole, so a colonless line could be mistaken for one already recorded. One definition, checked against the awk rule it replaces over the edge cases and a 4000-line fuzz. tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In hang mode the poll set DEADLINE to the real now plus a one-second budget, and when the second ticked before the first forge call the loop broke without ever calling gh: forge/calls was never written and the assertion failed reading a missing file. Freezing the clock in both modes removes the dependence on wall time; the bounded call is still cut by the real timeout, so the observation the test asserts still starts. Emission time stays optional on new status records, and legacy or malformed lines keep an unknown age. * no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion * no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs * test: fold emission-time snapshot coverage into the fixture case Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer touches workflows. Keep every emission-time assertion by folding it into test_fixture_snapshot_json. * no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch * no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape * no-mistakes(review): strip undelimited at-tags; correct brief stamp header * no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness * no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles * no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb * no-mistakes(review): read note and key past colon-bearing stamps * test(status): keep inactive reconcile assertions stamp-tolerant These two oracles were made stamp-tolerant while resolving one of the branch's merges from main. The rebase drops merge commits, so that adaptation was lost and both assertions went back to matching an exact substring that a stamped line no longer contains: the tag lands before the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...". Strip a well-formed tag before matching, as the branch's other oracles do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap * no-mistakes(document): correct stale unstamped status-line spellings in docs * no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments * no-mistakes(ci): rename subshell-local epoch in delivery-race stub The serialization test overrides fm_pending_reply_mark_delivered inside a (..) subshell. Its `epoch` local collided with the same name in status_line_at_epoch/status_stamp_line, which this branch added and this suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and failed Lint 2. The stub already prefixes its other locals with `pending_` for the same reason; `epoch` was the leftover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(bin): unify Lavish host and disconnect handling (#5060) * fix: ship clean Lavish host fixes * no-mistakes(review): Fix Lavish classifications and fail-closed host loading * no-mistakes(review): Restore Lavish host state across retries and launches * no-mistakes(review): Preserve destination Lavish host when configuration is absent * no-mistakes(document): Document Lavish status and host guarantees * feat: act on captain's away words during AFK supervision (#5076) * feat(afk): make the captain's away words the whole mandate Retire the clause fields, verb list, never-set scan, refused records, and the per-task merge-grant list from the away-posture record. The record is now version 2: the captain's words verbatim plus expected return, spend cap, and reach line; a version 1 record still validates, reads, and archives so a live away window is never broken by the upgrade. The supervision branch reads the words at the tail of every wake and acts on them by its own judgment through the guarded scripts under standing authority, never by analogy, holding for the return on doubt, and opens each such outcome summary with "per your away instructions:" so the return brief can render the words beside the session's account. While the record exists any green merge runs under away authority (ledger tag "away"); red merges, --allow-red, asynchronous and queued merges, and local-only landing stay refused. The branch may file a backlog item the words explicitly call for before dispatching it under the spend cap. Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and fm-pr-merge.sh as commands: version 2 written, version 1 read, retired flags and subcommands refused by name, green merges landing under the record, red and waived-red refused, the record lock still closing the authority-read window, and the Pi away tail carrying the words. * no-mistakes(review): carry the away read-back to the session verbatim * no-mistakes(review): match the exact away-action marker in the return brief * no-mistakes(review): refuse a words block truncated by a damaged line * no-mistakes(document): Refresh away-role contract documentation * fix(bin): render the remote charter's steering-inbox path host-local (#5049) * fix(bin): render the remote charter's steering-inbox path host-local A freshly provisioned remote secondmate read a parent-home absolute steering-inbox path in its charter - a location that exists on no route - and spent its first turn discovering the gap and filing a blocked decision for what was a render defect. The seed's remote-copy rewrite now maps the inbox to the route's host-local parent-route inbox, exactly as it already maps the reply-log path, so every mention - bare path, listing, and handled/ acknowledgement - lands host-local. Both rewrites also become plain assignments, because a quoted substitution nested inside a double-quoted printf argument leaks literal quotes into the replacement text on stock macOS bash. The lifecycle suite pins the corrected render both directions against the real seed, provisioning, and delivery route, sharing one fixture value between the render truth and the delivery truth. Closes #5012 * no-mistakes(document): document remote charter's host-local steering inbox * feat: route Lavish feedback directly to owning workers (#5099) * feat(procevent): route worker-owned Lavish rounds * no-mistakes(review): drop duplicate artifact field from task-owned registration * no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic * no-mistakes(review): keep worker board owned until terminal round acknowledged * no-mistakes(review): refuse every retirement of an open worker-owned round * no-mistakes(review): use real lavish reply flag, isolate reply generations * no-mistakes(review): drop .posted marker for best-effort reply posting * no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures * no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms * no-mistakes(review): re-arm only to acknowledge an open round * no-mistakes(review): conclude only a still-open terminal round * no-mistakes(review): record the acknowledgement before retiring the board * no-mistakes(review): retain the registration across a conclude, qualify terminal docs * no-mistakes(document): Document worker-owned Lavish round lifecycle * fix(bin): fit pull observation within the contribution poll budget (#5107) * fix(bin): reserve contribution observation budget * no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget * feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103) * feat(bin): add idempotent inbox orders, receipts, replies, and readiness Let a caller supply a request id when publishing a captain inbox note so a retry returns the original note instead of creating a second one, including across the crash window between save and wake announcement. Separate saved from announced so a failed wake is repairable without enqueueing again. Add bounded receipts JSON with omission disclosure, a durable primary reply against a note id, and a read-only readiness projection that can say unknown instead of inferring liveness from a lock file. * no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict * fix(bin): resolve ready from lock-holder ancestry; drop lock status --json Remove the extra JSON surface from fm-lock.sh so its human status still always exits zero. Have the readiness projection classify the inspected home from the lock-holder pid via fm-harness.sh ancestry, with an explicit FM_SUPERVISION_MODEL still winning and an unknown model when there is no holder. Prove the yes path when that ancestry names a known harness. * no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor * no-mistakes(document): Note read-only lock inspection in scripts inventory * no-mistakes(lint): Pass missing id argument to malformed-reply test printf --------- Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com> * fix(bin): stop harness footer rows below a composer from reading as pending text (#5118) * fix(composer): stop a harness footer row from reading as a composer holding text A harness draws its own furniture below the composer - a user statusLine, a permission-mode hint - and the cursorless "bottom-most shape wins" rule looks exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text everywhere else, so a statusLine opening with `→` was selected as a bare composer, swallowed the hint row beneath it as wrapped input, and answered `pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first doorbell and every retry were skipped and the worker never saw the steer. Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on Herdr 0.8.0 had genuinely empty composers and every one of them was refused. A separator pair that closed over a bare agent-glyph row is a proven composer container, so the contiguous non-blank rows below its closing rule are that composer's footer and are no longer composer candidates. The demotion is bounded by all three of its own preconditions: a blank row ends the zone, a pair that closed over no glyph row demotes nothing, and a shape with no separator pair at all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same composer, including a stray SGR mouse report left by a click in the pane, still reads `pending`. Pinned by two portable regressions and by a new cursorless arm on the live composer-matrix guard, which re-reads each harness's already-proven-idle pane the way every non-tmux backend reads it and fails naming the harness and version when that read is `pending`. * no-mistakes(review): make composer footer-zone demotion shape-independent * no-mistakes(review): make footer-zone demotion refuse-only and drop rescan * no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100 --------- Co-authored-by: Koen Muller <koen@catapult.nl> * feat(bin): append optional home-local include to briefs (#5115) Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com> * fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114) * fix(bin): stop misreading a no-run branch as an unreadable runs table Defect: when `no-mistakes axi status`'s overview is truncated (a task's own branch has zero rows among the shown ones), fm_nm_select_run's Python fallback derived the repo identity for its direct SQLite query from a `repo: <path>` line it expected in the overview text. The real CLI never emits that line, truncated or not (see the genuine capture at tests/captures/no-mistakes-v1.70.1/overview.toon, which has only `count:`/`runs[...]:`), so the lookup always failed and reported "unreadable runs table" for a task that simply has no run on its branch. On a fleet with many concurrent runs, every idle-branch task hits the truncated-overview path routinely, so this fired every few minutes and drowned genuine unreadable/blocked verdicts in noise. Fix: derive the repo identity from the task worktree path instead, which is exactly the value `no-mistakes` records as a repo's `working_path` (confirmed against the existing capped-overview test fixtures, which already register repos by worktree path). A worktree path that is not absolute cannot be matched and still reads as unreadable rather than being guessed at. Also raise the reader's SQLite busy timeout from 1s to 30s so ordinary lock contention on a busy fleet cannot masquerade as an unreadable database. Safety: every other verdict byte-for-byte unchanged - the repo lookup still requires exactly one matching row (a genuinely corrupt or mismatched repos table still reports unreadable, per the existing `repo` failure-mode test), the branch query and row validation are untouched, and a zero-row result for the branch still flows through the same recursive re-parse that already turns an empty `runs[0]{...}` table into `absent`. Added a regression test (test_capped_overview_without_repo_line_and_no_runs_reports_absent) that reproduces the real overview shape - capped, zero rows for the task's branch, no `repo: ` line - and asserts the crew state falls through to the pane/busy verdict instead of reporting unknown or "unreadable". Full fm-crew-state.test.sh suite passes unchanged otherwise. * fix: recovered same-branch inventory awk misreads empty result as unreadable fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:` overview and re-runs it through the same awk selection pass. When that rebuilt inventory has zero rows for the branch, the row-matching loop never executes, so its counters (`seen`) stay at awk's uninitialized empty string while `expected` and `shown` are plain strings parsed from the header text. Comparing an uninitialized value against a non-numeric string uses string comparison, so "" != "0" is true, and the END block takes the "unreadable runs table" branch instead of falling through to the correct "absent" verdict for a branch with genuinely zero runs. Coerce the affected END comparisons with `+0` so they are always numeric, matching seen/expected/shown/total regardless of whether awk classified them as strings or numeric strings. A truncated or genuinely malformed inventory still differs numerically and still reports unreadable. * no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup * no-mistakes(review): match recorded repo path first, tolerate duplicate spellings * no-mistakes(review): revert repo lookup to exact working_path match * no-mistakes(document): note state-db inventory read under crew-state nm timeout * fix(bin): require a non-draft pull request before a PR-based done report (#5141) * fix(bin): require a non-draft pull request before a PR-based done report A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge. The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done. bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before. The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged. Closes #4757 * fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata * fix: support quota-axi schema 6 snapshots (#4904) * fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more than one account: every provider row carries an accountKey and one provider id may appear on several rows. fm_quota_json_valid accepted only schema 5 with unique provider ids, so fm-dispatch-resolve.sh, fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live snapshot and quota-informed dispatch was dead against the current tool. - bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with accountKey required on every row and uniqueness on provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ is the one join every consumer uses: schema 5 binds by provider alone, schema 6 binds to the row keyed by the candidate's Pi lane, else the provider's default row, else no row (unmeasured, never blocked, never by position or summed across accounts). - bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey column, and joins through the shared function. - bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through the shared function; an expanded provider with no row for the candidate's account is reported as such. - tests: schema 6 fixtures shaped like the real snapshot, each paired with a schema 5 case on the same path; every new case fails on the previous scripts and passes now. - docs: the two sentences naming the row join describe the schema 6 key. * no-mistakes(review): Fix native Codex quota and expanded provider watches * no-mistakes(review): Align native Codex account matching across dispatch paths * no-mistakes(document): Align quota documentation with account-aware snapshots * no-mistakes(document): Align quota dispatch documentation with account matching * fix(bin): keep CI lint and the quota watch test portable - bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts that source this library, so full-mode ShellCheck reported SC2034 on the assignment; mark it alongside the existing SC2016 disable. - tests/fm-procevent-quota.test.sh: the schema 6 provider-watch assertions used rg, which CI runners do not install, so the case failed with 'rg: command not found' rather than on behavior; use grep like the rest of the file. * no-mistakes(document): Documented schema-version account-row compatibility * test: fix Claude session-start drain live E2E (#5165) * test: repair Claude live auto-arm regression * no-mistakes(review): Assert SessionStart digest completeness within its hook_response event * no-mistakes(document): Consolidate Claude live verification references * ci: pin the no-mistakes required check to v1.80.1 (#5195) Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies. * fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174) * fix: preserve Pi watcher ownership across session replacement * no-mistakes(document): Scope Pi predecessor retention away from omp * no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified * fix(bin): allow cleanup of windowless legacy task records (#5236) * fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown Catch-up correctly refuses while a leftover task record has no status file. Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers * no-mistakes(review): Validate windowless leftover identity via shared endpoint validator * no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity * no-mistakes(document): Clarify windowless teardown retry documentation --------- Co-authored-by: Cursor <cursoragent@cursor.com> * ci: exempt kunchenguid from the no-mistakes required check (#5256) * fix(bin): surface launches parked on an interactive prompt as not-started (#5250) * fix: surface parked launch prompts as not started * no-mistakes(document): docs: record launch-prompt busy backstop classification * no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop * fix: record away posture immediately on /afk (#5260) * feat(afk): make /afk itself the go with a same-turn record write Collapse the propose-then-confirm away entry into one 'enter' step that writes state/.afk-contract immediately and prints the announcement and read-back after the record exists, never asking for a go. The retired propose, confirm, and --proposal inputs are refused by name, and a stale proposal left by an older version is removed rather than promoted. Refresh and replace semantics, verbatim words, the single writer, the never-set, and per-harness launch behavior are unchanged. * no-mistakes(document): Refresh away-entry documentation evidence * fix(bin): recognize passed-with-override as a passing outcome (#5294) * fix(bin): map passed-with-override to done instead of unknown no-mistakes' axi status emits outcome: passed-with-override for a run that finished with an explicitly approved Test or CI exception. Both bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's pre-teardown terminal-run check only matched the literal passed and checks-passed tokens, so this outcome fell through to unknown/parked and a finished worker awaiting merge kept getting re-alerted as stale, while an abort race during teardown could also leave a finished run misreported as still parked. Map passed-with-override to the same done/terminal handling as a clean passed in both places. * fix(document): Replace stale outcome mapping with authoritative pointer * fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor * fix: clean up workers after their pull requests land (#5317) * fix: close landed workers from supervision in both postures and at return During the 2026-09-22 away window every exemption worker whose pull request had merged was left sitting for nine hours. The supervision branch received the stale wake, the merge-landed check, and the hourly inactive-outcome row for each of them, ran the recovery playbook, found nothing to recover, and reported "no further action". The branch prompt granted ordinary teardown of a confirmed-landed task without ever naming the moment or the command, and the playbook has no landed exit, so the stale path ended at "nothing to recover". The return brief then listed only blockers, decisions, and the latest five routine outcomes, so the landed workers stayed invisible after the captain came back. - bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale, inactive-outcome, or heartbeat row on a done task with a merged PR, as the moment to claim the lease and run bin/fm-teardown.sh with no flags; a refusal is reported, never forced or worked around. Add teardown to the handling tool list. - stuck-crewmate-recovery: a landed worker is not a recovery case; point at the ordinary teardown owner for each actor. - bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable records only (a live task record whose recorded PR carries the merge-notification marker), between could-not-fix and handled, without holding the gate; the afk skill's return step closes each listed task through ordinary teardown once the check clears. - tests: pin the prompt rule in fm-branch-supervision and the brief section in fm-afk-return through the real marker writer. * no-mistakes(document): Document landed-task cleanup ownership * fix: surface green no-mistakes PRs awaiting merge (#5327) * fix(bin): surface a green no-mistakes PR still in ci merge monitoring A green PR could sit unreported because neither the worker nor the supervisor could observe checks-green while the ci step kept monitoring for the merge. Supervisor read: fm_nm_select_run's capped-overview inventory reader looked the repository up by the task worktree path, but no-mistakes registers a repository once by its main clone path and resolves every linked worktree to it, so on every task copy of a busy repo the lookup matched no row and each read reported "complete same-branch run inventory unreadable". Key the lookup on the overview's own top-level `repo:` line, which every axi release emits as the resolved working_path. Even with a readable run, the ci-log classifier treated "base branch advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs a checks state only when it changes and a base advance does not clear readiness, so a green PR read as still validating for as long as main kept advancing. Stop treating that line as a marker, matching no-mistakes' own ci-log parser, and name the run's PR URL in the held-for-merge reading so the existing inactive-outcome path can act on it without a worker report. Worker contract: `axi status` never reports checks-passed while the ci step monitors for merge, so the definition of done no longer makes a status poll the wait for the next gate or outcome; the drive call's own return is the green signal, reattached with `no-mistakes axi run` after a bounded return. * no-mistakes(review): read the full ci log when checking checks-green * no-mistakes(review): correct stale ci log tail wording in docs * no-mistakes(document): Document checks-green supervisor fallback * fix: derive Lavish polling route from board session (#5334) * fix: derive Lavish polling server from its board session * no-mistakes(document): Document session-derived Lavish polling * no-mistakes(document): Correct Lavish routing verification claims * fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900) * fix(bin): ignore vanished state scratch files on secondmate relaunch Relaunch refused when find(1) exited non-zero while listing a secondmate home's state directory. A live watcher can delete scratch files between readdir and processing, which is not evidence that child *.meta records are unreadable. Prove the directory is listable from its mode and keep the existing readable-meta loop as the child-record guarantee. Fixes #4765. * no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root * fix(bin): stop each keyed answer from re-waking this home (#4907) * fix(bin): treat home-owned status closes as already read Self-announced bookkeeping appends now record their exact byte ranges. Later drains and signal scans skip those ranges, so two distinct --resolve-key answers after an OPEN DECISIONS fold do not each wake the supervisor. Worker-authored lines outside that ledger still signal. * no-mistakes(review): Keep owned closes in unread status; lock ledger writes * no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake * no-mistakes(review): Require real owned growth before ledger marks status seen * no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS * no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test * no-mistakes(review): Align ledger docs and scope ledger to wake path only * no-mistakes(review): Restore stranded historical-annotation test comment to its function * no-mistakes(review): Retire the home-appends lock alongside its ledger * no-mistakes(document): Note ledger's lock-helper dependency in classify library * no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions * no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger * no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger * no-mistakes(document): Note owned-append skip in watcher signal-scan comment * fix: deliver failed public follow-ups with updated AXI floors (#5350) * chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the floor-boundary test fixtures to the new versions. tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final expecting pr-merged, so add the regression test: a bound work that ends failed reports its honest outcome text through fm-public-followup-emit.sh, consume marks the commitment ready, and deliver posts that text exactly once. Also make two hang-guard tests in fm-backlog-atomicity portable to hosts without coreutils timeout, and stop an installed herdr from leaking into the secondmate-liveness husk classifier test. * no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test * no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures * no-mistakes(document): Document failed public-followup delivery behavior * no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed * fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878) * fix(bin): refuse ship done: when the named head lives only in the worker copy A ship done: is not current-state done until that exact commit is reachable outside the disposable copy. The check tests the named head, not whether some branch moved. * fix(bin): gate CI-ready ship done: on named-head reachability, not handoff Keep no-mistakes' first done: as the pipeline handoff, apply the same shared check when registering a PR and when a secondmate publishes ledger-first, treat a recorded merged PR as landed after prune, and name the PR head instead of scanning free-text SHAs. * no-mistakes(review): Bind named-head gate to recorded PR and forge heads * no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries * no-mistakes(review): Align worker done wording, test mapping, pending-retry test * no-mistakes(test): Raise watcher test time limit to stop load flake * no-mistakes(document): Restore ledger-path fact and name named-head gate coverage * ci: re-attest named-head ship-done gate for a fresh serial-3 verdict * no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery * no-mistakes(document): Name fm-crew-state among named-head gate callers * fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204) * fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows. * no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier * test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335) The re-arm recovery cases judged "the watcher stayed live instead of surfacing recovery" with fixed budgets below what a real stale-lock recovery costs on a contended host: the arm's default 10s confirmation deadline, a start helper that returned after about 4s whether or not the arm had confirmed its watcher, and an 80-poll exit wait. A changed-suite run beside other suites starves the recovery's many short-lived processes while this suite's sleeping poll loops keep their pace, so a watcher still surfacing its recovery read as one that stayed live (issue #3793). The original 0.25s window after confirmation was widened to 80 polls in #3837, which left the same race at a larger size. Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now gives the arm an explicit 30s confirmation budget and waits for its confirmation or exit within a ceiling that outlasts it, and every wait on a re-armed watcher uses one named iteration-counted ceiling that outlasts the same budget. A passing case returns as soon as the arm reports or exits, and a watcher that never surfaces its recovery still fails. A new case delays every mktemp and readlink the re-armed watcher runs after it publishes its beacon, so its first poll and exit take about 13s on any host. It fails with the reported symptom on the previous budgets and passes now. No bin/ change. * fix: stop watchers reliably during blocked polls (#5362) * fix(bin): let one TERM always stop the watcher on bash 5.2 Bash 5.2 runs a pending trap from the parser entry of the next command substitution it expands, where the trap body is parsed as the inside of that substitution and fails ("trap: line 2: unexpected EOF while looking for matching `)'") or is dropped silently, consuming the signal. The watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and keep polling while its stopper waited: the triage suite's reap waited forever (CI jobs cancelled at 30 minutes), and the arm's signal path and the away-mode daemon's shutdown wait for the watcher the same way. Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04. HUP and TERM now keep bash's native fatal-signal handling, which runs the EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps its trap because bash ignores a direct SIGINT while a child runs. The check-spawn deferral window no longer contains a command substitution. The triage suite's reap is now bounded and fails the case within 10s with process evidence instead of hanging the job, and a new regression test proves TERM stops a watcher blocked inside a poll's pane capture and still releases its lock and records an acknowledgeable stop. * no-mistakes(document): Clarify watcher stop-signal documentation * fix: submit stuck inbox doorbells instead of skipping them (#5374) * fix(bin): submit our own stuck doorbell instead of skipping every later ring * no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit * no-mistakes(document): Clarify doorbell retry and pending-composer documentation * feat: add opt-in fleet activity ledger (#5375) * feat(bin): add the opt-in fleet activity ledger Homes that create config/fleet-ledger get an append-only JSONL file, state/fleet-ledger.jsonl, recording task.dispatched, task.status, task.merged, and task.cleaned_up so outside tools can follow a fleet. With the flag absent each producer does one file test and nothing else. docs/fleet-ledger.md owns the record contract and its documented limits. * no-mistakes(review): Record task.status text verbatim after the first colon * no-mistakes(document): Clarify fleet ledger status and setup documentation * no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed * fix: validate public follow-up deliverables and wake on rejection (#5352) * fix(bin): format, validate, and surface public-followup deliverables brief pre-fills report_path=data/<work-id>/report.md and states the accepted format of every value it cannot know instead of a bare <value> placeholder. fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in both the direct and staged destinations, naming the key, value, and format. consume records the specific deliverable, outcome, or missing key behind a tasks-axi refusal, and each refusal wakes the owning home once through the existing relay poll. * no-mistakes(review): refuse emits missing a required deliverable in both destinations * no-mistakes(review): require promised deliverables and keep rejections recoverable * no-mistakes(review): mirror tasks-axi's canonical pull request URL rule * no-mistakes(review): keep a rejection wake whose line cannot be read * no-mistakes(review): key emit-time rules on the promise, not the outcome * no-mistakes(review): bound deliverable keys and values as tasks-axi does * no-mistakes(review): state rejection wakes as at-least-once and pin it * no-mistakes(review): enforce the promised contract tasks-axi holds at emit * no-mistakes(review): stop inferring a staged promise from its outcome * no-mistakes(document): Refresh public follow-up documentation * no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks * Revert unrelated CI auto-fix edits to the watcher and bearings board test The CI step's automatic repair changed bin/fm-watch.sh and tests/fm-bearings-board-render.test.sh to chase two intermittent CI failures that also occur on main and are not part of this change. Restore both files so this branch carries only the public-followup deliverable fix. * no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing * no-mistakes(document): Clarify public-followup validation and rejection-wake documentation * feat: add Devin CLI crewmate and scout adapter (#5380) * Add verified Devin CLI worker adapter * no-mistakes(review): Drop Devin resolver refusal and launch marker * no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs * no-mistakes(document): Document Devin sidecar, resume, and worker-only facts * no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals * fix(control): never pair Devin interrupt presses on an idle agent A fast double Escape on an idle Devin opens its /revert picker, where Enter reverts file changes. fm-control now sends the second press only after the first renders Devin's 'esc again to interrupt' armed hint, never sooner than 0.5 s, closes a revert picker a mistimed press opened with one Escape, and refuses to type the exit command while that picker is open. An unarmed interrupt reports cancel=not-running and leaves the busy record untouched. * fix(devin): disable Claude hook import and commit attribution for workers The per-task Devin config now forces read_config_from.claude=false, so a worker no longer runs the user's or project's Claude Code hooks (including Herdr's Claude agent-state hook), and attribution=false, so Devin adds no Co-Authored-By trailer or Generated-with line to commits and PRs. * test(devin): extend live guard and record Herdr and revert-picker evidence The credentialed live guard now fails if an imported Claude Code hook runs, if the worker's commit carries Devin attribution, if an idle interrupt sends more than one press or opens the revert picker, or if an open picker lets exit through or is closed with a revert. The Devin reference, agent-control doc, and verification records carry the 2026-09-22 tmux and Herdr lab results, including the Herdr exit refusal. * no-mistakes(document): Correct Devin documentation links and lifecycle guidance --------- Co-authored-by: Denis Beliaev <battler73@yandex.ru> * fix(bin): recognize passed-with-skips as a passing outcome (#5322) fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as unknown, so a finished worker awaiting merge is re-alerted as stale. The same blind spot lets fm-teardown's pre-teardown terminal-run check refuse a legitimate abort race that lands on this outcome. Map passed-with-skips to done in crew-state resolution, keeping the skipped publication/CI verification visible in the detail rather than reporting a clean pass, and recognize it as terminal during teardown. * fix(bin): refuse unavailable backend adapters before sourcing (#5382) * fix: refuse missing backend adapter before source * no-mistakes(review): Gate backend precheck under stock Bash * no-mistakes(document): Clarify adapter precheck docs * no-mistakes(lint): Suppress intentional child Bash ShellCheck warning * test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338) * fix(test): repair tmux liveness and calm follow-up loaded_off regressions Both self-tests fail on untouched main on a host whose coreutils are a multicall binary and whose Chrome has no pre-warmed profile, and each failure masks the other's file. tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under another name, but a multicall coreutils binary (uutils or busybox) resolves its applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs the wrong applet and exits immediately, so no foreground process exists and every positive case reads not-alive ("last verdict for liveness:agent was missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as the stand-in target, exactly the way the version-string case already builds its executable, and require the fallback target to demonstrably survive the rename before using it. Every assertion is untouched; the stand-in identity signal is unchanged (the kernel still records the symlink name as the executable identity). tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new `--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that pristine profile makes Chrome's first-run initialization never complete: the browser and its renderers start, but --dump-dom never returns, so all three bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never run ("could not render calm-mode HTML export DOM"). Chrome's own profile creation under a fresh HOME renders the same document in about a second, so the helper now gives Chrome a private per-attempt HOME instead of the explicit profile flag. Each attempt still gets an isolated profile, and every DOM assertion is unchanged. Root-cause evidence: a pristine --user-data-dir with `--headless=new --dump-dom` had not returned after 150s, while the same command with an empty HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an already-populated profile also returned it in ~1s. The render failure masked the rest of the file: with it repaired, the Pi follow-up loaded_off case passes unmodified against an installed @earendil-works/pi-coding-agent package. These two failures block downstream validation of every lane on hosts with multicall coreutils or a fresh Chrome profile. Verification: - timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok - timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok, including the Pi operational follow-up loaded_off case - bash -n and shellcheck clean on both touched files - rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded * fix(test): give wake-queue observation checkpoints the alerting ceiling tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground watcher checkpoints whose job is to record an observation, with the alerting checkpoint that follows asserting the stall. A checkpoint's exit publishes a downtime marker, and the next checkpoint consumes it only by reaching the end of the watcher's poll loop, where the recovery surfacing runs after the stall tick; the observation itself is recorded by that same stall tick. On a loaded host a 1s ceiling sits under the cost of that iteration (which includes a pane capture in the active-turn gate), so the observation was never recorded, the downtime marker stayed pending, and the alerting checkpoint surfaced `check: rearm-resurface` instead of the stall it asserts: not ok - a foreign queue with no progress did not alert: check: rearm-resurface not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface Give the observation checkpoints that feed a later alert the same 4s ceiling the file already documents for alerting checkpoints. The ceiling is only a bound - a checkpoint still returns on its first actionable wake - so no assertion is weakened, and the quiet windows get longer, not shorter. * no-mistakes(document): docs: correct export-DOM Chrome render root cause * no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup * chore: re-trigger fork workflow approval for triage --------- Co-authored-by: Captain <blackxwhite88@users.noreply.github.com> Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com> * fix: keep watcher status classification bounded to new log spans (#5383) * fix(bin): classify a status span without re-folding the whole log A watcher poll could take minutes, so its liveness beacon aged past the guard's 300s grace and the Stop auto-arm reported the watcher down. On the main home, cycles ended with beacon_age 91-235s while healthy and 534-706s while the laptop was CPU-starved. Cause: whenever a newly appended status span held a keyed needs-decision or blocked line, status_span_first_actionable_record re-read and re-folded the ENTIRE log to decide whether that opening was still live, forking several subshells per line. On a remote second mate's mirrored parent channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log per classification when idle, paid by every signal and heartbeat scan. Nothing regressed recently: subshell counts per classification were 20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and 13,188 from #3753 onward through HEAD. The cost grew with log size, since parent-channel logs only grow. Fix: fold only the captured span. An accepted opening does not depend on earlier lines and only later lines close or supersede it, and every later line lies inside the span, so the span fold names the same live openings at a cost bounded by the span. Old and new classification outputs are byte-identical across 51 span offsets of real-shaped secondmate and ship logs. A real-watcher regression test records every read the classification makes through the span-reader seam and asserts none reaches before the classified offset; it fails on the old code (5,157 bytes read from offset 0 to classify an 84-byte span). * no-mistakes(document): Clarify span classification and watcher regression coverage * test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381) * test: fix watcher timing flakes in fm-pr-check-security The bounded watcher's hang guard now counts only the watcher's own time: a case marks the intervals where it holds the watcher on injected work or makes it wait on concurrent work, and those no longer count against its budget. The budget itself stays at main's sixty seconds. The helper also stops forcing a one-second per-check timeout, which killed a correct merged poll whenever that poll took longer than a second, so the watcher only retried it or exited on a later check's wake without the merge. The concurrent-publication case pauses the guard while its arming is in flight, and its task now sorts before the contributions observer the arming also registers, so the watcher stops on the poll under test before running that unrelated fleet snapshot. The case also prints the watcher's stderr when it fails. The replacement case pauses the guard while the re-arm runs inside the watcher, runs that injected arming with the fixture root every other arming here uses, and waits on the replacement merge's process instead of a two-second cap. Merged-poll runs retire the contributions observer before the watcher starts, since no case here exercises it. The returned-descendant case no longer races a four-second sleep or a TERM landing at an arbitrary point in the watcher's idle loop: its descendant holds until killed, and a second check in the same cycle witnesses that it was drained and stops the watcher. * no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed * Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed" This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6. * feat: record fleet status immediately and emit PR-ready events (#5385) * feat: record task.pr_ready in the fleet ledger when a task PR is registered * feat: record worker status lines in the fleet ledger as they are written * no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger * no-mistakes(review): Resolve relative config override before embedding in worker command * no-mistakes(document): Clarify fleet ledger status capture timing * test: synchronize foreign queue stall checks with watcher progress (#5386) * test: synchronize foreign secondmate stall legs on the watcher's recorded observation Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg depends on the progress observation the previous leg's watcher recorded. Under load the watcher was killed before its first stall tick, the observation was never written, and the next leg treated its own sighting as the first one, so the stall alert never fired. Run the watcher directly and end each leg on its observable outcome: the progress marker recording the expected observation, or the watcher's own first wake. Also move a comment orphaned above this test back to the drain liveness test it describes. * no-mistakes(review): Wait for full stall reset before stopping watcher leg * test: isolate the bearings render fixture from the shared Lavish store (#5391) The listener resolves its server from that store before it polls…
* feat(bin): defer the wedge escalation for a lane parked at a supervisor-owed gate (#4974)
* fix(watch): recheck a gate awaiting a human instead of wedge-escalating it
A lane whose validation run is parked at a gate waiting on a human
decision is correctly quiet, but nothing in its status line says so: the
evidence is the pipeline's own gate state rather than anything the worker
wrote. The wedge timer read that silence as a suspected wedge and climbed
the escalation ladder for as long as the wait lasted, and each escalation
cost a supervising turn. The landed declared-wait consult does not reach
it, because a live ordinary crewmate never reports a declared pause, and
raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for
every lane by the same amount.
The threshold now reads a second, independent record when the status line
accounts for nothing: whether the crew's current state is a gate whose
answer is owed by a human. That is minted only from the gate's own
findings table, by a row whose `action` column is exactly `ask-user`,
located by position out of the table header the way nm_gate_step_row
already reads its row - never searched for over the run payload, where a
finding's free-text description or a branch name satisfies a search just
as well. A gate awaiting the CREWMATE's own answer keeps the unchanged
escalation schedule, reason and demand-deep-inspection wording, because a
crewmate that goes quiet before answering its own gate is exactly the
wedge the ladder exists to catch.
Each kind of wait now carries the human it is on, the action that clears
it, and whether that human is the captain as data alongside the verdict,
rather than as wording chosen per branch where the recheck is written, so
the deferral cannot word one kind of wait as another and a new kind
cannot ship without deciding all of them. A parked gate has no written
record of when its wait began, so its recheck publishes no wait age at
all rather than one read from the quiet window this deferral resets on
every pass, which would report the same small number for a gate of any
age. Like every other captain-facing recheck here it is absorbed in
silence while the away-posture record exists, arming no throttle, so the
recheck is owed in full the moment the record is archived.
The consult runs only in the at-threshold branch that was about to
escalate, beside the worktree walk already there, and only for lanes
whose status line explained nothing.
Closes #3055
* no-mistakes(review): require an unanswered decision before deferring a parked gate
* no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records
* test(watch): pass the pane hash wedge_timer_check now takes
Upstream gave wedge_timer_check a sixth <pane-hash> argument for its
dead-record probe. The malformed-wait-record rounds drive the real function
directly, so they pass one, and stub fm_backend_agent_state to a live agent so
the probe that runs after a refused deferral keeps the unchanged ladder rather
than reading a backend the child shell has none of.
* no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate
* no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling
* feat(watch): make the parked-gate wait deferral opt-in
The wedge timer deferring a lane parked at a validation gate is new
supervision behaviour rather than a restored one, and it decides which
lanes give up the escalation ladder, so it now ships as a default-off
per-home option instead of changing every home on upgrade.
config/wedge-defer-parked-gate arms it. The flag is read before the
decision fold, so an unconfigured home spends no fold or current-state
read, writes no record, and keeps the unchanged escalation schedule,
reasons and demand-deep-inspection wording; a test counts the reader
calls in both directions to pin that.
It is not inherited by secondmate homes: each home supervises its own
crew and owns that trade separately, the same reason
config/turnend-churn-absorb is home-local.
The away-posture absorb returns to leaving the idle timer alone, which
it had restarted only because the costly consult could reach it. A
parked-gate wait is owed to the supervisor rather than the captain, so
it never enters that branch, and the recheck owed on return is again
owed in full the moment the record is archived.
* test(watch): pin that the away-silenced hold leaves the idle timer alone
The absorb no longer restarts the timer, so the recheck owed on return is
owed in full rather than a cadence into the return. Nothing asserted
that, so a restart could be reintroduced silently.
* no-mistakes(review): document away-silence rationale, pin captured gate component
* no-mistakes(test): anchor gate row scan to the braced findings header
* no-mistakes(document): pin same-block gate row invariant in crew-state comment
* fix(bin): reclaim a task whose herdr endpoint was destroyed (#5007)
* fix(control): let the owning seat reclaim a task whose endpoint is gone
A destroyed pane or workspace made `missing` a terminal state. Relaunch
accepted only `dead` and said to stop the agent first; exit refused
`missing` and said to reconcile the task first; there is no reconcile
verb. Each command named the other as its prerequisite, so a task whose
terminal went away could not be reclaimed by anything, and a no-mistakes
approval it was parked on had no seat left to answer it.
`missing` is agent-free a fortiori: there is no endpoint, so there is no
agent in it. Widen the existing guards rather than add a verb.
- fm-spawn --relaunch accepts a positively proven `missing` and creates
one fresh endpoint in the recorded worktree; the record it already
republishes rebinds the task to it. A `dead` endpoint is still adopted
in place.
- fm-control exit reports `endpoint-gone` instead of dying, so the
relaunch transaction's stop step no longer dead-ends, and re-resolves
the endpoint from the record before verifying the replacement.
The duplicate-agent refusal is untouched: both verdicts come from the
same recovery-grade classifier, which claims `missing` only from positive
absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The
backends' own create paths refuse a live same-labeled endpoint as a
second independent guard. The worktree, its branch, commits, uncommitted
changes, armed poll and registration, record rows, and status log are all
untouched - a reclaim is a recovery, never a teardown.
A secondmate is excluded: its gone-endpoint recovery already has one
owner in the session-start liveness sweep, so relaunch refuses and names
it rather than becoming a second path to the same outcome.
Tests reproduce both halves of the deadlock, the reclaim succeeding,
unlanded work surviving it, and the refusals that still hold.
* no-mistakes(review): prove endpoint absence per backend before reclaim rebinds
* no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session
* no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly
* no-mistakes(review): stop refusals and docs asserting unestablished causes
* no-mistakes(review): stop herdr fixture helper losing tmp-root registration
* no-mistakes(review): document workspace drift and absence-probe server residue
* no-mistakes(review): correct rebind limitation to its one reachable case
* no-mistakes(review): stop claiming reclaim leaves instructions untouched
* no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage
* no-mistakes(rebase): read the staged launch file in the herdr fixture
Rebasing onto main picked up #4994, which stages a long worker launch
command into a script and delivers the short `. '<path>'` line instead of
the literal command. The tmux fake and tests/fixtures.sh were updated for
that; the herdr fake this branch adds was written before it and still
keyed "an agent now exists on this pane" off the literal
`encode launch-brief` text, so after the rebase it never marked the
rebound pane live and the reclaim's alive-wait read `dead`.
Dereference the staged file first, exactly as the tmux fake above does.
Test-fixture only; no production path changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(document): note reclaim placement in herdr and scripts inventories
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(bin): stamp status events with their emission time (#3764)
* test(status): reproduce missing event emission time
* wip(status): preserve optional event emission time
* test(status): document indirect clock stub invocation
* no-mistakes(review): Preserve historical status bytes during reply recovery
* no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies
* no-mistakes(review): Preserve captain regex overrides for timestamped status events
* no-mistakes(document): Clarify status event timing and publication contracts
* no-mistakes(lint): Quote literal done to satisfy ShellCheck
* no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed
* no-mistakes(test): Preserve terminal notifications with malformed timestamp tags
* no-mistakes(test): Stamp Rovo spawn failures with emission time
* no-mistakes(document): Verify status event documentation
* no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests
* no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed
* no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed
* no-mistakes(review): Stamp remote escalations at call sites, drop new flag
* no-mistakes(review): Accept stamped escalation and close lines in test assertions
* no-mistakes(review): Restore reserved-key answered-note guard for stamped closes
* test(status): accept optional emission time in PR-provenance assertions
The #4148 provenance test landed on main with exact unstamped greps.
Parent-channel lines from this branch carry [at=<epoch>], so strip only
that tag before the same exact match. No production change.
* no-mistakes(review): Accept stamped ready signal in PR fallback scrape
* no-mistakes(review): Drop relay flag, stamp parent events at call sites
* no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture
* no-mistakes(review): Accept optional stamp in live cmux drift guard
* no-mistakes(review): Restore original test invocation order in two suites
* no-mistakes(review): Strip only well-formed numeric status time tags
* no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc
* no-mistakes(review): Stamp agy spawn-failure status lines with event time
* fix(bin): normalize status event times in-shell and freeze the budget test clock
Two paths made a status event's emission time cost more than it should.
The captain-relevance fallback piped every line through awk to drop a
well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a
fork per line just to prepare a regex match. Shell parameter expansion does the
same strip with no fork, and the retry-dedup scan now reuses that one helper
instead of carrying a second copy of the rule in awk. The copies had already
drifted: the shell side stripped tags from lines with no colon, which the awk
rule left whole, so a colonless line could be mistaken for one already
recorded. One definition, checked against the awk rule it replaces over the
edge cases and a 4000-line fuzz.
tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In
hang mode the poll set DEADLINE to the real now plus a one-second budget, and
when the second ticked before the first forge call the loop broke without ever
calling gh: forge/calls was never written and the assertion failed reading a
missing file. Freezing the clock in both modes removes the dependence on wall
time; the bounded call is still cut by the real timeout, so the observation the
test asserts still starts.
Emission time stays optional on new status records, and legacy or malformed
lines keep an unknown age.
* no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion
* no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs
* test: fold emission-time snapshot coverage into the fixture case
Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer
touches workflows. Keep every emission-time assertion by folding it
into test_fixture_snapshot_json.
* no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch
* no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape
* no-mistakes(review): strip undelimited at-tags; correct brief stamp header
* no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness
* no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles
* no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb
* no-mistakes(review): read note and key past colon-bearing stamps
* test(status): keep inactive reconcile assertions stamp-tolerant
These two oracles were made stamp-tolerant while resolving one of the
branch's merges from main. The rebase drops merge commits, so that
adaptation was lost and both assertions went back to matching an exact
substring that a stamped line no longer contains: the tag lands before
the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...".
Strip a well-formed tag before matching, as the branch's other oracles do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap
* no-mistakes(document): correct stale unstamped status-line spellings in docs
* no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments
* no-mistakes(ci): rename subshell-local epoch in delivery-race stub
The serialization test overrides fm_pending_reply_mark_delivered inside a
(..) subshell. Its `epoch` local collided with the same name in
status_line_at_epoch/status_stamp_line, which this branch added and this
suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and
failed Lint 2. The stub already prefixes its other locals with `pending_`
for the same reason; `epoch` was the leftover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(bin): unify Lavish host and disconnect handling (#5060)
* fix: ship clean Lavish host fixes
* no-mistakes(review): Fix Lavish classifications and fail-closed host loading
* no-mistakes(review): Restore Lavish host state across retries and launches
* no-mistakes(review): Preserve destination Lavish host when configuration is absent
* no-mistakes(document): Document Lavish status and host guarantees
* feat: act on captain's away words during AFK supervision (#5076)
* feat(afk): make the captain's away words the whole mandate
Retire the clause fields, verb list, never-set scan, refused records, and
the per-task merge-grant list from the away-posture record. The record is
now version 2: the captain's words verbatim plus expected return, spend
cap, and reach line; a version 1 record still validates, reads, and
archives so a live away window is never broken by the upgrade.
The supervision branch reads the words at the tail of every wake and acts
on them by its own judgment through the guarded scripts under standing
authority, never by analogy, holding for the return on doubt, and opens
each such outcome summary with "per your away instructions:" so the
return brief can render the words beside the session's account. While the
record exists any green merge runs under away authority (ledger tag
"away"); red merges, --allow-red, asynchronous and queued merges, and
local-only landing stay refused. The branch may file a backlog item the
words explicitly call for before dispatching it under the spend cap.
Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and
fm-pr-merge.sh as commands: version 2 written, version 1 read, retired
flags and subcommands refused by name, green merges landing under the
record, red and waived-red refused, the record lock still closing the
authority-read window, and the Pi away tail carrying the words.
* no-mistakes(review): carry the away read-back to the session verbatim
* no-mistakes(review): match the exact away-action marker in the return brief
* no-mistakes(review): refuse a words block truncated by a damaged line
* no-mistakes(document): Refresh away-role contract documentation
* fix(bin): render the remote charter's steering-inbox path host-local (#5049)
* fix(bin): render the remote charter's steering-inbox path host-local
A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.
Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.
Closes #5012
* no-mistakes(document): document remote charter's host-local steering inbox
* feat: route Lavish feedback directly to owning workers (#5099)
* feat(procevent): route worker-owned Lavish rounds
* no-mistakes(review): drop duplicate artifact field from task-owned registration
* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic
* no-mistakes(review): keep worker board owned until terminal round acknowledged
* no-mistakes(review): refuse every retirement of an open worker-owned round
* no-mistakes(review): use real lavish reply flag, isolate reply generations
* no-mistakes(review): drop .posted marker for best-effort reply posting
* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures
* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms
* no-mistakes(review): re-arm only to acknowledge an open round
* no-mistakes(review): conclude only a still-open terminal round
* no-mistakes(review): record the acknowledgement before retiring the board
* no-mistakes(review): retain the registration across a conclude, qualify terminal docs
* no-mistakes(document): Document worker-owned Lavish round lifecycle
* fix(bin): fit pull observation within the contribution poll budget (#5107)
* fix(bin): reserve contribution observation budget
* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
* feat(bin): add idempotent inbox capture, replies, receipts, and readiness JSON (#5103)
* feat(bin): add idempotent inbox orders, receipts, replies, and readiness
Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.
* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict
* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json
Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.
* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor
* no-mistakes(document): Note read-only lock inspection in scripts inventory
* no-mistakes(lint): Pass missing id argument to malformed-reply test printf
---------
Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
* fix(bin): stop harness footer rows below a composer from reading as pending text (#5118)
* fix(composer): stop a harness footer row from reading as a composer holding text
A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.
Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.
A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.
Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.
* no-mistakes(review): make composer footer-zone demotion shape-independent
* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan
* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100
---------
Co-authored-by: Koen Muller <koen@catapult.nl>
* feat(bin): append optional home-local include to briefs (#5115)
Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
* fix(bin): report a branch with no validation run as absent instead of an unreadable runs table (#5114)
* fix(bin): stop misreading a no-run branch as an unreadable runs table
Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.
Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.
Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.
* fix: recovered same-branch inventory awk misreads empty result as unreadable
fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.
Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.
* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup
* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings
* no-mistakes(review): revert repo lookup to exact working_path match
* no-mistakes(document): note state-db inventory read under crew-state nm timeout
* fix(bin): require a non-draft pull request before a PR-based done report (#5141)
* fix(bin): require a non-draft pull request before a PR-based done report
A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.
The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.
Closes #4757
* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix: support quota-axi schema 6 snapshots (#4904)
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey
quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.
- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
accountKey required on every row and uniqueness on
provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
is the one join every consumer uses: schema 5 binds by provider alone,
schema 6 binds to the row keyed by the candidate's Pi lane, else the
provider's default row, else no row (unmeasured, never blocked, never
by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
the shared function; an expanded provider with no row for the
candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
with a schema 5 case on the same path; every new case fails on the
previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.
* no-mistakes(review): Fix native Codex quota and expanded provider watches
* no-mistakes(review): Align native Codex account matching across dispatch paths
* no-mistakes(document): Align quota documentation with account-aware snapshots
* no-mistakes(document): Align quota dispatch documentation with account matching
* fix(bin): keep CI lint and the quota watch test portable
- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
that source this library, so full-mode ShellCheck reported SC2034 on
the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
assertions used rg, which CI runners do not install, so the case
failed with 'rg: command not found' rather than on behavior; use grep
like the rest of the file.
* no-mistakes(document): Documented schema-version account-row compatibility
* test: fix Claude session-start drain live E2E (#5165)
* test: repair Claude live auto-arm regression
* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event
* no-mistakes(document): Consolidate Claude live verification references
* ci: pin the no-mistakes required check to v1.80.1 (#5195)
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
* fix(bin): retain Pi watcher predecessor to stop false down alarms (#5174)
* fix: preserve Pi watcher ownership across session replacement
* no-mistakes(document): Scope Pi predecessor retention away from omp
* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
* fix(bin): allow cleanup of windowless legacy task records (#5236)
* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown
Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.
Co-authored-by: Cursor <cursoragent@cursor.com>
* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers
* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator
* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity
* no-mistakes(document): Clarify windowless teardown retry documentation
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: exempt kunchenguid from the no-mistakes required check (#5256)
* fix(bin): surface launches parked on an interactive prompt as not-started (#5250)
* fix: surface parked launch prompts as not started
* no-mistakes(document): docs: record launch-prompt busy backstop classification
* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* fix: record away posture immediately on /afk (#5260)
* feat(afk): make /afk itself the go with a same-turn record write
Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.
* no-mistakes(document): Refresh away-entry documentation evidence
* fix(bin): recognize passed-with-override as a passing outcome (#5294)
* fix(bin): map passed-with-override to done instead of unknown
no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.
Map passed-with-override to the same done/terminal handling as a
clean passed in both places.
* fix(document): Replace stale outcome mapping with authoritative pointer
* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: clean up workers after their pull requests land (#5317)
* fix: close landed workers from supervision in both postures and at return
During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.
- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
inactive-outcome, or heartbeat row on a done task with a merged PR, as the
moment to claim the lease and run bin/fm-teardown.sh with no flags; a
refusal is reported, never forced or worked around. Add teardown to the
handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
records only (a live task record whose recorded PR carries the
merge-notification marker), between could-not-fix and handled, without
holding the gate; the afk skill's return step closes each listed task
through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
in fm-afk-return through the real marker writer.
* no-mistakes(document): Document landed-task cleanup ownership
* fix: surface green no-mistakes PRs awaiting merge (#5327)
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring
A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.
Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.
Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.
Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.
* no-mistakes(review): read the full ci log when checking checks-green
* no-mistakes(review): correct stale ci log tail wording in docs
* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling route from board session (#5334)
* fix: derive Lavish polling server from its board session
* no-mistakes(document): Document session-derived Lavish polling
* no-mistakes(document): Correct Lavish routing verification claims
* fix(bin): stop secondmate relaunch failing when watcher scratch files vanish (#4900)
* fix(bin): ignore vanished state scratch files on secondmate relaunch
Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.
Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes #4765.
* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
* fix(bin): stop each keyed answer from re-waking this home (#4907)
* fix(bin): treat home-owned status closes as already read
Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.
* no-mistakes(review): Keep owned closes in unread status; lock ledger writes
* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake
* no-mistakes(review): Require real owned growth before ledger marks status seen
* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS
* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test
* no-mistakes(review): Align ledger docs and scope ledger to wake path only
* no-mistakes(review): Restore stranded historical-annotation test comment to its function
* no-mistakes(review): Retire the home-appends lock alongside its ledger
* no-mistakes(document): Note ledger's lock-helper dependency in classify library
* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions
* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger
* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger
* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
* fix: deliver failed public follow-ups with updated AXI floors (#5350)
* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest
Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.
tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.
Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.
* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test
* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures
* no-mistakes(document): Document failed public-followup delivery behavior
* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
* fix(bin): refuse ship done: when the named head exists only in the worker copy (#4878)
* fix(bin): refuse ship done: when the named head lives only in the worker copy
A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.
* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff
Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.
* no-mistakes(review): Bind named-head gate to recorded PR and forge heads
* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries
* no-mistakes(review): Align worker done wording, test mapping, pending-retry test
* no-mistakes(test): Raise watcher test time limit to stop load flake
* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage
* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict
* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery
* no-mistakes(document): Name fm-crew-state among named-head gate callers
* fix(bin): ring a proven-idle secondmate before raising a wake-loop stall alarm (#5204)
* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm
A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.
* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
* test(watch-arm): size re-arm waits off the real loaded recovery cost (#5335)
The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue #3793).
The original 0.25s window after confirmation was widened to 80 polls in
#3837, which left the same race at a larger size.
Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.
A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix: stop watchers reliably during blocked polls (#5362)
* fix(bin): let one TERM always stop the watcher on bash 5.2
Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.
HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.
The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.
* no-mistakes(document): Clarify watcher stop-signal documentation
* fix: submit stuck inbox doorbells instead of skipping them (#5374)
* fix(bin): submit our own stuck doorbell instead of skipping every later ring
* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit
* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat: add opt-in fleet activity ledger (#5375)
* feat(bin): add the opt-in fleet activity ledger
Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.
* no-mistakes(review): Record task.status text verbatim after the first colon
* no-mistakes(document): Clarify fleet ledger status and setup documentation
* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
* fix: validate public follow-up deliverables and wake on rejection (#5352)
* fix(bin): format, validate, and surface public-followup deliverables
brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.
* no-mistakes(review): refuse emits missing a required deliverable in both destinations
* no-mistakes(review): require promised deliverables and keep rejections recoverable
* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule
* no-mistakes(review): keep a rejection wake whose line cannot be read
* no-mistakes(review): key emit-time rules on the promise, not the outcome
* no-mistakes(review): bound deliverable keys and values as tasks-axi does
* no-mistakes(review): state rejection wakes as at-least-once and pin it
* no-mistakes(review): enforce the promised contract tasks-axi holds at emit
* no-mistakes(review): stop inferring a staged promise from its outcome
* no-mistakes(document): Refresh public follow-up documentation
* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks
* Revert unrelated CI auto-fix edits to the watcher and bearings board test
The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.
* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing
* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
* feat: add Devin CLI crewmate and scout adapter (#5380)
* Add verified Devin CLI worker adapter
* no-mistakes(review): Drop Devin resolver refusal and launch marker
* no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs
* no-mistakes(document): Document Devin sidecar, resume, and worker-only facts
* no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals
* fix(control): never pair Devin interrupt presses on an idle agent
A fast double Escape on an idle Devin opens its /revert picker, where Enter
reverts file changes. fm-control now sends the second press only after the
first renders Devin's 'esc again to interrupt' armed hint, never sooner than
0.5 s, closes a revert picker a mistimed press opened with one Escape, and
refuses to type the exit command while that picker is open. An unarmed
interrupt reports cancel=not-running and leaves the busy record untouched.
* fix(devin): disable Claude hook import and commit attribution for workers
The per-task Devin config now forces read_config_from.claude=false, so a
worker no longer runs the user's or project's Claude Code hooks (including
Herdr's Claude agent-state hook), and attribution=false, so Devin adds no
Co-Authored-By trailer or Generated-with line to commits and PRs.
* test(devin): extend live guard and record Herdr and revert-picker evidence
The credentialed live guard now fails if an imported Claude Code hook runs,
if the worker's commit carries Devin attribution, if an idle interrupt sends
more than one press or opens the revert picker, or if an open picker lets
exit through or is closed with a revert. The Devin reference, agent-control
doc, and verification records carry the 2026-09-22 tmux and Herdr lab results,
including the Herdr exit refusal.
* no-mistakes(document): Correct Devin documentation links and lifecycle guidance
---------
Co-authored-by: Denis Beliaev <battler73@yandex.ru>
* fix(bin): recognize passed-with-skips as a passing outcome (#5322)
fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as
unknown, so a finished worker awaiting merge is re-alerted as stale. The
same blind spot lets fm-teardown's pre-teardown terminal-run check refuse
a legitimate abort race that lands on this outcome.
Map passed-with-skips to done in crew-state resolution, keeping the
skipped publication/CI verification visible in the detail rather than
reporting a clean pass, and recognize it as terminal during teardown.
* fix(bin): refuse unavailable backend adapters before sourcing (#5382)
* fix: refuse missing backend adapter before source
* no-mistakes(review): Gate backend precheck under stock Bash
* no-mistakes(document): Clarify adapter precheck docs
* no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
* test: repair base-red liveness, export-DOM, and wake-queue self-tests (#5338)
* fix(test): repair tmux liveness and calm follow-up loaded_off regressions
Both self-tests fail on untouched main on a host whose coreutils are a
multicall binary and whose Chrome has no pre-warmed profile, and each failure
masks the other's file.
tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were
symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under
another name, but a multicall coreutils binary (uutils or busybox) resolves its
applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs
the wrong applet and exits immediately, so no foreground process exists and
every positive case reads not-alive ("last verdict for liveness:agent was
missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as
the stand-in target, exactly the way the version-string case already builds its
executable, and require the fallback target to demonstrably survive the rename
before using it. Every assertion is untouched; the stand-in identity signal is
unchanged (the kernel still records the symlink name as the executable
identity).
tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new
`--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that
pristine profile makes Chrome's first-run initialization never complete: the
browser and its renderers start, but --dump-dom never returns, so all three
bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never
run ("could not render calm-mode HTML export DOM"). Chrome's own profile
creation under a fresh HOME renders the same document in about a second, so the
helper now gives Chrome a private per-attempt HOME instead of the explicit
profile flag. Each attempt still gets an isolated profile, and every DOM
assertion is unchanged.
Root-cause evidence: a pristine --user-data-dir with `--headless=new
--dump-dom` had not returned after 150s, while the same command with an empty
HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an
already-populated profile also returned it in ~1s. The render failure masked
the rest of the file: with it repaired, the Pi follow-up loaded_off case passes
unmodified against an installed @earendil-works/pi-coding-agent package.
These two failures block downstream validation of every lane on hosts with
multicall coreutils or a fresh Chrome profile.
Verification:
- timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok
- timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok,
including the Pi operational follow-up loaded_off case
- bash -n and shellcheck clean on both touched files
- rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded
* fix(test): give wake-queue observation checkpoints the alerting ceiling
tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground
watcher checkpoints whose job is to record an observation, with the alerting
checkpoint that follows asserting the stall. A checkpoint's exit publishes a
downtime marker, and the next checkpoint consumes it only by reaching the end of
the watcher's poll loop, where the recovery surfacing runs after the stall tick;
the observation itself is recorded by that same stall tick. On a loaded host a
1s ceiling sits under the cost of that iteration (which includes a pane capture
in the active-turn gate), so the observation was never recorded, the downtime
marker stayed pending, and the alerting checkpoint surfaced
`check: rearm-resurface` instead of the stall it asserts:
not ok - a foreign queue with no progress did not alert: check: rearm-resurface
not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface
Give the observation checkpoints that feed a later alert the same 4s ceiling the
file already documents for alerting checkpoints. The ceiling is only a bound - a
checkpoint still returns on its first actionable wake - so no assertion is
weakened, and the quiet windows get longer, not shorter.
* no-mistakes(document): docs: correct export-DOM Chrome render root cause
* no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup
* chore: re-trigger fork workflow approval for triage
---------
Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
* fix: keep watcher status classification bounded to new log spans (#5383)
* fix(bin): classify a status span without re-folding the whole log
A watcher poll could take minutes, so its liveness beacon aged past the
guard's 300s grace and the Stop auto-arm reported the watcher down. On the
main home, cycles ended with beacon_age 91-235s while healthy and 534-706s
while the laptop was CPU-starved.
Cause: whenever a newly appended status span held a keyed needs-decision
or blocked line, status_span_first_actionable_record re-read and re-folded
the ENTIRE log to decide whether that opening was still live, forking
several subshells per line. On a remote second mate's mirrored parent
channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log
per classification when idle, paid by every signal and heartbeat scan.
Nothing regressed recently: subshell counts per classification were
20,272 from #3268 (2026-08-29, which introduced the whole-log fold) and
13,188 from #3753 onward through HEAD. The cost grew with log size, since
parent-channel logs only grow.
Fix: fold only the captured span. An accepted opening does not depend on
earlier lines and only later lines close or supersede it, and every later
line lies inside the span, so the span fold names the same live openings
at a cost bounded by the span. Old and new classification outputs are
byte-identical across 51 span offsets of real-shaped secondmate and ship
logs.
A real-watcher regression test records every read the classification
makes through the span-reader seam and asserts none reaches before the
classified offset; it fails on the old code (5,157 bytes read from
offset 0 to classify an 84-byte span).
* no-mistakes(document): Clarify span classification and watcher regression coverage
* test: close pr-check watcher test gaps (original flake already fixed by #5362 and #4878) (#5381)
* test: fix watcher timing flakes in fm-pr-check-security
The bounded watcher's hang guard now counts only the watcher's own time: a
case marks the intervals where it holds the watcher on injected work or makes
it wait on concurrent work, and those no longer count against its budget. The
budget itself stays at main's sixty seconds. The helper also stops forcing a
one-second per-check timeout, which killed a correct merged poll whenever that
poll took longer than a second, so the watcher only retried it or exited on a
later check's wake without the merge.
The concurrent-publication case pauses the guard while its arming is in
flight, and its task now sorts before the contributions observer the arming
also registers, so the watcher stops on the poll under test before running
that unrelated fleet snapshot. The case also prints the watcher's stderr when
it fails.
The replacement case pauses the guard while the re-arm runs inside the
watcher, runs that injected arming with the fixture root every other arming
here uses, and waits on the replacement merge's process instead of a
two-second cap. Merged-poll runs retire the contributions observer before the
watcher starts, since no case here exercises it.
The returned-descendant case no longer races a four-second sleep or a TERM
landing at an arbitrary point in the watcher's idle loop: its descendant holds
until killed, and a second check in the same cycle witnesses that it was
drained and stops the watcher.
* no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed
* Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed"
This reverts commit 6a59859b2e2a3778f9b46faeea42d6de37468cd6.
* feat: record fleet status immediately and emit PR-ready events (#5385)
* feat: record task.pr_ready in the fleet ledger when a task PR is registered
* feat: record worker status lines in the fleet ledger as they are written
* no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger
* no-mistakes(review): Resolve relative config override before embedding in worker command
* no-mistakes(document): Clarify fleet ledger status capture timing
* test: synchronize foreign queue stall checks with watcher progress (#5386)
* test: synchronize foreign secondmate stall legs on the watcher's recorded observation
Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once
ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg
depends on the progress observation the previous leg's watcher recorded. Under
load the watcher was killed before its first stall tick, the observation was
never written, and the next leg treated its own sighting as the first one, so
the stall alert never fired.
Run the watcher directly and end each leg on its observable outcome: the
progress marker recording the expected observation, or the watcher's own first
wake. Also move a comment orphaned above this test back to the drain liveness
test it describes.
* no-mistakes(review): Wait for full stall reset before stopping watcher leg
* test: isolate the bearings render fixture from the shared Lavish store (#5391)
The listener resolves its server from that store before it polls. Without a session for this bo…
Brings in four upstream fixes: - kunchenguid#6192 ci: rebalance portable test groups and enforce a packing budget - kunchenguid#6213 fix(bin): let a stale record on a reassigned slot retire records-only - kunchenguid#6216 fix(bin): run no repository hook when core.hooksPath is empty - kunchenguid#6240 fix(bin): keep the steering doorbell short under deep homes Naive merge would pick a stale history merge base because fork main carries the #11 and #14 syncs as squashes; an explicit content merge base of b3d4133 (last upstream commit whose content fork main shared) is used instead. Two genuine conflicts (bin/fm-teardown.sh, docs/architecture.md) were 3-way merged: upstream kunchenguid#6213's records-only code path survives, fork #13's claim-first semantics and endpoint-cleared behavior are kept, and the overlapping header/doc prose keeps the fork's newer wording. Content delta vs fork main is exactly the 22 files of kunchenguid#6192+kunchenguid#6213+kunchenguid#6216+kunchenguid#6240 (docs/architecture.md is byte-identical to fork main after resolution).
…Path refusal cases
|
| for path in "$root"/*.claim; do | ||
| [ -f "$path" ] && [ ! -L "$path" ] || continue | ||
| fm_claim_read "$path" || continue | ||
| if [ "$FM_CLAIM_HOME" = "$HOME" ] && [ "$FM_CLAIM_TASK" = "$TASK" ]; then | ||
| if rm -f -- "$path"; then |
There was a problem hiding this comment.
Replacement claims can be deleted When one home releases a claim while another acquires the same target,
release checks ownership before taking the target lock, and release-task removes matching files without taking that lock. Either command can delete the new holder’s claim. A third home can then claim the target and dispatch conflicting work. Recheck ownership under the target lock immediately before removing each claim.
| while [ "$path" != "${path#./}" ]; do path=${path#./}; done | ||
| path=$(printf '%s' "$path" | tr -s '/') | ||
| path=${path#/} | ||
| while [ "$path" != "${path%/}" ]; do path=${path%/}; done | ||
| [ -n "$path" ] || return 1 | ||
| printf 'area:%s:%s\n' "$project" "$path" |
There was a problem hiding this comment.
Equivalent areas get separate claims The area key keeps internal
. and .. components. For example, area:proj:src/../docs and area:proj:docs name the same directory but produce different claim files. Two homes can therefore both acquire that area, contrary to the documented one-key-per-target contract. Normalize those components or refuse ambiguous paths.
| if [ "$EXAMINED" -ge "$MAX_HOLDS" ] || budget_exhausted; then | ||
| DEFERRED=$((DEFERRED + 1)) | ||
| continue |
There was a problem hiding this comment.
Deferred holds are never checked Each sweep sorts aged holds oldest-first and examines only the first
MAX_HOLDS—12 by default. The next sweep starts at the same first row, so when there are more than 12, the remaining captain decisions are never re-verified. Old Done rows can also occupy the limited slots. Rotate or page through the deferred holds.
| LANE_CAP_EXCLUDE=$ID | ||
| [ -n "$LANE_CAP_MODEL" ] || LANE_CAP_MODEL=$(fm_meta_get "$RELAUNCH_META" model) | ||
| fi | ||
| fm_provider_cap_refuse "$STATE" "$CONFIG" "$HARNESS" "$LANE_CAP_MODEL" "$LANE_CAP_EXCLUDE" || exit 1 |
There was a problem hiding this comment.
OpenHands checks wrong provider cap If an OpenHands spawn relies on ambient
LLM_MODEL instead of --model, this check runs while MODEL is still empty. The model is resolved afterward, so the spawn can be admitted against the OpenHands fallback bucket rather than the model’s billing provider. That provider can exceed its configured lane cap. Resolve the effective model before checking capacity.
| if ! holds=$(snapshot_holds); then | ||
| SNAPSHOT_ERROR="could not read the aged-hold projection" | ||
| else |
There was a problem hiding this comment.
Failed reads erase the docket When the backlog projection fails, the sweep still replaces the previous docket with zero examined holds and an empty findings list. The wake line reports the error, but someone reading the docket loses the last valid assessment and sees an empty one instead. Keep the previous docket or record the read failure in it.
Intent
Second fork sync PR for keenvc/firstmate (PR #15): bring in four upstream fixes merged since PR #14's sync point (#6192 rebalance portable test groups and packing budget, #6213 stale record on a reassigned slot retires records-only, #6216 run no repository hook when core.hooksPath is empty, #6240 keep the steering doorbell short under deep homes). Fork main carries the #11 and #14 syncs as squashes, so the merge uses an explicit content merge base b3d4133. Two genuine conflicts (bin/fm-teardown.sh, docs/architecture.md) were 3-way resolved: upstream #6213's records-only code hunk survives, the fork's #13 claim-first semantics and endpoint-cleared behavior are kept, and the overlapping header/doc prose keeps the fork's newer wording. A prior pipeline run validated this content end-to-end and opened PR #15 on the fork, where all 19 substantive CI checks pass. The branch was then manually cleaned: the document step had accidentally committed 75 untracked local .swarm/ runtime-state files which broke the fork's documentation-audience check; the branch was rebuilt at b9ebe4c dropping exactly those files (restored untracked; .git/info/exclude covers .swarm/). This run re-attests the cleaned head so the PR-must-be-raised-via-no-mistakes gate binds to b9ebe4c. The pipeline agent is now pi because the configured Claude Max seat hit its weekly limit (resets Oct 3) and the gate refuses grok for this repo; the user explicitly approved reconfiguring to a seat with headroom.
What Changed
bin/fm-test-run.shrebalances the portable test groups and enforces a packing budget with coverage reporting,bin/fm-teardown.shlets a stale record on a reassigned pool slot retire records-only,bin/fm-git-strip-ai-trailers.shruns no repository hook whencore.hooksPathis empty instead of refusing, and the steering doorbell names$FM_TASK_INBOXrather than an absolute path so its line no longer grows with home depth (bin/fm-task-inbox-lib.sh,bin/fm-spawn.sh).bin/fm-teardown.shkeeps upstream's records-only code hunk alongside the fork's claim-first semantics and endpoint-cleared behavior, anddocs/architecture.mdkeeps the fork's newer wording.core.hooksPathrefusal cases and corrected fork-specific portable serial-lane coverage claims.Risk Assessment
✅ Low: This is a bounded sync merge whose source delta is four upstream-reviewed fixes with small, semantically consistent conflict resolutions; I reconstructed the changed control flow and the coverage/packing arithmetic and found no concrete reachable defect, with the only caveat being the documented, guard-enforced need for the fork to measure its own 10 unmeasured serial members.
Testing
Drove all four upstream fixes and both conflict resolutions against the real firstmate product on head b9ebe4c. #6192: the live runner's
--check-coveragereports a complete, disjoint 9-shard portable serial partition packed under its 1200000ms budget, and its exact-budget/one-ms-over boundary test passes. #6216: a real commit in a repo with an emptycore.hooksPathsucceeds, runs no repository hook, and still strips the AI trailer; the broken-config cases are refused by git itself on this git 2.43 host (documented). #6213: a realfm-teardown.shrun retires a stale record on a slot claimed by another task records-only, then the claimant tears down and returns the slot, while contested and unclaimed collisions still refuse before mutation. #6240: realfm-send.shinto a real tmux pane rings the identical 166-character doorbell for shallow and deep homes, with no absolute path and the inbox named once. The fork's #13 claim-first and endpoint-cleared semantics survive the merge (own/absent claims tear down, cleared endpoints complete without force while unlanded work still refuses). The branch is clean: no tracked .swarm files and the documentation-audience check passes. The only failure seen,tests/fm-session-lock-ancestry.test.sh(which also fails one case oftests/fm-test-run.test.sh), is a hostsystemd --usersubreaper reparenting mismatch in a file unchanged by the sync, not a regression. Three scenarios could not be driven live here: the live-agent doorbell e2e (needs an installed harness seat and real tokens; the configured Claude Max seat is exhausted), the CI workflow/runner lane-agreement test (needsruby, absent on this host), and the merged architecture doc's slot-ownership wording (a committed-artifact prose check read from HEAD, not driven against a running product).bin/fm-test-run.sh --check-coverage->FM_TEST_COVERAGE ok ... serial_shards=9 serial_max_ms=1074843 serial_budget_ms=1200000; artifact 01-portable-test-groups.txtbin/fm-git-strip-ai-trailers.sh install+ realgit commitwithcore.hooksPath='': rc=0, repo pre-commit did not run, commit body has no Co-authored-by; artifact 02-empty-hookspath.txtbin/fm-teardown.sh stale-task --forcerc=0 with the claimant's slot sentinel and record intact, thenlive-task --forcerc=0 returning the slot; artifact 03-teardown-reassigned-slot.txtbin/fm-send.shinto a real tmux pane: shallow and deep homes both produced the identical 166-char line with no absolute path;FM_TASK_INBOXlisted the deep inbox from/; artifact 04-doorbel…treehouse <return>); reassigned-slot teardown leaves the claimant's slot alone; artifacts 03 and 05cleared:<reason>, E2 refuses and mutates nothing, E3 refuses with the ordinary missing-window error; artifact 05-endpoint-cleared.txtgit show HEAD:docs/architecture.md) as a committed-artifact/prose check and never drove a running product, so the…bin/fm-doc-audience-check.sh->ok surfaces=122 local_links=716;git ls-files -- '.swarm/**' | wc -l-> 0; artifact 06-cleanliness.txtrubyis not installed on this host (only python3/pyyaml are present), and #6192 did not modify .github/workflows/ci.yml. The runner-side lane membership and budget were validated live via --check-co…Evidence: #6192 portable test groups + packing budget (live)
Source: #6192 portable test groups + packing budget (live)
Evidence: #6216 empty and broken core.hooksPath (live)
Source: #6216 empty and broken core.hooksPath (live)
Evidence: #6213 stale record on a reassigned slot (live)
Source: #6213 stale record on a reassigned slot (live)
Evidence: #6240 doorbell stays short under a deep home (live)
Source: #6240 doorbell stays short under a deep home (live)
Evidence: Fork #13 endpoint-cleared teardown semantics (live)
Source: Fork #13 endpoint-cleared teardown semantics (live)
Evidence: Branch cleanliness and documentation-audience check
Source: Branch cleanliness and documentation-audience check
Evidence: Focused hook/doorbell/inbox test suites
Source: Focused hook/doorbell/inbox test suites
Evidence: Teardown endpoint-safety and full teardown suites
Source: Teardown endpoint-safety and full teardown suites
Evidence: Runner acceptance cases and host reparenting proof
Source: Runner acceptance cases and host reparenting proof
Evidence: Sync content present at the re-attested head
Source: Sync content present at the re-attested head
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
.agents/skills/agent-skill-trigger-index/SKILL.md- branch carries 48 commit(s) that exist on your local main branch but were never pushed to origin/main; these may be unintended bundled work (proposed PR changes 98 file(s)):Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.
🔧 No changes applied.
1 warning still open:
.agents/skills/agent-skill-trigger-index/SKILL.md- branch carries 48 commit(s) that exist on your local main branch but were never pushed to origin/main; these may be unintended bundled work (proposed PR changes 98 file(s)):Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.
no changes applied: bundled local-default commits require manual separation or explicit approval; the rebase conflict resolver cannot safely select commits to discard.
✅ **Review** - passed
✅ No issues found.
tests/fm-session-lock-ancestry.test.sh:772-tests/fm-session-lock-ancestry.test.shfails on this host withthe pty-host was not reparented to init after the daemon ended, and that failure also makes the--jobs on a proven family must be admittedcase insidetests/fm-test-run.test.shexit 1, because the runner's concurrency case executes the watcher-wake-lock family. This host runs asystemd --usersubreaper at PID 1285, so an orphaned process reparents to 1285, not PID 1 (verified:( sleep 6 & echo $! )yields PPID 1285). The test file is byte-identical between base 23e5584 and target b9ebe4c (git diffempty), so this is a host capability mismatch, not a regression introduced by the sync. Reported so a green local run is not misread and so the CI-only coverage is known. No action for this change.bin/fm-test-run.sh --check-coverage->FM_TEST_COVERAGE ok ... serial_shards=9 serial_max_ms=1074843 serial_budget_ms=1200000; artifact 01-portable-test-groups.txtbin/fm-git-strip-ai-trailers.sh install+ realgit commitwithcore.hooksPath='': rc=0, repo pre-commit did not run, commit body has no Co-authored-by; artifact 02-empty-hookspath.txtbin/fm-teardown.sh stale-task --forcerc=0 with the claimant's slot sentinel and record intact, thenlive-task --forcerc=0 returning the slot; artifact 03-teardown-reassigned-slot.txtbin/fm-send.shinto a real tmux pane: shallow and deep homes both produced the identical 166-char line with no absolute path;FM_TASK_INBOXlisted the deep inbox from/; artifact 04-doorbel…treehouse <return>); reassigned-slot teardown leaves the claimant's slot alone; artifacts 03 and 05cleared:<reason>, E2 refuses and mutates nothing, E3 refuses with the ordinary missing-window error; artifact 05-endpoint-cleared.txtgit show HEAD:docs/architecture.md) as a committed-artifact/prose check and never drove a running product, so the…bin/fm-doc-audience-check.sh->ok surfaces=122 local_links=716;git ls-files -- '.swarm/**' | wc -l-> 0; artifact 06-cleanliness.txtrubyis not installed on this host (only python3/pyyaml are present), and #6192 did not modify .github/workflows/ci.yml. The runner-side lane membership and budget were validated live via --check-co…bin/fm-test-run.sh --check-coverage(live, exit 0: total=251, serial_shards=9, serial_max_ms=1074843 <= serial_budget_ms=1200000, disjoint/complete cover)bin/fm-test-run.sh --list-lanes(9 portable serial shards present)bash tests/fm-test-run.test.sh(all #6192 acceptance cases pass: coverage guard, serial packing budget, exact-budget boundary, disjoint shard cover; file exits 1 only on the environment-caused ancestry child)Live #6216 drive: realbin/fm-git-strip-ai-trailers.sh install+ realgit commitin a repo withcore.hooksPath=''(commit rc=0, no repo hook ran, trailer stripped, empty stderr)Live #6216 adversarial drive: real commit in repos with unresolvable and valuelesscore.hooksPath(refused, HEAD unmoved, git error once)bash tests/fm-git-strip-ai-trailers.test.sh(exit 0)Live #6213 drive: realbin/fm-teardown.sh --forceagainst a treehouse pool fixture where a stale record and the claimant both name the slot (stale retires records-only, claimant then tears down and returns the slot)Live #6213 adversarial drive: contested slot (own claim + second live record) and unclaimed slot with two records both refuse before any mutationbash tests/fm-teardown-endpoint-safety.test.sh(exit 0, includesa stale record on a claimed slot retires, then the claimant tears down)bash tests/fm-teardown.test.sh(exit 0, 107 ok)Live #6240 drive: realbin/fm-send.shthrough real tmux into a real pane, shallow vs deep home (identical 166-char doorbell, no absolute path, inbox named once, FM_TASK_INBOX resolves the deep inbox from/)bash tests/fm-send-inbox.test.shandbash tests/fm-task-inbox.test.sh(both exit 0)Live fork #13 endpoint-cleared drive: realbin/fm-teardown.shon a windowless record withendpoint_cleared=(completes without --force/--legacy-record), with unlanded work (refuses, mutates nothing), and without a stamp (ordinary missing-window refusal)bin/fm-doc-audience-check.sh(exit 0: ok surfaces=122 local_links=716) andgit ls-files -- '.swarm/**'(0)git merge-base --is-ancestorfor eb77f02b, bd744684, 65c75b0d, 23e55847 against HEAD (all yes)bash tests/fm-session-lock-ancestry.test.sh(exit 1; host subreaper, documented as finding)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.