feat(pi): accept native Codex ultra effort with progress-aware supervision - #4038
Conversation
…d by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
|
Speaking as Kun's firstmate: Triage: HEAD Classing evidence:
VISION (per rule):
contract-class: opt-in. New unconfigured default path does not select ultra; progress/native-tools activate only for native Codex paths. Security: no tip beyond ordinary adapter surface. Auto-merge now. |
|
Speaking as Kun's firstmate: this is merged. Thank you @3264studios — really appreciate you taking the time on this. |
* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)
* fix(bin): refuse the behavior suite in the repository primary checkout
A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.
bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.
bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.
* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT
* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(bin): bound stale alarms for backlog captain holds (#3842)
* fix(bin): bound a stale alarm with the backlog hold, not only the status line
A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.
Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.
Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.
Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.
Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.
* fix(bin): identify the captain call a stale alarm is bounded by
Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.
The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.
The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.
The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.
Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.
* fix(review): Compose stale throttles with captain-call lifecycle identity
* fix(review): Preserve bounded same-hash captain-call resurfacing
* revert(bin): narrow the captain-hold stale bound to its observed defect
Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.
Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.
Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.
History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.
* fix(review): Document secondmate captain-hold scope boundary
* fix(document): Document captain-hold stale alarm scope
* fix(bin): bind the stale throttle to the captain call, not the status log
The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.
That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.
Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:
base c499f84 call-1 first=ALARM call-1 churn=ALARM new call first sight=ALARM
before this fix call-1 first=ALARM call-1 churn=absorbed new call first sight=absorbed
after call-1 first=ALARM call-1 churn=absorbed new call first sight=ALARM
Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.
`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.
The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.
* fix(document): Document captain-call throttle lifecycle scope
* fix(ci): isolate the Herdr restart fixtures from a claimed worktree
The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.
The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
* feat(pi): resolve extension-registered providers in the supervision branch (#3871)
* Let the supervision branch resolve extension-registered providers
The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".
Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.
Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.
* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc
* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered
* ci: retrigger flaky Herdr/serial-1 lanes
* no-mistakes(document): docs already cover branch extension-provider copy
* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)
* fix(watch): detect stalled secondmate queue progress
* no-mistakes(review): gate secondmate stall on active turns and progress episodes
* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector
* no-mistakes(document): clarify active-turn gate in wake-stall config docs
* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree
* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
* test(calm): harden the export-DOM render step and record Pi 0.85.1 evidence (#3952)
* fix(tests): make the Calm export-DOM render step retry and report
The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.
What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.
Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.
The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo
* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change
* no-mistakes(review): tests: drop inert chrome flags, report render timeouts
* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): make every counted wake queue row presentable or retired (#3950)
* fix(bin): make every counted wake queue row presentable or retired
A wake row could be counted as queued while no drain would ever present
it, leaving the operator told to "drain them before anything else" by a
command that printed nothing and offered no acknowledgement.
Two independent paths produced that state.
A row reserved by a live supervision-branch grant is excluded from a main
drain by design, but fm-guard.sh counted the whole queue, so main was
warned about rows only the branch could present, on every guarded command
for as long as the grant was held.
A row that lost its five appended fields or its numeric sequence can never
be claimed, presented, or named by an --ack-through cutoff, yet it still
counted as queued, wedging the queue permanently.
The guard now counts only the rows the calling actor can itself present or
retire, and a main drain retires unusable rows under the queue lock,
reporting them in bounded escaped form before removal so the evidence
survives for the separate row-generating defects. A retirement failure is
reported loudly and never suppresses unrelated consumable work. A main
drain whose remaining rows are all branch-held says so in one bounded line
instead of exiting silently. Grant row-list and owner-record reads move
into fm-wake-lib.sh so the drain, the grant publisher, and the guard share
one implementation.
Ownership is unchanged: a branch drain still touches nothing outside its
grant and never retires a row, and main still cannot present or acknowledge
an active grant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1
* no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass
* no-mistakes(review): add guard advisory for branch-held wake rows
* no-mistakes(document): document per-actor wake counting and unusable-row retirement
* no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use system stat for Darwin BSD formats (#3305)
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing
* fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main
* test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state
* ci: re-trigger after Chrome headless timeout in calm HTML export test
* ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test
* test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat
* no-mistakes(document): Refresh stat and shard docs
* fix(spawn): carry attribution-off policy in every claude launch (#3945)
The captain's attribution policy (no Co-Authored-By trailer, no
Claude-Session link, no generated-with line) lives in Claude Code's `user`
settings scope. A spawned worker's settings sources are not guaranteed to
load that scope, so a launched worker could write attribution trailers into
its commits and PR bodies regardless of the captain's own configuration.
launch_template()'s claude case now carries the same policy
("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in
its inline --settings JSON, so every claude launch keeps attribution off
independent of which settings scopes end up loaded. Tests assert the policy
on the rendered launch command for both a crewmate and a secondmate spawn.
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
* fix(bin): derive watcher beacon staleness grace from poll cadence (#3946)
* fix(bin): derive the away-mode beacon grace from the poll cadence
fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.
Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.
fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.
Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.
* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh
* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace
---------
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
* fix(procevent): reap orphaned runners and prevent launch storms (#3904)
* fix(procevent): bind a source runner to the session that owns it
A process-event source runner is detached into its own process group so a
persistent source survives the turn that armed it. Nothing bounded that
detachment, so a runner could reparent to init and keep its blocking child -
and every process that child spawned - running with nothing left to reap it.
One such runner outlived its home for about a day; the cost was not the runner
but the exec churn of the poll stubs under it, which stalled every fresh
process launch on the host.
Each runner now starts a small guard beside it, in a separate process group,
that re-reads its home's process-event lease and stops the runner's whole
process group once that lease can no longer be proved fresh. Every ordinary
entry point an owning session runs refreshes the lease, and the watcher's
reconcile cycle keeps it fresh in a live home; nothing a runner spawns can
refresh it, so a source cannot certify its own owner. Scope is the owning state
root and one runner generation, never a script or process name, so a live
source in another home is untouched and a live home simply starts a
replacement runner on its next cycle.
The test scaffolding that starts real runners could not reap them either: the
bearings-board and board-render suites tracked their homes in a shell array
appended to inside a command substitution, so the array was always empty and
every listener they started survived the run. Home registration moves to a
`$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path,
now including HUP and QUIT, and the blocking fixture stubs stop themselves at a
bound so an escaped one cannot keep spawning processes indefinitely.
Adds a regression test that reproduces the orphan shape - a reparented listener
with a live descendant tree under it - and proves the whole group and its
process churn stop once its session is gone, that an identical listener in a
home whose session is still there is untouched, and that retirement still
reaches a reparented listener and everything under it.
* no-mistakes(review): Bound source launches and fail closed on guard startup
* no-mistakes(document): Document runner lease and storm containment
* no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite
* test(procevent): pin that an unprovable stop is retried, not abandoned
The owner guard used to call stop_runner_pid and exit unconditionally, so a
stop it could not prove - a descendant still finishing uninterruptible work
outlives even the group signal, and an unreadable process identity proves
nothing - left a still-running expired runner with nothing watching it. That is
the best-effort reaping this mechanism exists to remove, and the fix that made
the guard retry landed without a test holding it in place.
The unprovable attempt is injected through the signal the real path actually
reads: `ps` answers exactly one process-group query for the runner with a group
it does not lead, which is how a stop that cannot be proved is reported, and
every other call is the real command. The test also asserts that the injected
attempt happened, so it cannot pass vacuously if the fixture stops arming.
Fails against the exit-after-one-attempt guard, where the runner survives its
expired lease, and passes once the guard retries on its check cadence.
* docs(procevent): scope the no-self-refresh rule to confused-agent grade
The runner-lease documentation asserted as an absolute that nothing a runner
spawns can refresh the lease, so a source cannot certify its own owner.
That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually
enforces.
The marker holds at confused-agent grade: a runner and its ordinary children
inherit it and skip every refresh, which is exactly the accidental case this
boundary exists for.
A source that deliberately strips the marker from its environment can still
refresh, so adversarial-grade unforgeability is explicitly out of scope and
tracked as separate follow-up design work.
This states the real scope in docs/configuration.md, which owns the operating
contract, and corrects the two matching comments in bin/fm-procevent.sh.
The process-event-sources skill keeps its cross-reference and gains one line
in its never-to-be-claimed list so the overclaim is not reintroduced from the
agent-facing side.
The lease mechanism itself is unchanged.
* no-mistakes(review): Fix process-event lease and launch pacing edge cases
* no-mistakes(review): Scope launch pacing and clarify lease boundaries
* no-mistakes(review): Reap leftover groups and use monotonic launch pacing
* no-mistakes(review): Keep guards alive across runner PID reuse
* no-mistakes(review): Use monotonic leases and simplify launch generation identity
* no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards
* no-mistakes(review): Preserve active pacing state on failed registration
* no-mistakes(review): Reap reused runner groups with registration evidence
* no-mistakes(review): Avoid ambiguous group kills and encode pacing identities
* no-mistakes(review): Abort kill escalation after runner identity reuse
* no-mistakes(review): Gate group signals and prune stale pacing state
* no-mistakes(review): Document bounded PID reuse signaling safety
* no-mistakes(review): Align leaderless group ambiguity guidance
* no-mistakes(review): Expire reboot stamps and preserve publication success
* no-mistakes(review): Bind owner leases to physical state roots
* no-mistakes(document): Clarify process-event lease and pacing contracts
* fix(procevent): drop a platform-dependent post-TERM test assertion
CI ran red on two lanes that the local gate could not see.
Lint failed with SC2034 on two reads in cmd_owner_watchdog that
destructure the state-root identity into five fields while using only the
device and inode.
Local changed-file mode suppresses the cross-file codes that need
--external-sources, so the warning cleared the pre-push lint step and
failed CI's full analysis, exactly as bin/fm-lint.sh's header describes.
The unused fields now read into `_`.
The behavior shard failed on this suite's own post-TERM assertion, which
required the stubbed identity source to be consulted more than once.
Whether that happens is platform-dependent: where the runner leader keeps
waiting on its TERM-ignoring source child, the post-TERM check sees a live
leader whose identity no longer matches, and where the leader dies
promptly it sees a leaderless group carrying the same numeric id.
fm_procevent_pid_state reaches that second verdict without consulting
process identity at all, so the identity source is never read twice and
the count assertion fails through no fault of the behavior.
The case now asserts the invariant both forms share: retirement refuses,
and the ambiguous group is not signalled.
Scoping a mutation to this fixture and making the refusal signal instead
confirms the case still fails, so dropping the count does not leave it
passing vacuously.
* no-mistakes(review): Prevent superseded runners recreating stale pacing stamps
* no-mistakes(document): Document pacing and ambiguity boundaries
* fix(procevent): retire under the recorded identity source and state the home-scoped lease
The reused-group case started its runner with the proc-root override in
place, so the runner recorded a ps-derived identity, then retired it
without that override.
Where /proc exists the retirement read identity from a different source
than the one recorded, the guard correctly refused an identity it could
not confirm, and the case failed on Linux while passing on macOS.
It now retires under the same source, and clearing the stub marker first
turns that cleanup into the complementary assertion: once the ambiguity is
gone, retirement reaps the whole group instead of leaving it behind.
The lease prose claimed a runner is bound to the session that owns it,
while the mechanism binds it to the home.
That gap is what makes a replacement session or an inspection command look
like a defect: any activity in the same home refreshes the lease.
The granularity is deliberate, because a persistent source is meant to
outlive the session that armed it, and binding a runner to that session
would stop the sources this mechanism exists to keep running.
A runner whose source is no longer wanted in a live home is stopped by
reconcile when that source is retired, independently of the lease, so the
lease is the backstop for a home that is gone - the torn-down sandbox this
change bounds - and the residual is recorded as a known limit.
* no-mistakes(review): Rate-limit polls and skip superseded runner launches
* fix(procevent): build the claim-only sweep case as a runnerless owned claim
A superseded generation now observes the registration-identity mismatch,
self-retires, and releases its claim, which is the behavior we want: it
clears its own residue rather than leaving a claim with no runner for the
home sweep to find.
The claim-only sweep case was built by deleting a registration out from
under a live runner, which used to leave that runner in place. It now
makes the runner retire itself, so the sweep raced that exit and retired
one source or two depending on which won. The case failed three runs in
four, alternating between a preflight-count failure and `attempted=1`.
It now builds the state it means to test: kill the runner's group so it
cannot run its own cleanup, assert the owned claim survived that kill, and
only then drop the registration. Coverage is unchanged - a runnerless
owned claim must still be swept - and the result no longer depends on
whether the runner had exited yet. Three consecutive runs pass.
The superseded exit also skipped the runner-marker cleanup the normal path
performs. The marker is written before the launch floor is waited on, and
a home sweep counts a marker with no owned claim as a preflight failure,
so exiting without clearing it would make that home refuse to sweep.
`FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI
runs, though not under the changed-file mode the pre-push gate uses.
* test(procevent): retire a quiet reparented listener instead of racing a storm
Explicit retirement was exercised against the spawn-churning stub, which
made it nondeterministic. Retirement refuses rather than signalling when it
cannot confirm the runner's identity, that identity is read through `ps`,
and the stub's 0.1s spawn loop starves that read often enough that a single
attempt is a race - the suite failed on this case roughly one run in four,
reporting `cannot confirm runner identity; source remains registered`.
The refusal is correct: it is the documented preserve-for-retry contract,
and a separate case already asserts it. So this is a fixture problem, not a
behavior problem.
The storm is still covered where the evidence for it lives. The owner-loss
home keeps the churning stub and still asserts its tick log stops, which is
what proves the churn ended rather than one pid going away. The retirement
home never asserted ticks; it only ever read the descendant pid, so the
spawn loop bought this case nothing while costing it determinism.
It now uses a quiet stub that still reparents and still holds a real
descendant in its process group, so the assertions are unchanged: retiring
the source must reap the reparented listener's whole group and the
descendant under it. Four consecutive runs pass.
* no-mistakes(review): Serialize registration replacement through source child launch
* chore(no-mistakes): require honest test-step scenario marking
The test step recorded scenarios as passing that were only reached through
a stubbed dependency or the executable suite, and its validator refused
them, because `pass` asserts a scenario was verified against the real live
product.
That refusal is correct, so the fix is to mark honestly rather than to
weaken the gate: a scenario driven live stays a pass and cites its live
transcript, while one reached only through a stub or the suite is recorded
as untested with the reason and a pointer to its executable coverage.
Untested scenarios are reported rather than treated as failures, so real
coverage stays visible without claiming verification that did not happen.
The instruction also forbids dropping a scenario to avoid marking it
untested, since that would hide the gap instead of stating it.
* no-mistakes(review): Remove unrelated test scenario policy
* no-mistakes(document): Clarify process-event home lease documentation
* no-mistakes(document): Correct owner guard failure wording
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority re…
* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)
* fix(bin): refuse the behavior suite in the repository primary checkout
A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.
bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.
bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.
* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT
* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(bin): bound stale alarms for backlog captain holds (#3842)
* fix(bin): bound a stale alarm with the backlog hold, not only the status line
A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.
Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.
Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.
Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.
Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.
* fix(bin): identify the captain call a stale alarm is bounded by
Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.
The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.
The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.
The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.
Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.
* fix(review): Compose stale throttles with captain-call lifecycle identity
* fix(review): Preserve bounded same-hash captain-call resurfacing
* revert(bin): narrow the captain-hold stale bound to its observed defect
Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.
Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.
Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.
History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.
* fix(review): Document secondmate captain-hold scope boundary
* fix(document): Document captain-hold stale alarm scope
* fix(bin): bind the stale throttle to the captain call, not the status log
The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.
That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.
Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:
base c499f84 call-1 first=ALARM call-1 churn=ALARM new call first sight=ALARM
before this fix call-1 first=ALARM call-1 churn=absorbed new call first sight=absorbed
after call-1 first=ALARM call-1 churn=absorbed new call first sight=ALARM
Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.
`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.
The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.
* fix(document): Document captain-call throttle lifecycle scope
* fix(ci): isolate the Herdr restart fixtures from a claimed worktree
The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.
The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
* feat(pi): resolve extension-registered providers in the supervision branch (#3871)
* Let the supervision branch resolve extension-registered providers
The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".
Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.
Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.
* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc
* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered
* ci: retrigger flaky Herdr/serial-1 lanes
* no-mistakes(document): docs already cover branch extension-provider copy
* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)
* fix(watch): detect stalled secondmate queue progress
* no-mistakes(review): gate secondmate stall on active turns and progress episodes
* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector
* no-mistakes(document): clarify active-turn gate in wake-stall config docs
* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree
* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
* test(calm): harden the export-DOM render step and record Pi 0.85.1 evidence (#3952)
* fix(tests): make the Calm export-DOM render step retry and report
The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.
What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.
Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.
The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo
* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change
* no-mistakes(review): tests: drop inert chrome flags, report render timeouts
* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): make every counted wake queue row presentable or retired (#3950)
* fix(bin): make every counted wake queue row presentable or retired
A wake row could be counted as queued while no drain would ever present
it, leaving the operator told to "drain them before anything else" by a
command that printed nothing and offered no acknowledgement.
Two independent paths produced that state.
A row reserved by a live supervision-branch grant is excluded from a main
drain by design, but fm-guard.sh counted the whole queue, so main was
warned about rows only the branch could present, on every guarded command
for as long as the grant was held.
A row that lost its five appended fields or its numeric sequence can never
be claimed, presented, or named by an --ack-through cutoff, yet it still
counted as queued, wedging the queue permanently.
The guard now counts only the rows the calling actor can itself present or
retire, and a main drain retires unusable rows under the queue lock,
reporting them in bounded escaped form before removal so the evidence
survives for the separate row-generating defects. A retirement failure is
reported loudly and never suppresses unrelated consumable work. A main
drain whose remaining rows are all branch-held says so in one bounded line
instead of exiting silently. Grant row-list and owner-record reads move
into fm-wake-lib.sh so the drain, the grant publisher, and the guard share
one implementation.
Ownership is unchanged: a branch drain still touches nothing outside its
grant and never retires a row, and main still cannot present or acknowledge
an active grant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1
* no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass
* no-mistakes(review): add guard advisory for branch-held wake rows
* no-mistakes(document): document per-actor wake counting and unusable-row retirement
* no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added
---------
Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use system stat for Darwin BSD formats (#3305)
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing
* fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main
* test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state
* ci: re-trigger after Chrome headless timeout in calm HTML export test
* ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test
* test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat
* no-mistakes(document): Refresh stat and shard docs
* fix(spawn): carry attribution-off policy in every claude launch (#3945)
The captain's attribution policy (no Co-Authored-By trailer, no
Claude-Session link, no generated-with line) lives in Claude Code's `user`
settings scope. A spawned worker's settings sources are not guaranteed to
load that scope, so a launched worker could write attribution trailers into
its commits and PR bodies regardless of the captain's own configuration.
launch_template()'s claude case now carries the same policy
("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in
its inline --settings JSON, so every claude launch keeps attribution off
independent of which settings scopes end up loaded. Tests assert the policy
on the rendered launch command for both a crewmate and a secondmate spawn.
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
* fix(bin): derive watcher beacon staleness grace from poll cadence (#3946)
* fix(bin): derive the away-mode beacon grace from the poll cadence
fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.
Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.
fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.
Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.
* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh
* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace
---------
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
* fix(procevent): reap orphaned runners and prevent launch storms (#3904)
* fix(procevent): bind a source runner to the session that owns it
A process-event source runner is detached into its own process group so a
persistent source survives the turn that armed it. Nothing bounded that
detachment, so a runner could reparent to init and keep its blocking child -
and every process that child spawned - running with nothing left to reap it.
One such runner outlived its home for about a day; the cost was not the runner
but the exec churn of the poll stubs under it, which stalled every fresh
process launch on the host.
Each runner now starts a small guard beside it, in a separate process group,
that re-reads its home's process-event lease and stops the runner's whole
process group once that lease can no longer be proved fresh. Every ordinary
entry point an owning session runs refreshes the lease, and the watcher's
reconcile cycle keeps it fresh in a live home; nothing a runner spawns can
refresh it, so a source cannot certify its own owner. Scope is the owning state
root and one runner generation, never a script or process name, so a live
source in another home is untouched and a live home simply starts a
replacement runner on its next cycle.
The test scaffolding that starts real runners could not reap them either: the
bearings-board and board-render suites tracked their homes in a shell array
appended to inside a command substitution, so the array was always empty and
every listener they started survived the run. Home registration moves to a
`$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path,
now including HUP and QUIT, and the blocking fixture stubs stop themselves at a
bound so an escaped one cannot keep spawning processes indefinitely.
Adds a regression test that reproduces the orphan shape - a reparented listener
with a live descendant tree under it - and proves the whole group and its
process churn stop once its session is gone, that an identical listener in a
home whose session is still there is untouched, and that retirement still
reaches a reparented listener and everything under it.
* no-mistakes(review): Bound source launches and fail closed on guard startup
* no-mistakes(document): Document runner lease and storm containment
* no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite
* test(procevent): pin that an unprovable stop is retried, not abandoned
The owner guard used to call stop_runner_pid and exit unconditionally, so a
stop it could not prove - a descendant still finishing uninterruptible work
outlives even the group signal, and an unreadable process identity proves
nothing - left a still-running expired runner with nothing watching it. That is
the best-effort reaping this mechanism exists to remove, and the fix that made
the guard retry landed without a test holding it in place.
The unprovable attempt is injected through the signal the real path actually
reads: `ps` answers exactly one process-group query for the runner with a group
it does not lead, which is how a stop that cannot be proved is reported, and
every other call is the real command. The test also asserts that the injected
attempt happened, so it cannot pass vacuously if the fixture stops arming.
Fails against the exit-after-one-attempt guard, where the runner survives its
expired lease, and passes once the guard retries on its check cadence.
* docs(procevent): scope the no-self-refresh rule to confused-agent grade
The runner-lease documentation asserted as an absolute that nothing a runner
spawns can refresh the lease, so a source cannot certify its own owner.
That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually
enforces.
The marker holds at confused-agent grade: a runner and its ordinary children
inherit it and skip every refresh, which is exactly the accidental case this
boundary exists for.
A source that deliberately strips the marker from its environment can still
refresh, so adversarial-grade unforgeability is explicitly out of scope and
tracked as separate follow-up design work.
This states the real scope in docs/configuration.md, which owns the operating
contract, and corrects the two matching comments in bin/fm-procevent.sh.
The process-event-sources skill keeps its cross-reference and gains one line
in its never-to-be-claimed list so the overclaim is not reintroduced from the
agent-facing side.
The lease mechanism itself is unchanged.
* no-mistakes(review): Fix process-event lease and launch pacing edge cases
* no-mistakes(review): Scope launch pacing and clarify lease boundaries
* no-mistakes(review): Reap leftover groups and use monotonic launch pacing
* no-mistakes(review): Keep guards alive across runner PID reuse
* no-mistakes(review): Use monotonic leases and simplify launch generation identity
* no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards
* no-mistakes(review): Preserve active pacing state on failed registration
* no-mistakes(review): Reap reused runner groups with registration evidence
* no-mistakes(review): Avoid ambiguous group kills and encode pacing identities
* no-mistakes(review): Abort kill escalation after runner identity reuse
* no-mistakes(review): Gate group signals and prune stale pacing state
* no-mistakes(review): Document bounded PID reuse signaling safety
* no-mistakes(review): Align leaderless group ambiguity guidance
* no-mistakes(review): Expire reboot stamps and preserve publication success
* no-mistakes(review): Bind owner leases to physical state roots
* no-mistakes(document): Clarify process-event lease and pacing contracts
* fix(procevent): drop a platform-dependent post-TERM test assertion
CI ran red on two lanes that the local gate could not see.
Lint failed with SC2034 on two reads in cmd_owner_watchdog that
destructure the state-root identity into five fields while using only the
device and inode.
Local changed-file mode suppresses the cross-file codes that need
--external-sources, so the warning cleared the pre-push lint step and
failed CI's full analysis, exactly as bin/fm-lint.sh's header describes.
The unused fields now read into `_`.
The behavior shard failed on this suite's own post-TERM assertion, which
required the stubbed identity source to be consulted more than once.
Whether that happens is platform-dependent: where the runner leader keeps
waiting on its TERM-ignoring source child, the post-TERM check sees a live
leader whose identity no longer matches, and where the leader dies
promptly it sees a leaderless group carrying the same numeric id.
fm_procevent_pid_state reaches that second verdict without consulting
process identity at all, so the identity source is never read twice and
the count assertion fails through no fault of the behavior.
The case now asserts the invariant both forms share: retirement refuses,
and the ambiguous group is not signalled.
Scoping a mutation to this fixture and making the refusal signal instead
confirms the case still fails, so dropping the count does not leave it
passing vacuously.
* no-mistakes(review): Prevent superseded runners recreating stale pacing stamps
* no-mistakes(document): Document pacing and ambiguity boundaries
* fix(procevent): retire under the recorded identity source and state the home-scoped lease
The reused-group case started its runner with the proc-root override in
place, so the runner recorded a ps-derived identity, then retired it
without that override.
Where /proc exists the retirement read identity from a different source
than the one recorded, the guard correctly refused an identity it could
not confirm, and the case failed on Linux while passing on macOS.
It now retires under the same source, and clearing the stub marker first
turns that cleanup into the complementary assertion: once the ambiguity is
gone, retirement reaps the whole group instead of leaving it behind.
The lease prose claimed a runner is bound to the session that owns it,
while the mechanism binds it to the home.
That gap is what makes a replacement session or an inspection command look
like a defect: any activity in the same home refreshes the lease.
The granularity is deliberate, because a persistent source is meant to
outlive the session that armed it, and binding a runner to that session
would stop the sources this mechanism exists to keep running.
A runner whose source is no longer wanted in a live home is stopped by
reconcile when that source is retired, independently of the lease, so the
lease is the backstop for a home that is gone - the torn-down sandbox this
change bounds - and the residual is recorded as a known limit.
* no-mistakes(review): Rate-limit polls and skip superseded runner launches
* fix(procevent): build the claim-only sweep case as a runnerless owned claim
A superseded generation now observes the registration-identity mismatch,
self-retires, and releases its claim, which is the behavior we want: it
clears its own residue rather than leaving a claim with no runner for the
home sweep to find.
The claim-only sweep case was built by deleting a registration out from
under a live runner, which used to leave that runner in place. It now
makes the runner retire itself, so the sweep raced that exit and retired
one source or two depending on which won. The case failed three runs in
four, alternating between a preflight-count failure and `attempted=1`.
It now builds the state it means to test: kill the runner's group so it
cannot run its own cleanup, assert the owned claim survived that kill, and
only then drop the registration. Coverage is unchanged - a runnerless
owned claim must still be swept - and the result no longer depends on
whether the runner had exited yet. Three consecutive runs pass.
The superseded exit also skipped the runner-marker cleanup the normal path
performs. The marker is written before the launch floor is waited on, and
a home sweep counts a marker with no owned claim as a preflight failure,
so exiting without clearing it would make that home refuse to sweep.
`FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI
runs, though not under the changed-file mode the pre-push gate uses.
* test(procevent): retire a quiet reparented listener instead of racing a storm
Explicit retirement was exercised against the spawn-churning stub, which
made it nondeterministic. Retirement refuses rather than signalling when it
cannot confirm the runner's identity, that identity is read through `ps`,
and the stub's 0.1s spawn loop starves that read often enough that a single
attempt is a race - the suite failed on this case roughly one run in four,
reporting `cannot confirm runner identity; source remains registered`.
The refusal is correct: it is the documented preserve-for-retry contract,
and a separate case already asserts it. So this is a fixture problem, not a
behavior problem.
The storm is still covered where the evidence for it lives. The owner-loss
home keeps the churning stub and still asserts its tick log stops, which is
what proves the churn ended rather than one pid going away. The retirement
home never asserted ticks; it only ever read the descendant pid, so the
spawn loop bought this case nothing while costing it determinism.
It now uses a quiet stub that still reparents and still holds a real
descendant in its process group, so the assertions are unchanged: retiring
the source must reap the reparented listener's whole group and the
descendant under it. Four consecutive runs pass.
* no-mistakes(review): Serialize registration replacement through source child launch
* chore(no-mistakes): require honest test-step scenario marking
The test step recorded scenarios as passing that were only reached through
a stubbed dependency or the executable suite, and its validator refused
them, because `pass` asserts a scenario was verified against the real live
product.
That refusal is correct, so the fix is to mark honestly rather than to
weaken the gate: a scenario driven live stays a pass and cites its live
transcript, while one reached only through a stub or the suite is recorded
as untested with the reason and a pointer to its executable coverage.
Untested scenarios are reported rather than treated as failures, so real
coverage stays visible without claiming verification that did not happen.
The instruction also forbids dropping a scenario to avoid marking it
untested, since that would hide the gap instead of stating it.
* no-mistakes(review): Remove unrelated test scenario policy
* no-mistakes(document): Clarify process-event home lease documentation
* no-mistakes(document): Correct owner guard failure wording
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority re…
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.
NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.
Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.
Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.
* fix(procevent): repair our decimal-interval regression and enforce the timing phase
REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.
Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".
THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.
Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.
Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.
PROVEN BY MUTATION, each one run against this code:
- remove the decimal normalisation -> the interval case fails on 08;
- a full interval between the two reads -> "the guard exceeded its bound:
group still running 17.1s ... against a documented bound of 15s";
- a full interval WITH startup forced to ~2.5s, which is exactly the condition
the old construction pin could not survive -> still red, same message;
- the same ~2.5s startup with the correct guard -> still passes, 13.0s against
the 15s bound, so the delay alone does not break the case;
- phase evidence made unavailable -> "could not establish the required
pre-expiry guard-read phase", a refusal rather than a pass, even though the
group stopped quickly;
- act on one failed read -> the debounce case fails and the bound case passes
FASTER, 9.5s against 12.7s, which is why these remain separate cases.
Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.
* fix(document): Correct process-event timing and debounce comments
* fix(backlog): route lifecycle transitions through configured adapters (#3417)
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` clos…
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(bin): derive watcher beacon staleness grace from poll cadence (#3946)
* fix(bin): derive the away-mode beacon grace from the poll cadence
fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.
Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.
fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.
Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.
* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh
* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace
---------
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
* fix(procevent): reap orphaned runners and prevent launch storms (#3904)
* fix(procevent): bind a source runner to the session that owns it
A process-event source runner is detached into its own process group so a
persistent source survives the turn that armed it. Nothing bounded that
detachment, so a runner could reparent to init and keep its blocking child -
and every process that child spawned - running with nothing left to reap it.
One such runner outlived its home for about a day; the cost was not the runner
but the exec churn of the poll stubs under it, which stalled every fresh
process launch on the host.
Each runner now starts a small guard beside it, in a separate process group,
that re-reads its home's process-event lease and stops the runner's whole
process group once that lease can no longer be proved fresh. Every ordinary
entry point an owning session runs refreshes the lease, and the watcher's
reconcile cycle keeps it fresh in a live home; nothing a runner spawns can
refresh it, so a source cannot certify its own owner. Scope is the owning state
root and one runner generation, never a script or process name, so a live
source in another home is untouched and a live home simply starts a
replacement runner on its next cycle.
The test scaffolding that starts real runners could not reap them either: the
bearings-board and board-render suites tracked their homes in a shell array
appended to inside a command substitution, so the array was always empty and
every listener they started survived the run. Home registration moves to a
`$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path,
now including HUP and QUIT, and the blocking fixture stubs stop themselves at a
bound so an escaped one cannot keep spawning processes indefinitely.
Adds a regression test that reproduces the orphan shape - a reparented listener
with a live descendant tree under it - and proves the whole group and its
process churn stop once its session is gone, that an identical listener in a
home whose session is still there is untouched, and that retirement still
reaches a reparented listener and everything under it.
* no-mistakes(review): Bound source launches and fail closed on guard startup
* no-mistakes(document): Document runner lease and storm containment
* no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite
* test(procevent): pin that an unprovable stop is retried, not abandoned
The owner guard used to call stop_runner_pid and exit unconditionally, so a
stop it could not prove - a descendant still finishing uninterruptible work
outlives even the group signal, and an unreadable process identity proves
nothing - left a still-running expired runner with nothing watching it. That is
the best-effort reaping this mechanism exists to remove, and the fix that made
the guard retry landed without a test holding it in place.
The unprovable attempt is injected through the signal the real path actually
reads: `ps` answers exactly one process-group query for the runner with a group
it does not lead, which is how a stop that cannot be proved is reported, and
every other call is the real command. The test also asserts that the injected
attempt happened, so it cannot pass vacuously if the fixture stops arming.
Fails against the exit-after-one-attempt guard, where the runner survives its
expired lease, and passes once the guard retries on its check cadence.
* docs(procevent): scope the no-self-refresh rule to confused-agent grade
The runner-lease documentation asserted as an absolute that nothing a runner
spawns can refresh the lease, so a source cannot certify its own owner.
That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually
enforces.
The marker holds at confused-agent grade: a runner and its ordinary children
inherit it and skip every refresh, which is exactly the accidental case this
boundary exists for.
A source that deliberately strips the marker from its environment can still
refresh, so adversarial-grade unforgeability is explicitly out of scope and
tracked as separate follow-up design work.
This states the real scope in docs/configuration.md, which owns the operating
contract, and corrects the two matching comments in bin/fm-procevent.sh.
The process-event-sources skill keeps its cross-reference and gains one line
in its never-to-be-claimed list so the overclaim is not reintroduced from the
agent-facing side.
The lease mechanism itself is unchanged.
* no-mistakes(review): Fix process-event lease and launch pacing edge cases
* no-mistakes(review): Scope launch pacing and clarify lease boundaries
* no-mistakes(review): Reap leftover groups and use monotonic launch pacing
* no-mistakes(review): Keep guards alive across runner PID reuse
* no-mistakes(review): Use monotonic leases and simplify launch generation identity
* no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards
* no-mistakes(review): Preserve active pacing state on failed registration
* no-mistakes(review): Reap reused runner groups with registration evidence
* no-mistakes(review): Avoid ambiguous group kills and encode pacing identities
* no-mistakes(review): Abort kill escalation after runner identity reuse
* no-mistakes(review): Gate group signals and prune stale pacing state
* no-mistakes(review): Document bounded PID reuse signaling safety
* no-mistakes(review): Align leaderless group ambiguity guidance
* no-mistakes(review): Expire reboot stamps and preserve publication success
* no-mistakes(review): Bind owner leases to physical state roots
* no-mistakes(document): Clarify process-event lease and pacing contracts
* fix(procevent): drop a platform-dependent post-TERM test assertion
CI ran red on two lanes that the local gate could not see.
Lint failed with SC2034 on two reads in cmd_owner_watchdog that
destructure the state-root identity into five fields while using only the
device and inode.
Local changed-file mode suppresses the cross-file codes that need
--external-sources, so the warning cleared the pre-push lint step and
failed CI's full analysis, exactly as bin/fm-lint.sh's header describes.
The unused fields now read into `_`.
The behavior shard failed on this suite's own post-TERM assertion, which
required the stubbed identity source to be consulted more than once.
Whether that happens is platform-dependent: where the runner leader keeps
waiting on its TERM-ignoring source child, the post-TERM check sees a live
leader whose identity no longer matches, and where the leader dies
promptly it sees a leaderless group carrying the same numeric id.
fm_procevent_pid_state reaches that second verdict without consulting
process identity at all, so the identity source is never read twice and
the count assertion fails through no fault of the behavior.
The case now asserts the invariant both forms share: retirement refuses,
and the ambiguous group is not signalled.
Scoping a mutation to this fixture and making the refusal signal instead
confirms the case still fails, so dropping the count does not leave it
passing vacuously.
* no-mistakes(review): Prevent superseded runners recreating stale pacing stamps
* no-mistakes(document): Document pacing and ambiguity boundaries
* fix(procevent): retire under the recorded identity source and state the home-scoped lease
The reused-group case started its runner with the proc-root override in
place, so the runner recorded a ps-derived identity, then retired it
without that override.
Where /proc exists the retirement read identity from a different source
than the one recorded, the guard correctly refused an identity it could
not confirm, and the case failed on Linux while passing on macOS.
It now retires under the same source, and clearing the stub marker first
turns that cleanup into the complementary assertion: once the ambiguity is
gone, retirement reaps the whole group instead of leaving it behind.
The lease prose claimed a runner is bound to the session that owns it,
while the mechanism binds it to the home.
That gap is what makes a replacement session or an inspection command look
like a defect: any activity in the same home refreshes the lease.
The granularity is deliberate, because a persistent source is meant to
outlive the session that armed it, and binding a runner to that session
would stop the sources this mechanism exists to keep running.
A runner whose source is no longer wanted in a live home is stopped by
reconcile when that source is retired, independently of the lease, so the
lease is the backstop for a home that is gone - the torn-down sandbox this
change bounds - and the residual is recorded as a known limit.
* no-mistakes(review): Rate-limit polls and skip superseded runner launches
* fix(procevent): build the claim-only sweep case as a runnerless owned claim
A superseded generation now observes the registration-identity mismatch,
self-retires, and releases its claim, which is the behavior we want: it
clears its own residue rather than leaving a claim with no runner for the
home sweep to find.
The claim-only sweep case was built by deleting a registration out from
under a live runner, which used to leave that runner in place. It now
makes the runner retire itself, so the sweep raced that exit and retired
one source or two depending on which won. The case failed three runs in
four, alternating between a preflight-count failure and `attempted=1`.
It now builds the state it means to test: kill the runner's group so it
cannot run its own cleanup, assert the owned claim survived that kill, and
only then drop the registration. Coverage is unchanged - a runnerless
owned claim must still be swept - and the result no longer depends on
whether the runner had exited yet. Three consecutive runs pass.
The superseded exit also skipped the runner-marker cleanup the normal path
performs. The marker is written before the launch floor is waited on, and
a home sweep counts a marker with no owned claim as a preflight failure,
so exiting without clearing it would make that home refuse to sweep.
`FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI
runs, though not under the changed-file mode the pre-push gate uses.
* test(procevent): retire a quiet reparented listener instead of racing a storm
Explicit retirement was exercised against the spawn-churning stub, which
made it nondeterministic. Retirement refuses rather than signalling when it
cannot confirm the runner's identity, that identity is read through `ps`,
and the stub's 0.1s spawn loop starves that read often enough that a single
attempt is a race - the suite failed on this case roughly one run in four,
reporting `cannot confirm runner identity; source remains registered`.
The refusal is correct: it is the documented preserve-for-retry contract,
and a separate case already asserts it. So this is a fixture problem, not a
behavior problem.
The storm is still covered where the evidence for it lives. The owner-loss
home keeps the churning stub and still asserts its tick log stops, which is
what proves the churn ended rather than one pid going away. The retirement
home never asserted ticks; it only ever read the descendant pid, so the
spawn loop bought this case nothing while costing it determinism.
It now uses a quiet stub that still reparents and still holds a real
descendant in its process group, so the assertions are unchanged: retiring
the source must reap the reparented listener's whole group and the
descendant under it. Four consecutive runs pass.
* no-mistakes(review): Serialize registration replacement through source child launch
* chore(no-mistakes): require honest test-step scenario marking
The test step recorded scenarios as passing that were only reached through
a stubbed dependency or the executable suite, and its validator refused
them, because `pass` asserts a scenario was verified against the real live
product.
That refusal is correct, so the fix is to mark honestly rather than to
weaken the gate: a scenario driven live stays a pass and cites its live
transcript, while one reached only through a stub or the suite is recorded
as untested with the reason and a pointer to its executable coverage.
Untested scenarios are reported rather than treated as failures, so real
coverage stays visible without claiming verification that did not happen.
The instruction also forbids dropping a scenario to avoid marking it
untested, since that would hide the gap instead of stating it.
* no-mistakes(review): Remove unrelated test scenario policy
* no-mistakes(document): Clarify process-event home lease documentation
* no-mistakes(document): Correct owner guard failure wording
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so r…
* test: make timestamp fixtures portable across macOS and Linux (#4037)
* test(lib): set fixture mtimes through one portable epoch helper
On macOS the visible symptom was ONE red case in the turn-end guard suite. The
actual damage was TWO cases that had quietly stopped testing their subject. The
red one was the harmless half - people read one red case as one broken thing,
and here that intuition is wrong.
`touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves
the file at its current mtime. So on macOS the three away-mode beacon cases
never aged their beacon at all.
The 400s case exists to pin that 400s is stale under the flat 300s default but
fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it
guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was
worth on this platform:
pre-fix input (beacon left at now): ok - passes with the feature DELETED
post-fix input (beacon 400s old): not ok - expected exit 0, got 2
It was green while measuring nothing, and could not have caught a regression in
the grace it names. Only the 700s case broke loudly.
`touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the
only host-specific step left is formatting the epoch into that stamp, which
date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that
probe once and fails loudly rather than leaving an unset timestamp behind - the
failure mode that caused this. Verified on BSD touch/date here and on GNU
coreutils 9.7 in a container.
Three real sites, and one consistency change - not four fixes. The stale
destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect:
its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and
BSD touch accepts that space-separated form anyway. `touch -t` takes the date
directly on both platforms, so the branch goes rather than standing as a second
copy of the same platform assumption.
Known limit: the third case (away mode off) is only HALF recovered here. It now
receives the input its name claims, but it is still insensitive after this fix -
its verdict is identical with a 0s and a 400s beacon, because the fixture
records a daemon lock and no watcher lock, and with away mode off the daemon
lock proves nothing. Not fixed here; tracked separately, with the requirement
that any fix be shown to FAIL when the protection is removed.
FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The
turn-end guard and remote-handoff suites are clean. Attribution was established
by running the four failures at the base commit and at this head on an idle
machine, because base-idle against head-under-load moves two variables at once:
script head/loaded base/idle head/idle verdict
fm-calm-pi-extension red red red pre-existing
fm-backlog-atomicity red red red pre-existing
fm-procevent red red red pre-existing
fm-startup-network red green green cause unestablished,
load-sensitive under
a full run
Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS
because Chrome is absent instead of declaring the capability it needs and
standing aside, so its verdict is about the machine rather than its subject -
the same family as the defect above, with the red at least announcing itself.
Neighbouring class, reported not changed: `file_mode()` - a verbatim
`uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least
five test scripts plus a `reread_mode` variant, and epoch-mtime reads are
open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as
the defect above. `git init` without `-b main` depends on the host's
init.defaultBranch in several scripts (branch-name case, tracked elsewhere).
timeout, sha256sum and sed -i uses are all correctly guarded where checked.
Observed while building the check rather than the fix: the first watcher I
wrote to wait for the suite matched its own command line, so it was waiting on
its own existence and could never fire. Same shape as the cases above -
machinery answering confidently about something other than its subject, by
including itself in the evidence it was meant to judge. The file sentinel it
was replaced with cannot be produced by the observer that reads it.
* fix(review): Pin fixture timestamps to UTC across DST transitions
* fix(document): Clarify shared fixture suite coverage
* feat(pi): accept native Codex ultra effort with progress-aware supervision (#4038)
* fix(pi): preserve native Codex effort and guarded supervision
* test(pi): identify native compatibility guard versions
* no-mistakes(review): share native-main follow rule between build and picker
* no-mistakes(document): document native progress marker and ultra effort owners
* no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file
---------
Co-authored-by: Talon Stark <talonstark@gmail.com>
* fix(herdr): bypass stale clients rejected by running servers (#4041)
* fix(herdr): step around a stale client the running server refuses
A remote host can carry a self-updated herdr in ~/.local/bin beside a
package-managed one, and the fixed remote-job PATH resolves ~/.local/bin
first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2
client (protocol 20) was answered with protocol_mismatch on every command,
which the read classifiers folded into `unreadable`: the live remote
secondmate read unknown, every doorbell into it failed, and both the spawn
and relaunch recovery paths refused, so the defect trapped itself.
The adapter's session-scoped CLI wrapper now recognizes that refusal, reads
status per session from each distinct herdr on PATH, adopts the first one
the running server reports compatible, retries once, and keeps it for the
process. The happy path makes no extra call and no other failure reselects.
An endpoint that still reads unreadable names the refused client, both
protocols, and the fix on stderr; the remote state read, fm-crew-state, and
the launch refusal carry that reason, and fm-remote-doctor reports the
selected client and rebinds the launch agent to it.
Regression coverage: fake two-client hosts in the herdr unit suite, the
doctor suite, the crew-state remote arm, and the real host-local control
script in the remote lifecycle e2e; the real-herdr smoke refreshes the
status shape the selection reads.
* no-mistakes(review): Reselect Herdr client after every protocol mismatch
* no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding
* no-mistakes(review): Scope cached Herdr clients to their selected session
* no-mistakes(review): Restrict herdr client selection to reactive CLI calls
* no-mistakes(document): Document session-scoped Herdr client reselection
* no-mistakes(document): Clarify Herdr client selection documentation
* no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat: add durable AFK posture lifecycle (#4048)
* feat(afk): record the away posture and its lifecycle (phase 1)
Away mode becomes a posture of the one supervision session, recorded in
state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the
record schema, the mandate-clause grammar and compiler, refusal naming the
missing part, the read-back rendering, the entry announcement (hold-for-return
only, no phone channel), and the archive at return. This release records
clauses and does not execute them; the announcement and return brief say so.
bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any
daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives
the record last on stop. bin/fm-afk-return.sh snapshots supervisor health
before shutdown, renders the return brief (health, mandate, waiting on the
captain, could not fix, handled, cost) from the archived record, the outcome
store, the held set, and the status logs, and shrinks the blocker gate to what
the away session could not fix.
While the record exists the watcher and the daemon never recheck an item held
for the captain. Declared external waits get a four-hour default cadence and
honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by
FM_PAUSE_UNTIL_MAX_SECS.
The /afk skill, AGENTS.md's layout and away-mode stub, the session-start
digest, and the architecture, Pi branch, configuration, and scripts docs
describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a
real Pi primary; its verification record carries the 2026-09-08 run.
* no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating
* no-mistakes(review): Harden AFK authority and posture lifecycle
* no-mistakes(review): Preserve AFK history and tighten authority grammar
* refactor(afk): record clause fields with no natural-language parser
By the captain's mandate the away-posture record keeps no static parser
that tries to understand natural language. A mandate clause is now given
as explicit fields (--action, --object, --when, optional --stop) that
bin/fm-afk-contract.sh records verbatim. The structural check asserts
only that the action, object, and precondition fields are present and
that the action is a listed verb; whether a precondition holds is the
supervision session's judgment at execution time in a later phase.
The never-set stays as a forbidden-concept safety scan: fields mentioning
credentials, passwords, logins, legal or financial acceptance, payments,
invoices, one-time codes, or an attended prompt are refused, matched at
token prefixes after punctuation normalization so compound and plural
spellings are caught. The red-check grammar, class-word rejection,
unconditional-word detection, clause-reference resolution, and condition
aliases are removed. --words-file keeps the captain's words verbatim,
trailing newline included.
The skill, docs, launcher help, and tests describe the field form.
* no-mistakes(review): Preserve AFK words and tighten safety refusals
* no-mistakes(review): Preserve clause bytes and honor declared waits
* no-mistakes(review): Harden deny-list and gate unreadable outcomes
* no-mistakes(review): Demote never-set scan and clarify authority
* no-mistakes(review): Gate return on unreadable held and status data
* no-mistakes(review): Validate posture archives and enforce Pi detection
* fix(afk): make the never-set a non-refusing flag and keep return fail-safe
Per the captain's decision the never-set scan is a coarse best-effort
flag, never a refusal and never the gate: a clause naming a listed
concept is still recorded with a flag the read-back, announcement, and
return brief show, and the scan matches listed terms exactly or with a
plain inflection at punctuation-delimited token boundaries, so unrelated
names such as ping-service or tokenize-worker are never flagged and
joined compounds remain a documented miss. Authoritative never-set and
forbidden-action enforcement is the supervision session's judgment at
execution time in phase 4.
A replacement copies the superseded record through a temporary name and
renames it atomically so a failed copy leaves no partial archive, the
record owner gains validate and flags subcommands, and the return keeps
catch-up gated when a superseded archive cannot be read.
* no-mistakes(review): Harden AFK record validation and return reconciliation
* no-mistakes(review): Harden AFK record validation and simplify commands
* no-mistakes(review): Harden mandate validation and retain missing records
* no-mistakes(review): Refuse blank explicit mandate stops
* no-mistakes(review): Recover restored posture epoch before return
* no-mistakes(review): Prevent return brief status symlink reads
* no-mistakes(document): Refresh AFK posture documentation
* no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks
* no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
* feat(bin): add IMAP/SMTP mail plane with standing poll (#3765)
Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off.
Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix: launch remote Herdr through the user login shell (#4061)
* fix(remote): start the fm-remote Herdr agent through a login shell
Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot.
* fix(remote): start fm-remote Herdr via the account login shell
Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render.
* no-mistakes(review): Fix launch-agent shell fallback resolution
* no-mistakes(review): Preserve and escape Directory Services shell paths
* no-mistakes(document): Document login-shell LaunchAgent behavior
* no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check
* no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks
* no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass
* no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
* fix: distinguish landed deliveries from resolved captain calls (#3710)
* fix(bearings): keep captain-approved deliveries in Recently Landed
A closed task is never held: tasks-axi clears the held flag when a task
closes and keeps hold-kind and the hold reason as the record of the call
that was made. Recently Landed excluded every Done row whose hold-kind was
captain, so the marker it treated as "closed while still waiting on the
captain" was in fact the proof that the captain had approved the work. Every
merge routed through a captain decision disappeared from the list of what
shipped, including under --all-landed.
The selector now asks whether the closed row delivered something. Recently
Landed is merged PRs, completed scouts, and finished local-only merges, so a
row carrying one of those artifacts belongs there whoever approved it. A
captain question closes with an answer and no artifact of its own, and that
is what still stays out, so an answered question is never rendered as
shipped work.
The same rule was written twice - the bearings projection selects this
home's Done rows and the fleet snapshot selects each secondmate home's Done
rows into the roll-up the same section merges in - which is why one defect
hid deliveries in every home. Both now share bin/fm-landed-lib.sh.
* fix(review): Normalize landed evidence and exclude answered captain questions
* fix(review): Normalize captain delivery evidence across relocated data
* fix(review): Record authoritative delivery provenance with legacy fallback
* fix(review): Harden delivery provenance across forced and pruned completions
* fix(review): Replace premature merge closure with existing release contract
* fix(review): Document provenance-based Recently Landed selection
* fix(document): Align documentation with completion provenance
* fix(lint): Fix targeted ShellCheck warnings
* fix(ci): order the pinned tasks-axi install before its stock-Bash consumers
In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both
stock-Bash consumers, and the Bearings expectation is updated from 49 to 50
tests.
Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup
1/1. Full repository lint and all three workflow validations pass, and
`git diff --check` is clean.
* fix(review): Make completion provenance unambiguous
* fix(review): Make completion verdict authoritative over quoted provenance
* fix(review): Preserve retained artifacts through resumed captain closes
* fix(review): Unify completion provenance ordering across writer and reader
* fix(review): Preserve artifacts across failed captain closes
* fix(review): Refresh v1 assertions; provenance authority remains unresolved
* fix(review): Remove unreliable provenance while preserving landed deliveries
* fix(review): Reject stale home summaries visibly
* fix(review): Restore retained deliverable recording
* fix(review): Match landed artifacts and restore retention documentation
* fix(review): Disambiguate captain calls and restore landed artifact matching
* fix(review): Persist retained report and PR artifacts
* fix(review): Preserve staged artifacts before captain answers
* fix(review): Avoid wedging answers on unsupported report paths
* fix(review): Exclude unreleased captain-held pull requests
* fix(review): Exclude held local-only answers from landed
* fix(review): Preserve retained scout reports across snapshot rendering
* fix(document): Align landed lifecycle documentation with release semantics
* fix(review): Enforce landed artifact-kind ownership
* fix(review): Infer canonical task kinds in snapshots
* fix(review): Require captain-hold release before merges
* fix(review): Qualify merge lifecycle regression evidence
* fix(review): Serialize captain holds with merge operations
* fix(review): Document merge cleanup residuals honestly
* fix(test): Replace vacuous Bearings regression with behavioral cases
* fix(document): Align Bearings verification and merge lifecycle documentation
* fix(review): Serialize merges and exclude captain calls from landed
* fix(review): Harden merge identity and landed selection
* fix(document): Clarify landed selector compatibility filtering
* fix(bin): keep merge entrypoints usable on records without an incarnation
The merge identity guard refused any task record with no spawn_gen field.
That field identifies one exact incarnation, so comparing it across the wait
for the merge lock is what catches a task relaunched while the merge was
queued. Requiring it to be present is a different rule, and it refused every
record written before the field existed: a legacy task could no longer be
merged at all, and five behaviour suites refused before reaching the check
they were written to exercise.
The comparison only needs to notice a change. An absent field is now read as
an empty incarnation and compared like any other value, so a record that
gains, loses, or alters one is still refused, while a record that simply
predates the field merges. An ambiguous or unreadable field stays an error,
because a record that cannot name one incarnation cannot be compared. The
missing-record message each entrypoint had before the guard is restored, so
a genuinely absent record still says so in its own words.
The role partition now precedes reading the record. Refusing the supervision
branch is a statement about the actor, not about the task, so it cannot
depend on a record the wrong actor may not have.
A backlog file that does not exist meant "no longer an open captain call".
For a caller that asked to tell absence apart it now means absent, so a board
card whose home carries no backlog stays visible instead of being dropped as
resolved.
Fixture repositories pin their initial branch instead of inheriting
init.defaultBranch, which resolved to main on a developer machine and master
on a runner, so a fixture naming main failed only in CI.
* fix(review): read local-only note from body; surface pending-close failures
* fix(review): keep kindless local-only landings in Recently Landed
* fix(review): bind local-only note scan to the tasks-axi note line
* fix(review): Guard unavailable captain-hold authority records
* fix(document): Document unreadable authority predicate outcome
* fix(bin): read an absent backlog as absence, not an unreadable record
The merge gate refused every task whose home carries no backlog file. A
backlog that does not exist holds no captain call, so nothing can be held and
the merge is safe; only a backlog that exists and cannot be read may hide a
live hold. Those two states were collapsed into one refusal, which stopped
merges in any home that keeps no backlog.
The predicate now reports a missing backlog file as absence, alongside a row
the backlog does not carry. A record that exists but cannot be read still
leaves by the existing cannot-tell path, which both merge entrypoints already
refuse, so the restrictive direction is unchanged.
That leaves no way to reach the separate unavailable-record result, so the
result and the two branches that handled it are removed rather than left
describing an outcome that can no longer occur. The lifecycle documentation
loses the same claim.
Regressions cover both directions in each entrypoint: a home with a task
record and no backlog merges, and a backlog present but unreadable refuses
without reaching the forge.
* fix(review): Fail closed unreadable backend configuration
* fix(tests): pin the bare origin's initial branch in the remote seed fixture
The fixture created its bare origin with no initial branch, so that
repository's HEAD followed init.defaultBranch while the source repository
pushed the branch fm_git_init_commit pins. On a host that still defaults to
master the two disagreed: the bare origin's HEAD named a branch the push never
created, cloning it warned that the remote HEAD referred to a nonexistent ref
and checked out nothing, and the seed assertion for the cloned README failed.
A machine whose default is already main paired the two by accident and hid it,
which is why the fixture passed locally and failed on the runner.
Pinning the bare origin to the same branch removes the dependency on the
ambient default from both sides. Verified under both conditions: with
init.defaultBranch set to master, and set to main, the suite passes 26 of 26.
* fix(review): Fail closed unreadable user backend configuration
* fix(bin): republish the home summary as v1 and record two load-bearing rules
The published home-summary schema had moved to v3, which routed every
secondmate home still emitting the earlier version to the stale branch: their
landed rows, open decisions and holds all came back empty and their state read
as unknown until each home was updated. The payload never justified that. Its
field set, field order, truncations and the landed array construction are
byte-identical to v1, so only which rows the selector places in landed
differs, and a v1 consumer reads that the same way.
Republishing as v1 removes the rollout regression and, with it, the tolerance
machinery that existed only to soften the bump: the stale-schema predicate,
its two collection branches, the flag and its provenance branch, the omitted
surface that can no longer be reached, and the fixtures and assertions that
covered them.
Two rules that a scope review proposed removing are kept, each now carrying
the reason it exists, because both were measured to be load-bearing:
The artifact-kind ownership clause is what keeps an explicit scout that
recorded no report out of Recently Landed. Without it such a row has none of
the three artifacts, satisfies the compatibility fallback and renders as
shipped work with an empty artifact.
The kind fallback is needed because tasks-axi omits the kind metadata
entirely when a title begins with a canonical keyword. Without it a scout
titled "SCOUT ..." reports no kind, its recorded report stops counting as a
delivery, and it drops out of the section this selector exists to repair.
* fix(bin): move the scout guard note onto the rule and drop two dead pieces
The LOAD-BEARING note sat on an unreachable branch. Measured in both
directions: removing that branch together with the kind-is-not-scout guards
lets an explicit reportless scout into Recently Landed and fails
tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone
leaves that suite passing at 49 assertions. The guards carry the rule, so the
note now sits on them and the unreachable branch is gone. A note pointing a
later reader at the wrong line is the hazard this change corrects elsewhere.
summary_file_has_schema lost its only caller when the stale-schema machinery
was removed, so it goes with it.
* fix(review): Fix legacy report artifacts and canonical keyword boundaries
* fix(review): Update pinned Bearings test count to 56
* fix(document): Clarify landed summary compatibility documentation
* fix(review): Preserve unreadable backend configuration errors
* fix(review): Honor backend resolution errors at existing call sites
* fix(test): Stabilize remote collector tests under host load
* fix(document): Document backend resolution failure contracts
* fix(lint): Suppress intentional deferred probe expansion warnings
* fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass
* fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
* fix(remote): keep fm-remote Herdr servers in the Aqua session (#4090)
* fix(remote): let the Aqua launch agent own the fm-remote Herdr session
A herdr server keeps the macOS audit session of whatever started it, and
only the Aqua login session (gui/<uid>) can read the login keychain
without a prompt. Herdr's SSH remote attach starts the fm-remote server
as its own child when it finds none, wins the socket at boot because sshd
accepts connections before the login session exists, and every claude
pane under that server then gets `security` exit 36, falls back to a stale
plaintext credentials file, and reports "Login expired". launchd's own job
lost the socket on every KeepAlive retry and the doctor still reported the
session ready because it only asked whether any server answered.
- Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start
the server in the foreground when nothing owns the socket, exit 0 when an
Aqua-born server does, and otherwise stop the foreign server, wait for the
socket, and exec the server at once.
- Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner
discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth
markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or
remote-client-bridge ancestry matched on argv[0] and whole arguments).
- Render the agent as the login shell exec'ing the guard with
KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded
job's successful-exit semaphore, and report a session served outside the
Aqua login session as fixable so --fix retakes it through launchd; the
reload waits for an Aqua-born owner rather than any running server.
- Correct the doctor and docs: the launch shell provides environment parity,
the launchd domain provides keychain access.
- Pin the guard's decision table and the doctor's verdicts against real
marker-carrying processes, and record the dated audit-session evidence.
* no-mistakes(review): Verify Aqua ownership through launchd domains
* no-mistakes(document): Document macOS lsof ownership requirement
* fix(bin): prefer a live no-mistakes run over a terminal one (#2881)
* fix(bin): prefer a live no-mistakes run over a terminal one
A worktree can bind to more than one recorded no-mistakes run at once.
The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an
exact-equal commit and a worktree-is-an-ancestor match, but never stated
which wins when both bind, so the tie fell to whichever candidate the
caller reached first.
Observed on a live fleet: a crashed validation daemon left a FAILED run
at the worktree's own commit while the live run that replaced it
validated a descendant commit on the same branch. Bare `axi status`
answers with the most-recently-touched run - the corpse - and it bound by
the equal-commit rule, so every recomputation reported `failed` for a
task whose real run was healthy. The same label had also read `failed`
earlier while the work was genuinely stalled, so the signal was wrong in
both directions.
State the live-over-terminal policy in the matching rule's own contract,
where the equal-commit and ancestor rules already live, and add
fm_nm_run_status_class as the one classifier that decides liveness from a
recorded status word. fm-crew-state.sh applies it on both selection
paths: the runs listing now scans past a terminal row for a live one, and
a terminal `axi status` answer is provisional until the listing has been
asked whether this worktree also has a live run.
Same-liveness-class candidates keep the listing's newest-first
precedence, and a status word the classifier cannot place keeps the
caller's own ordering rather than displacing a known result, so a
single-run task and a task whose runs are all terminal are unchanged.
Regression coverage reproduces the proven case (terminal run at the
worktree's exact commit plus a live run descending from it) and its
runs-list twin; both fail under the old tie-break. Two companion cases
pin the no-widening half - two terminal rows still resolve newest-first,
and a terminal run with no live sibling keeps its full run-step detail -
and both pass before and after the change.
* no-mistakes(review): accept unfetched live sibling anchored at exact worktree head
* docs(bin): name both ledger reads behind the runs-limit setting
The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described
the runs ledger as scanned only by the cross-branch fallback, but the
live-over-terminal fix also consults it as the live-sibling probe behind a
terminal axi status answer. Point the comment at docs/configuration.md as the
setting's owner instead of restating a second copy.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(bin): repair process-event shutdown and tighten guard timing (#4009)
* fix(procevent): make the ordinary stop signal actually stop a runner
The owner guard that shipped in #3904 half-reaps. Against a poll child that
handles the ordinary stop signal and keeps waiting, the guard signals the group,
loses the runner leader to its own signal, then reads that success as a
leaderless group and exits without escalating. It destroys the only proof of
ownership that would have authorised the forced signal, so the survivor becomes
unreachable by retire, reconcile, sweep-home and the guard alike. A guard that
turns a leaking-but-identifiable generation into a permanently unreachable one
is worse than no guard at all.
Two defects, and they hid each other:
- The escalation re-derived ownership from the leader. `runner_group_signal`
now takes a `proved` mode, passed only by the escalation inside the stop that
already proved and signalled that exact generation moments earlier. A leader
dying to our own signal is the ordinary outcome, not fresh ambiguity.
- Every stop held the per-source lock across its wait while the runner's own
exit cleanup waited unboundedly for that same lock. That circular wait was
broken only by the forced signal, so the forced signal silently became the
normal path - and, by keeping the leader alive through the whole window, it
masked the escalation defect above. The runner's exit cleanup now refuses that
lock instead of waiting for it, which is what its existing `return 0` already
said it did.
Fixing the lock alone would have turned every stop of a signal-proof child into
a refusal that leaves it running, so both land together and the tests pin that.
Measured on macOS with a stand-in poll child that traps TERM, INT and HUP:
the guard left it running past 70s and now clears the group within the lease
plus one check; retiring a healthy runner fell from ~2.8s with a forced group
signal every time to ~0.6s on the ordinary signal alone.
Unchanged and stated deliberately: a leader lost to anything other than the
stop's own signal still leaves a group that retire, reconcile, sweep-home and
the guard all refuse, permanently - and that source stops listening without
saying so. Whether such a group may ever be signalled is an open decision and
is not answered here.
* fix(review): Fix proved escalation race and stop regression assertions
* fix(review): Preserve proved escalation through transient identity failures
* fix(review): Simplify proved escalation and correct guard timing documentation
* fix(document): Clarify process-event stop ownership and cleanup limits
* fix(document): Clarify process-event stop ownership and fixture comments
* revert(skills): restore the leaderless-ambiguity limit to the loaded skill
An automatic documentation step in this branch's validation edited
.agents/skills/process-event-sources/SKILL.md, which no instruction in this
change asked it to touch. That file is not documentation about the code: it is
the agent-loaded instruction surface, what an agent reads to know what it is
permitted to do.
The step deleted this line:
- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling
or replacement, as owned by the operating contract in
[`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent);
and folded it, with its neighbour, into a generic "registration and ownership
transitions, stop authority, and claim reclamation follow the operating
contract".
That deleted line states a PROHIBITION - that such a group is preserved WITHOUT
SIGNALLING - and it is the exact limit an open captain decision currently rests
on. Folded into a pointer, an agent reading the skill to learn what it may do
would have to chase a second document to discover it may not signal. A
prohibition that requires a second lookup is not a prohibition. The effect was
to weaken, in the instructions themselves, the boundary that keeps one home from
signalling another's process group - while the question of whether that boundary
should move at all is still open.
This is a deliberate revert, not an oversight, and it restores the file exactly
to its pre-branch state. The full statement also survives in
docs/configuration.md; that does not rescue it, because the agent handling a
process-event wake loads the skill and not the documentation.
* revert(procevent): restore the open-question marking beside the escalation
The same automatic documentation step that edited the loaded skill also removed
this from the comment above runner_group_signal:
A leaderless group nobody in this call ever proved remains refused too, for
every caller. That untouched refusal is what makes a crashed leader's group
permanent, and relaxing it is a separate open question, not something this
path assumes.
and replaced it with a pointer to docs/configuration.md.
This one fails differently from the skill deletion, which is why it is restored
separately. There, a prohibition was moved out of the reader's path, and a
missing prohibition gets violated. Here the prohibition survives in code - the
unproved path still refuses - and what was removed is the fact that the limit is
UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as
settled design, and settled design gets relied on, extended, and eventually
relaxed by someone confident they understand why it is there. That question is
open right now.
The rule this branch's four instances produce, stated once here because this is
the point of decision: an unresolved question must be marked unresolved AT THE
POINT OF DECISION, not only where the contract is documented. A reader who does
not know something is open will treat it as closed, and that default is stronger
than any pointer overcomes.
The pointer added by that step is kept alongside; this restores what it replaced
rather than reverting it.
* docs(verification): restore the measured guard bound and its reason
The document step's rewrite of this record dropped the concrete figure while
keeping the surrounding measurements. What went missing was the bound itself -
lease plus two consecutive failed checks plus the stop's grace, roughly 630
seconds at the shipped 600-second lease and 15-second check - together with the
reason there are two checks rather than one: a single unreadable read must not
be enough to kill a live runner.
The mechanism survived elsewhere and the reason survived in
docs/configuration.md, so nothing was lost from the repository. The concreteness
was, and that is what this restores. A number recorded without why it is that
number is the one a later reader shortens; the reason is the whole safety
argument for the debounce, and the debounce is what stops the reaper killing a
live runner on one bad read.
* fix(document): Replace stale stop-authority summaries with owner pointers
* test(procevent): make the guard-bound case able to fail for its own reason
An automated reviewer observed that this case allowed sixty seconds for a bound
of roughly eight, so it could not go red for the reason it names: it would have
passed a guard that took fifty-five seconds. That is correct, and it is the same
family as the defect the case exists to defend against - a check that is green
because it cannot fail, rather than because the thing it guards is working.
The deadline is now derived from the bound itself - the lease, plus the two
consecutive failed checks the guard debounces on, plus the stop's own ordinary
and forced signal windows - rather than from a flat wall-clock number, and the
shortened lease and check the fixtures run under have a single definition so a
derived deadline cannot silently diverge from the settings the guard is given.
The doubling that remains is a load allowance and is documented as one; widening
it to make a slow guard pass would convert the assertion back into decoration.
Proven by mutation rather than by argument. Against the repaired case:
correct code ok
guard debounces on 20 misses instead of 2 not ok - "still holding
the group after 16s,
against a documented
bound of 8s"
proved escalation removed (the original defect) not ok - same
code restored ok
The previous sixty-second version passes every one of those mutations.
The reviewer's other claim, that the guard can survive past the announced bound
when an owner disappears immediately after a check, was measured and does not
hold against what this branch announces. Sweeping the phase deliberately at
0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s
and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed
checks plus stop grace, which is up to 8s at those settings. The mechanism the
reviewer describes is real and is the announced mechanism; the bound it was
measured against is a phrasing this branch no longer carries.
* revert(scope): return the instruction surfaces to their base state
This delivery is being split. It carries the two proven process fixes alone; the
instruction text travels separately, through a run that removes the
documentation step rather than refusing it at its gate.
Two surfaces are therefore returned to exactly what the base branch has, so this
delivery neither adds to them nor removes from them:
.agents/skills/process-event-sources/SKILL.md - identical to base again. Three
bullets an automatic documentation step had folded into a pointer, including
that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT
SIGNALLING and that there is one identity-matched owner per canonical source
across homes sharing one store.
The header comment block of bin/fm-procevent.sh, which is what the script
prints as its own help. Seven lines were removed from it: that a live owner is
never displaced, that only a claim whose stale owner and independently absent
process group prove its whole generation gone is reclaimed, that a crashed
leader or reused pid whose process group still has members cannot relax
ownership cleanup, and that reconcile signals only a live identity-matched
runner group and otherwise keeps the claim without starting a replacement.
The help output is now byte-identical to base.
Neither removal was requested by any instruction in this change, and both were
made to text that predates it. Returning them is scoping, not a third
restoration: nothing is being added to those files here.
* fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor
* test(procevent): make the post-TERM cases report what they saw when they fail
On the failure path only, these cases now print what they actually saw: the
identity recorded at claim time, the identity readable at that moment, the size
of the signals file, the leader's state and wchan, every live member of the
runner's process group with its own state and wchan, the elapsed time since the
stop began, and what retire said. None of it runs when a case passes.
WHY THIS IS KEPT, stated accurately rather than by its original reason. It was
written to make an unexplained CI failure verifiable. That failure is now
explained - it was a fixture deadlock, diagnosed and repaired in the preceding
commit - so that justification has expired and is not the reason given here.
The reason it stays is smaller and independent of that failure: it is already
written, it is small, it sits in the file whose assertion this change reworked,
and an assertion that could not say why it failed cost most of a morning to
diagnose from the outside. The next failure will not be this one.
WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already
observed many times, not proof that anything is fixed. Only a failure carrying
the evidence above establishes a cause.
* fix(document): Clarify process-event fixture diagnostic rationale
* fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor
* fix(procevent): bound owner-guard cleanup at one check interval, not two
A THIRD WAY, not a capitulation to the reviewer and not a refusal of it.
The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the
number of observations the guard makes before it acts. It asked for a single
read because that was the only route it could see to an acceptable bound. There
was another route, and this change takes it: the bound is reached and both reads
are kept.
TIGHTENED - the SPACING of the guard's two reads, not their number. The owner
watchdog now sleeps half the configured check interval and still requires two
consecutive failing reads, so the pair completes inside one check interval
instead of costing two. Worst-case detection falls from the lease term plus TWO
check intervals to the lease term plus ONE. At the shipped 600s lease and 15s
interval the stated bound falls from ~635s to ~620s.
PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is
untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root
identity read, and either can fail transiently on a live, healthy home. Acting
on the first failure would let one isolated unreadable read kill a live service.
Requiring a second, independent read is what makes that impossible, and it is a
protection rather than padding. Nothing was traded away to reach the bound.
Both properties are now guarded by their own case, and each was proven by
MUTATION rather than asserted:
- putting a full interval back between the two reads fails the bound case:
"still running 17.0s after the last owner activity, against a documented
bound of 15s";
- acting on one failed read fails the new debounce case: "one unreadable lease
read ended a runner whose home was still alive" - while the bound case then
passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one
is exactly why these are two cases: one elapsed-time case would have
registered the removal of the protection as an improvement.
MEASURED, sampling the phase between the guard's check clock and the lease clock
across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned
listener whose home stopped refreshing its lease:
lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after
lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after
The 4s configuration is the informative one: the gap is about one check
interval, which is precisely the term that was removed.
A previously unstated term of the bound surfaced while measuring: the lease age
is compared in whole seconds, so a configured lease of N is honoured until that
age reads N+1. It is now part of the documented bound and of the regression's
derivation instead of being absorbed into a fudge factor.
The bound regression derives its deadline from the documented bound instead of a
flat number, and PINS the phase between the guard's check clock and the lease
clock rather than sampling it, because with a sampled phase a guard spending two
intervals passes about half the time on a lucky alignment. Its load slack is
additive and stays under half a check interval, so an extra whole interval
cannot hide inside it. The two flat deadlines that were there before (40s and
20s) and the doubling allowance on the derived one are gone; that looseness was
the reviewer's third complaint.
The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the
forced one. It is a ceiling paid only by a group that outlives the signal it was
sent, not a delay every stop pays - a healthy runner's whole retire measures
0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is
unreachable by any implementation, since signalling a process and giving it any
chance to exit takes non-zero time; detection now meets it and the stop runs
inside its own ceiling, and the contract says so rather than glossing it.
NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs
start rather than after they report. On the previous head, "Behavior portable
serial 1" and "Behavior portable serial 4" were both CANCELLED at the job
ceiling, independently of this finding. A new head triggers fresh runs, so those
two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING
DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly
and sometimes is not; the shard-packing repair for it is open separately. Do not
reread a lucky pass here as a resolution.
Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh
was NOT updated even though the two new cases add ~19s of wall clock.
docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI
timing artifacts of green runs, and that repair is the open request doing it; a
hand-edited estimate here would collide with it and silently repack the shards.
This suite runs in portable serial shard 3, which was green in the last run.
Verification: tests/fm-procevent.test.sh green, plus
tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and
bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was
removed" case flaked in 4 of 7 local full runs; an isolated 20-trial
reproduction measured it at 13/20 unclean before this change and 11/20 after, so
it is issue 4080 and is not aggravated here.
* fix(procevent): repair our decimal-interval regression and enforce the timing phase
REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added
by the previous commit read a zero-prefixed interval as octal: 010 halved to 4
instead of 5, and 08 was not a number at all, so the owner guard died before
reporting ready and the runner failed closed and never listened. The validator
accepts those values and `[` compares them as decimal, so this broke a
configuration that worked before. Introduced by this delivery, found in review,
repaired here. Forcing base ten before the arithmetic is the whole runtime fix.
Proven by driving it rather than by reading the source: a new case starts a real
listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and
5s. Removing the normalisation turns that case red with "a zero-prefixed decimal
interval (08) prevented the listener from starting".
THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case
pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced
nothing. Review was right that this is not enough: once startup reaches about two
seconds the expiry lands in a different part of the interval and the case
silently stops rejecting a two-interval guard while still reporting success. A
bound that cannot fail for the reason it names is the defect this whole delivery
exists to correct, so it must not ship inside the fix for it.
Now the lease is synchronised to the guard's own FIRST observed lease read,
every later real read is recorded, and the case REFUSES unless one recorded read
proves the required phase: it read the synchronised reference, it was still
fresh, and it began late enough that two further full intervals could not finish
before the deadline. An unestablished precondition refuses; it does not proceed
on trust. The derived deadline, the two-read debounce and the additive slack are
unchanged, and the slack invariant is now asserted rather than left to a comment.
Review also found the deadline was only ever checked while the group was still
alive, so a sampler descheduled past it would see the group gone and certify
success. The observed completion time is now checked too.
PROVEN BY MUTATION, each one run against this code:
- remove the decimal normalisation -> the interval case fails on 08;
- a full interval between the two reads -> "the guard exceeded its bound:
group still running 17.1s ... against a documented bound of 15s";
- a full interval WITH startup forced to ~2.5s, which is exactly the condition
the old construction pin could not survive -> still red, same message;
- the same ~2.5s startup with the correct guard -> still passes, 13.0s against
the 15s bound, so the delay alone does not break the case;
- phase evidence made unavailable -> "could not establish the required
pre-expiry guard-read phase", a refusal rather than a pass, even though the
group stopped quickly;
- act on one failed read -> the debounce case fails and the bound case passes
FASTER, 9.5s against 12.7s, which is why these remain separate cases.
Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean.
* fix(document): Correct process-event timing and debounce comments
* fix(backlog): route lifecycle transitions through configured adapters (#3417)
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* no-mistakes(review): Harden markdown lifecycle routing and close recovery
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting
* no-mistakes(review): validate tasks config before exemption; fix lint quote gap
* fix(backlog): address the markdown backlog as <data>/backlog.md
Resolving the markdown backlog through a configured `[markdown] path` was
scope this task never asked for. It is absent from main, which addresses
`<data>/backlog.md` everywhere, and it came from an earlier review round
rather than the task brief.
Making it effective on the transition path alone put that path at odds
with every other consumer of the same backlog - fm-captain-hold.sh,
fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh,
fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In
fm-captain-hold.sh the split was live: its reads had already moved to the
shared gate while its writes had not, so the two could address different
files.
Address `<data>/backlog.md` from the shared gate, delete the unused
resolver, and drop the two tests that pinned the withdrawn behaviour.
What this task actually changes is unaffected: a configured non-markdown
adapter is still addressed by its own root, without `--file`.
* fix(backlog): honor configured task adapters
* no-mistakes(review): Harden backend purity lint against prefixed Beads calls
* no-mistakes(document): Document configured backend lifecycle transitions
* fix(backlog): preserve markdown exemptions
* no-mistakes(review): Enforce backend purity for explicit lint paths
* no-mistakes(document): Update lifecycle backend documentation
* no-mistakes(lint): Remove redundant backend lint pattern
* fix(backlog): close adapter routing gaps
* no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint
* no-mistakes(document): Document environment-selected backlog adapters
* no-mistakes(lint): Fix empty local variable assignment
* fix(backlog): close quoted path gaps
* no-mistakes(review): Reject partially quoted direct Beads commands
* no-mistakes(document): Align lifecycle documentation with configured adapters
* test(backlog): keep structural cases markdown-only
* fix(lint): catch dollar-quoted beads commands
* no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P
* no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING
* no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases)
* fix(bin): preserve captain calls during teardown (#3595)
* fix(bin): never close a captain call during cleanup
A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. …
Intent
Firstmate documents worker effort as a concrete harness axis, but the shared enum rejects native Codex
ultrabefore Pi can launch it.Native activity inside a long Pi turn also fails to refresh the watcher age bound, and native tool execution cannot discover Firstmate's guarded controls.
Accept
ultraonly for explicitcodex-native/<model>selections on Pi or Pi-signed and emit--codex-effort ultra, never--thinking ultraormax.Preserve that effort through batch dispatch, guarded relaunch, and local or remote secondmate restart; refuse unsupported and raw launches before lifecycle changes.
Refresh a generation-bound progress marker independently of semantic busy state and turn completion, so observable native activity prevents false wedge escalation while stale incarnations remain unable to refresh their replacements.
Expose the registered Firstmate tools and operational-message types through
firstmate:native-tools, retaining the original execute callbacks and ownership checks without exposing Pi built-ins.When the primary itself runs native Codex, its background supervision branch remains an independent ordinary Pi agent.
It follows the same model through
openai-codex, never inherits the main native thread, refuses acodex-nativebranch pin, and returns the wake to main if the ordinary model is unavailable.The branch builder and
/supervision-modelpicker share one resolution rule, so the picker reports the model that will run or the exact refusal.Reproduction and validation
On upstream base
b84e0e362face25f3dd8945297a3df1320d7668c, the native Ultra spawn fails with--effort must be one of low, medium, high, xhigh, max.Running the expanded dispatch test with the base
bin/also fails its new Ultra cases.On this branch, the expanded suite passes: native Pi and Pi-signed launches emit
--codex-effort ultra, metadata retainseffort=ultra, and invalid profiles and raw launches are refused.Targeted spawn, relaunch, restart, bootstrap, busy-state, adapter, watcher, and Pi extension tests pass.
Shell syntax, canonical ShellCheck, workflow lint, strict Pi typechecking, and documentation checks pass.
The default-on native guard passes against installed Pi 0.85.1 and pi-codex-native 0.2.1, including initial, operational, and resumed Ultra turns; observable adapter progress; guarded tool execution; and durable outcome acknowledgement with duplicate refusal.
Its App Server peer and watcher-close process are deterministic fixtures; it does not test a live model or real backend.
The broad portable run completed 151 scripts with six capability skips and the pre-existing macOS fixture failure disclosed below.
A separate live-harness run completed 26 scripts with zero failures and 24 capability or opt-in skips; real Herdr lifecycle checks were not exercised.
The responsiveness guard now copies
fm-native-contract.tsinto its isolated Pi fixture.The missing dependency caused Pi to reject the extension; the corrected fixture loads successfully, and its serial CI job passes.
Baseline disclosures
The unchanged turn-end fixture in
tests/fm-turnend-guard.test.shfails under BSD touch on macOS because GNU-styletouch -d @<epoch>is rejected, leaving the stale-beacon fixture fresh.The complete suite passes with GNU touch selected through an isolated PATH.
The test and guard sources match upstream, and this branch does not change that fixture or fix its portability issue.
The opt-in Pi branch SDK guard also fails identically on upstream main and this branch against Pi 0.85.1: its unpinned session-restore assertion observes
gpt-5.5instead of fixture modelfm-live-a.This branch does not change or fix that pre-existing failure.
Neither baseline failure is reported as a pass.
Upstream reconciliation
The busy-age documentation overlapped newer upstream pause and foreign-queue stall guidance.
The reconciliation preserves upstream pause behavior and the 180-second no-progress foreign-stall bound while adding native progress to the busy-age description.
Updates from git push no-mistakes