feat(bin): sync upstream to batch 19 tip (afk, quiet mode, mail plane) - #53
Merged
Merged
Conversation
…henguid#3904) * fix(procevent): bind a source runner to the session that owns it A process-event source runner is detached into its own process group so a persistent source survives the turn that armed it. Nothing bounded that detachment, so a runner could reparent to init and keep its blocking child - and every process that child spawned - running with nothing left to reap it. One such runner outlived its home for about a day; the cost was not the runner but the exec churn of the poll stubs under it, which stalled every fresh process launch on the host. Each runner now starts a small guard beside it, in a separate process group, that re-reads its home's process-event lease and stops the runner's whole process group once that lease can no longer be proved fresh. Every ordinary entry point an owning session runs refreshes the lease, and the watcher's reconcile cycle keeps it fresh in a live home; nothing a runner spawns can refresh it, so a source cannot certify its own owner. Scope is the owning state root and one runner generation, never a script or process name, so a live source in another home is untouched and a live home simply starts a replacement runner on its next cycle. The test scaffolding that starts real runners could not reap them either: the bearings-board and board-render suites tracked their homes in a shell array appended to inside a command substitution, so the array was always empty and every listener they started survived the run. Home registration moves to a `$$`-keyed registry in tests/lib.sh, which sweeps it from every cleanup path, now including HUP and QUIT, and the blocking fixture stubs stop themselves at a bound so an escaped one cannot keep spawning processes indefinitely. Adds a regression test that reproduces the orphan shape - a reparented listener with a live descendant tree under it - and proves the whole group and its process churn stop once its session is gone, that an identical listener in a home whose session is still there is untouched, and that retirement still reaches a reparented listener and everything under it. * no-mistakes(review): Bound source launches and fail closed on guard startup * no-mistakes(document): Document runner lease and storm containment * no-mistakes(ci): Fixed the Greptile watchdog finding: failed runner cleanup now retries on each watchdog tick instead of abandoning the orphaned process group. The shared state-root lease behavior remains unchanged because it is an explicitly accepted ownership policy. Verified with bash syntax checks, git diff checks, and the complete fm-procevent test suite * test(procevent): pin that an unprovable stop is retried, not abandoned The owner guard used to call stop_runner_pid and exit unconditionally, so a stop it could not prove - a descendant still finishing uninterruptible work outlives even the group signal, and an unreadable process identity proves nothing - left a still-running expired runner with nothing watching it. That is the best-effort reaping this mechanism exists to remove, and the fix that made the guard retry landed without a test holding it in place. The unprovable attempt is injected through the signal the real path actually reads: `ps` answers exactly one process-group query for the runner with a group it does not lead, which is how a stop that cannot be proved is reported, and every other call is the real command. The test also asserts that the injected attempt happened, so it cannot pass vacuously if the fixture stops arming. Fails against the exit-after-one-attempt guard, where the runner survives its expired lease, and passes once the guard retries on its check cadence. * docs(procevent): scope the no-self-refresh rule to confused-agent grade The runner-lease documentation asserted as an absolute that nothing a runner spawns can refresh the lease, so a source cannot certify its own owner. That overclaims what the inherited FM_PROCEVENT_IN_RUNNER marker actually enforces. The marker holds at confused-agent grade: a runner and its ordinary children inherit it and skip every refresh, which is exactly the accidental case this boundary exists for. A source that deliberately strips the marker from its environment can still refresh, so adversarial-grade unforgeability is explicitly out of scope and tracked as separate follow-up design work. This states the real scope in docs/configuration.md, which owns the operating contract, and corrects the two matching comments in bin/fm-procevent.sh. The process-event-sources skill keeps its cross-reference and gains one line in its never-to-be-claimed list so the overclaim is not reintroduced from the agent-facing side. The lease mechanism itself is unchanged. * no-mistakes(review): Fix process-event lease and launch pacing edge cases * no-mistakes(review): Scope launch pacing and clarify lease boundaries * no-mistakes(review): Reap leftover groups and use monotonic launch pacing * no-mistakes(review): Keep guards alive across runner PID reuse * no-mistakes(review): Use monotonic leases and simplify launch generation identity * no-mistakes(review): Prevent pacing identity reuse and bound reused-group guards * no-mistakes(review): Preserve active pacing state on failed registration * no-mistakes(review): Reap reused runner groups with registration evidence * no-mistakes(review): Avoid ambiguous group kills and encode pacing identities * no-mistakes(review): Abort kill escalation after runner identity reuse * no-mistakes(review): Gate group signals and prune stale pacing state * no-mistakes(review): Document bounded PID reuse signaling safety * no-mistakes(review): Align leaderless group ambiguity guidance * no-mistakes(review): Expire reboot stamps and preserve publication success * no-mistakes(review): Bind owner leases to physical state roots * no-mistakes(document): Clarify process-event lease and pacing contracts * fix(procevent): drop a platform-dependent post-TERM test assertion CI ran red on two lanes that the local gate could not see. Lint failed with SC2034 on two reads in cmd_owner_watchdog that destructure the state-root identity into five fields while using only the device and inode. Local changed-file mode suppresses the cross-file codes that need --external-sources, so the warning cleared the pre-push lint step and failed CI's full analysis, exactly as bin/fm-lint.sh's header describes. The unused fields now read into `_`. The behavior shard failed on this suite's own post-TERM assertion, which required the stubbed identity source to be consulted more than once. Whether that happens is platform-dependent: where the runner leader keeps waiting on its TERM-ignoring source child, the post-TERM check sees a live leader whose identity no longer matches, and where the leader dies promptly it sees a leaderless group carrying the same numeric id. fm_procevent_pid_state reaches that second verdict without consulting process identity at all, so the identity source is never read twice and the count assertion fails through no fault of the behavior. The case now asserts the invariant both forms share: retirement refuses, and the ambiguous group is not signalled. Scoping a mutation to this fixture and making the refusal signal instead confirms the case still fails, so dropping the count does not leave it passing vacuously. * no-mistakes(review): Prevent superseded runners recreating stale pacing stamps * no-mistakes(document): Document pacing and ambiguity boundaries * fix(procevent): retire under the recorded identity source and state the home-scoped lease The reused-group case started its runner with the proc-root override in place, so the runner recorded a ps-derived identity, then retired it without that override. Where /proc exists the retirement read identity from a different source than the one recorded, the guard correctly refused an identity it could not confirm, and the case failed on Linux while passing on macOS. It now retires under the same source, and clearing the stub marker first turns that cleanup into the complementary assertion: once the ambiguity is gone, retirement reaps the whole group instead of leaving it behind. The lease prose claimed a runner is bound to the session that owns it, while the mechanism binds it to the home. That gap is what makes a replacement session or an inspection command look like a defect: any activity in the same home refreshes the lease. The granularity is deliberate, because a persistent source is meant to outlive the session that armed it, and binding a runner to that session would stop the sources this mechanism exists to keep running. A runner whose source is no longer wanted in a live home is stopped by reconcile when that source is retired, independently of the lease, so the lease is the backstop for a home that is gone - the torn-down sandbox this change bounds - and the residual is recorded as a known limit. * no-mistakes(review): Rate-limit polls and skip superseded runner launches * fix(procevent): build the claim-only sweep case as a runnerless owned claim A superseded generation now observes the registration-identity mismatch, self-retires, and releases its claim, which is the behavior we want: it clears its own residue rather than leaving a claim with no runner for the home sweep to find. The claim-only sweep case was built by deleting a registration out from under a live runner, which used to leave that runner in place. It now makes the runner retire itself, so the sweep raced that exit and retired one source or two depending on which won. The case failed three runs in four, alternating between a preflight-count failure and `attempted=1`. It now builds the state it means to test: kill the runner's group so it cannot run its own cleanup, assert the owned claim survived that kill, and only then drop the registration. Coverage is unchanged - a runnerless owned claim must still be swept - and the result no longer depends on whether the runner had exited yet. Three consecutive runs pass. The superseded exit also skipped the runner-marker cleanup the normal path performs. The marker is written before the launch floor is waited on, and a home sweep counts a marker with no owned claim as a preflight failure, so exiting without clearing it would make that home refuse to sweep. `FM_LAVISH_POLL_RETRY_DELAY= ` trips SC1007 under the full analysis CI runs, though not under the changed-file mode the pre-push gate uses. * test(procevent): retire a quiet reparented listener instead of racing a storm Explicit retirement was exercised against the spawn-churning stub, which made it nondeterministic. Retirement refuses rather than signalling when it cannot confirm the runner's identity, that identity is read through `ps`, and the stub's 0.1s spawn loop starves that read often enough that a single attempt is a race - the suite failed on this case roughly one run in four, reporting `cannot confirm runner identity; source remains registered`. The refusal is correct: it is the documented preserve-for-retry contract, and a separate case already asserts it. So this is a fixture problem, not a behavior problem. The storm is still covered where the evidence for it lives. The owner-loss home keeps the churning stub and still asserts its tick log stops, which is what proves the churn ended rather than one pid going away. The retirement home never asserted ticks; it only ever read the descendant pid, so the spawn loop bought this case nothing while costing it determinism. It now uses a quiet stub that still reparents and still holds a real descendant in its process group, so the assertions are unchanged: retiring the source must reap the reparented listener's whole group and the descendant under it. Four consecutive runs pass. * no-mistakes(review): Serialize registration replacement through source child launch * chore(no-mistakes): require honest test-step scenario marking The test step recorded scenarios as passing that were only reached through a stubbed dependency or the executable suite, and its validator refused them, because `pass` asserts a scenario was verified against the real live product. That refusal is correct, so the fix is to mark honestly rather than to weaken the gate: a scenario driven live stays a pass and cites its live transcript, while one reached only through a stub or the suite is recorded as untested with the reason and a pointer to its executable coverage. Untested scenarios are reported rather than treated as failures, so real coverage stays visible without claiming verification that did not happen. The instruction also forbids dropping a scenario to avoid marking it untested, since that would hide the gap instead of stating it. * no-mistakes(review): Remove unrelated test scenario policy * no-mistakes(document): Clarify process-event home lease documentation * no-mistakes(document): Correct owner guard failure wording
…nguid#4037) * test(lib): set fixture mtimes through one portable epoch helper On macOS the visible symptom was ONE red case in the turn-end guard suite. The actual damage was TWO cases that had quietly stopped testing their subject. The red one was the harmless half - people read one red case as one broken thing, and here that intuition is wrong. `touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves the file at its current mtime. So on macOS the three away-mode beacon cases never aged their beacon at all. The 400s case exists to pin that 400s is stale under the flat 300s default but fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was worth on this platform: pre-fix input (beacon left at now): ok - passes with the feature DELETED post-fix input (beacon 400s old): not ok - expected exit 0, got 2 It was green while measuring nothing, and could not have caught a regression in the grace it names. Only the 700s case broke loudly. `touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the only host-specific step left is formatting the epoch into that stamp, which date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that probe once and fails loudly rather than leaving an unset timestamp behind - the failure mode that caused this. Verified on BSD touch/date here and on GNU coreutils 9.7 in a container. Three real sites, and one consistency change - not four fixes. The stale destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect: its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and BSD touch accepts that space-separated form anyway. `touch -t` takes the date directly on both platforms, so the branch goes rather than standing as a second copy of the same platform assumption. Known limit: the third case (away mode off) is only HALF recovered here. It now receives the input its name claims, but it is still insensitive after this fix - its verdict is identical with a 0s and a 400s beacon, because the fixture records a daemon lock and no watcher lock, and with away mode off the daemon lock proves nothing. Not fixed here; tracked separately, with the requirement that any fix be shown to FAIL when the protection is removed. FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The turn-end guard and remote-handoff suites are clean. Attribution was established by running the four failures at the base commit and at this head on an idle machine, because base-idle against head-under-load moves two variables at once: script head/loaded base/idle head/idle verdict fm-calm-pi-extension red red red pre-existing fm-backlog-atomicity red red red pre-existing fm-procevent red red red pre-existing fm-startup-network red green green cause unestablished, load-sensitive under a full run Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS because Chrome is absent instead of declaring the capability it needs and standing aside, so its verdict is about the machine rather than its subject - the same family as the defect above, with the red at least announcing itself. Neighbouring class, reported not changed: `file_mode()` - a verbatim `uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least five test scripts plus a `reread_mode` variant, and epoch-mtime reads are open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as the defect above. `git init` without `-b main` depends on the host's init.defaultBranch in several scripts (branch-name case, tracked elsewhere). timeout, sha256sum and sed -i uses are all correctly guarded where checked. Observed while building the check rather than the fix: the first watcher I wrote to wait for the suite matched its own command line, so it was waiting on its own existence and could never fire. Same shape as the cases above - machinery answering confidently about something other than its subject, by including itself in the evidence it was meant to judge. The file sentinel it was replaced with cannot be produced by the observer that reads it. * fix(review): Pin fixture timestamps to UTC across DST transitions * fix(document): Clarify shared fixture suite coverage
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
…guid#4041) * fix(herdr): step around a stale client the running server refuses A remote host can carry a self-updated herdr in ~/.local/bin beside a package-managed one, and the fixed remote-job PATH resolves ~/.local/bin first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2 client (protocol 20) was answered with protocol_mismatch on every command, which the read classifiers folded into `unreadable`: the live remote secondmate read unknown, every doorbell into it failed, and both the spawn and relaunch recovery paths refused, so the defect trapped itself. The adapter's session-scoped CLI wrapper now recognizes that refusal, reads status per session from each distinct herdr on PATH, adopts the first one the running server reports compatible, retries once, and keeps it for the process. The happy path makes no extra call and no other failure reselects. An endpoint that still reads unreadable names the refused client, both protocols, and the fix on stderr; the remote state read, fm-crew-state, and the launch refusal carry that reason, and fm-remote-doctor reports the selected client and rebinds the launch agent to it. Regression coverage: fake two-client hosts in the herdr unit suite, the doctor suite, the crew-state remote arm, and the real host-local control script in the remote lifecycle e2e; the real-herdr smoke refreshes the status shape the selection reads. * no-mistakes(review): Reselect Herdr client after every protocol mismatch * no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding * no-mistakes(review): Scope cached Herdr clients to their selected session * no-mistakes(review): Restrict herdr client selection to reactive CLI calls * no-mistakes(document): Document session-scoped Herdr client reselection * no-mistakes(document): Clarify Herdr client selection documentation * no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat(afk): record the away posture and its lifecycle (phase 1) Away mode becomes a posture of the one supervision session, recorded in state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the record schema, the mandate-clause grammar and compiler, refusal naming the missing part, the read-back rendering, the entry announcement (hold-for-return only, no phone channel), and the archive at return. This release records clauses and does not execute them; the announcement and return brief say so. bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives the record last on stop. bin/fm-afk-return.sh snapshots supervisor health before shutdown, renders the return brief (health, mandate, waiting on the captain, could not fix, handled, cost) from the archived record, the outcome store, the held set, and the status logs, and shrinks the blocker gate to what the away session could not fix. While the record exists the watcher and the daemon never recheck an item held for the captain. Declared external waits get a four-hour default cadence and honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by FM_PAUSE_UNTIL_MAX_SECS. The /afk skill, AGENTS.md's layout and away-mode stub, the session-start digest, and the architecture, Pi branch, configuration, and scripts docs describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a real Pi primary; its verification record carries the 2026-09-08 run. * no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating * no-mistakes(review): Harden AFK authority and posture lifecycle * no-mistakes(review): Preserve AFK history and tighten authority grammar * refactor(afk): record clause fields with no natural-language parser By the captain's mandate the away-posture record keeps no static parser that tries to understand natural language. A mandate clause is now given as explicit fields (--action, --object, --when, optional --stop) that bin/fm-afk-contract.sh records verbatim. The structural check asserts only that the action, object, and precondition fields are present and that the action is a listed verb; whether a precondition holds is the supervision session's judgment at execution time in a later phase. The never-set stays as a forbidden-concept safety scan: fields mentioning credentials, passwords, logins, legal or financial acceptance, payments, invoices, one-time codes, or an attended prompt are refused, matched at token prefixes after punctuation normalization so compound and plural spellings are caught. The red-check grammar, class-word rejection, unconditional-word detection, clause-reference resolution, and condition aliases are removed. --words-file keeps the captain's words verbatim, trailing newline included. The skill, docs, launcher help, and tests describe the field form. * no-mistakes(review): Preserve AFK words and tighten safety refusals * no-mistakes(review): Preserve clause bytes and honor declared waits * no-mistakes(review): Harden deny-list and gate unreadable outcomes * no-mistakes(review): Demote never-set scan and clarify authority * no-mistakes(review): Gate return on unreadable held and status data * no-mistakes(review): Validate posture archives and enforce Pi detection * fix(afk): make the never-set a non-refusing flag and keep return fail-safe Per the captain's decision the never-set scan is a coarse best-effort flag, never a refusal and never the gate: a clause naming a listed concept is still recorded with a flag the read-back, announcement, and return brief show, and the scan matches listed terms exactly or with a plain inflection at punctuation-delimited token boundaries, so unrelated names such as ping-service or tokenize-worker are never flagged and joined compounds remain a documented miss. Authoritative never-set and forbidden-action enforcement is the supervision session's judgment at execution time in phase 4. A replacement copies the superseded record through a temporary name and renames it atomically so a failed copy leaves no partial archive, the record owner gains validate and flags subcommands, and the return keeps catch-up gated when a superseded archive cannot be read. * no-mistakes(review): Harden AFK record validation and return reconciliation * no-mistakes(review): Harden AFK record validation and simplify commands * no-mistakes(review): Harden mandate validation and retain missing records * no-mistakes(review): Refuse blank explicit mandate stops * no-mistakes(review): Recover restored posture epoch before return * no-mistakes(review): Prevent return brief status symlink reads * no-mistakes(document): Refresh AFK posture documentation * no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks * no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
) Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off. Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix(remote): start the fm-remote Herdr agent through a login shell Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot. * fix(remote): start fm-remote Herdr via the account login shell Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render. * no-mistakes(review): Fix launch-agent shell fallback resolution * no-mistakes(review): Preserve and escape Directory Services shell paths * no-mistakes(document): Document login-shell LaunchAgent behavior * no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check * no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks * no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass * no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
…enguid#3710) * fix(bearings): keep captain-approved deliveries in Recently Landed A closed task is never held: tasks-axi clears the held flag when a task closes and keeps hold-kind and the hold reason as the record of the call that was made. Recently Landed excluded every Done row whose hold-kind was captain, so the marker it treated as "closed while still waiting on the captain" was in fact the proof that the captain had approved the work. Every merge routed through a captain decision disappeared from the list of what shipped, including under --all-landed. The selector now asks whether the closed row delivered something. Recently Landed is merged PRs, completed scouts, and finished local-only merges, so a row carrying one of those artifacts belongs there whoever approved it. A captain question closes with an answer and no artifact of its own, and that is what still stays out, so an answered question is never rendered as shipped work. The same rule was written twice - the bearings projection selects this home's Done rows and the fleet snapshot selects each secondmate home's Done rows into the roll-up the same section merges in - which is why one defect hid deliveries in every home. Both now share bin/fm-landed-lib.sh. * fix(review): Normalize landed evidence and exclude answered captain questions * fix(review): Normalize captain delivery evidence across relocated data * fix(review): Record authoritative delivery provenance with legacy fallback * fix(review): Harden delivery provenance across forced and pruned completions * fix(review): Replace premature merge closure with existing release contract * fix(review): Document provenance-based Recently Landed selection * fix(document): Align documentation with completion provenance * fix(lint): Fix targeted ShellCheck warnings * fix(ci): order the pinned tasks-axi install before its stock-Bash consumers In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both stock-Bash consumers, and the Bearings expectation is updated from 49 to 50 tests. Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup 1/1. Full repository lint and all three workflow validations pass, and `git diff --check` is clean. * fix(review): Make completion provenance unambiguous * fix(review): Make completion verdict authoritative over quoted provenance * fix(review): Preserve retained artifacts through resumed captain closes * fix(review): Unify completion provenance ordering across writer and reader * fix(review): Preserve artifacts across failed captain closes * fix(review): Refresh v1 assertions; provenance authority remains unresolved * fix(review): Remove unreliable provenance while preserving landed deliveries * fix(review): Reject stale home summaries visibly * fix(review): Restore retained deliverable recording * fix(review): Match landed artifacts and restore retention documentation * fix(review): Disambiguate captain calls and restore landed artifact matching * fix(review): Persist retained report and PR artifacts * fix(review): Preserve staged artifacts before captain answers * fix(review): Avoid wedging answers on unsupported report paths * fix(review): Exclude unreleased captain-held pull requests * fix(review): Exclude held local-only answers from landed * fix(review): Preserve retained scout reports across snapshot rendering * fix(document): Align landed lifecycle documentation with release semantics * fix(review): Enforce landed artifact-kind ownership * fix(review): Infer canonical task kinds in snapshots * fix(review): Require captain-hold release before merges * fix(review): Qualify merge lifecycle regression evidence * fix(review): Serialize captain holds with merge operations * fix(review): Document merge cleanup residuals honestly * fix(test): Replace vacuous Bearings regression with behavioral cases * fix(document): Align Bearings verification and merge lifecycle documentation * fix(review): Serialize merges and exclude captain calls from landed * fix(review): Harden merge identity and landed selection * fix(document): Clarify landed selector compatibility filtering * fix(bin): keep merge entrypoints usable on records without an incarnation The merge identity guard refused any task record with no spawn_gen field. That field identifies one exact incarnation, so comparing it across the wait for the merge lock is what catches a task relaunched while the merge was queued. Requiring it to be present is a different rule, and it refused every record written before the field existed: a legacy task could no longer be merged at all, and five behaviour suites refused before reaching the check they were written to exercise. The comparison only needs to notice a change. An absent field is now read as an empty incarnation and compared like any other value, so a record that gains, loses, or alters one is still refused, while a record that simply predates the field merges. An ambiguous or unreadable field stays an error, because a record that cannot name one incarnation cannot be compared. The missing-record message each entrypoint had before the guard is restored, so a genuinely absent record still says so in its own words. The role partition now precedes reading the record. Refusing the supervision branch is a statement about the actor, not about the task, so it cannot depend on a record the wrong actor may not have. A backlog file that does not exist meant "no longer an open captain call". For a caller that asked to tell absence apart it now means absent, so a board card whose home carries no backlog stays visible instead of being dropped as resolved. Fixture repositories pin their initial branch instead of inheriting init.defaultBranch, which resolved to main on a developer machine and master on a runner, so a fixture naming main failed only in CI. * fix(review): read local-only note from body; surface pending-close failures * fix(review): keep kindless local-only landings in Recently Landed * fix(review): bind local-only note scan to the tasks-axi note line * fix(review): Guard unavailable captain-hold authority records * fix(document): Document unreadable authority predicate outcome * fix(bin): read an absent backlog as absence, not an unreadable record The merge gate refused every task whose home carries no backlog file. A backlog that does not exist holds no captain call, so nothing can be held and the merge is safe; only a backlog that exists and cannot be read may hide a live hold. Those two states were collapsed into one refusal, which stopped merges in any home that keeps no backlog. The predicate now reports a missing backlog file as absence, alongside a row the backlog does not carry. A record that exists but cannot be read still leaves by the existing cannot-tell path, which both merge entrypoints already refuse, so the restrictive direction is unchanged. That leaves no way to reach the separate unavailable-record result, so the result and the two branches that handled it are removed rather than left describing an outcome that can no longer occur. The lifecycle documentation loses the same claim. Regressions cover both directions in each entrypoint: a home with a task record and no backlog merges, and a backlog present but unreadable refuses without reaching the forge. * fix(review): Fail closed unreadable backend configuration * fix(tests): pin the bare origin's initial branch in the remote seed fixture The fixture created its bare origin with no initial branch, so that repository's HEAD followed init.defaultBranch while the source repository pushed the branch fm_git_init_commit pins. On a host that still defaults to master the two disagreed: the bare origin's HEAD named a branch the push never created, cloning it warned that the remote HEAD referred to a nonexistent ref and checked out nothing, and the seed assertion for the cloned README failed. A machine whose default is already main paired the two by accident and hid it, which is why the fixture passed locally and failed on the runner. Pinning the bare origin to the same branch removes the dependency on the ambient default from both sides. Verified under both conditions: with init.defaultBranch set to master, and set to main, the suite passes 26 of 26. * fix(review): Fail closed unreadable user backend configuration * fix(bin): republish the home summary as v1 and record two load-bearing rules The published home-summary schema had moved to v3, which routed every secondmate home still emitting the earlier version to the stale branch: their landed rows, open decisions and holds all came back empty and their state read as unknown until each home was updated. The payload never justified that. Its field set, field order, truncations and the landed array construction are byte-identical to v1, so only which rows the selector places in landed differs, and a v1 consumer reads that the same way. Republishing as v1 removes the rollout regression and, with it, the tolerance machinery that existed only to soften the bump: the stale-schema predicate, its two collection branches, the flag and its provenance branch, the omitted surface that can no longer be reached, and the fixtures and assertions that covered them. Two rules that a scope review proposed removing are kept, each now carrying the reason it exists, because both were measured to be load-bearing: The artifact-kind ownership clause is what keeps an explicit scout that recorded no report out of Recently Landed. Without it such a row has none of the three artifacts, satisfies the compatibility fallback and renders as shipped work with an empty artifact. The kind fallback is needed because tasks-axi omits the kind metadata entirely when a title begins with a canonical keyword. Without it a scout titled "SCOUT ..." reports no kind, its recorded report stops counting as a delivery, and it drops out of the section this selector exists to repair. * fix(bin): move the scout guard note onto the rule and drop two dead pieces The LOAD-BEARING note sat on an unreachable branch. Measured in both directions: removing that branch together with the kind-is-not-scout guards lets an explicit reportless scout into Recently Landed and fails tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone leaves that suite passing at 49 assertions. The guards carry the rule, so the note now sits on them and the unreachable branch is gone. A note pointing a later reader at the wrong line is the hazard this change corrects elsewhere. summary_file_has_schema lost its only caller when the stale-schema machinery was removed, so it goes with it. * fix(review): Fix legacy report artifacts and canonical keyword boundaries * fix(review): Update pinned Bearings test count to 56 * fix(document): Clarify landed summary compatibility documentation * fix(review): Preserve unreadable backend configuration errors * fix(review): Honor backend resolution errors at existing call sites * fix(test): Stabilize remote collector tests under host load * fix(document): Document backend resolution failure contracts * fix(lint): Suppress intentional deferred probe expansion warnings * fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass * fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
…nguid#4090) * fix(remote): let the Aqua launch agent own the fm-remote Herdr session A herdr server keeps the macOS audit session of whatever started it, and only the Aqua login session (gui/<uid>) can read the login keychain without a prompt. Herdr's SSH remote attach starts the fm-remote server as its own child when it finds none, wins the socket at boot because sshd accepts connections before the login session exists, and every claude pane under that server then gets `security` exit 36, falls back to a stale plaintext credentials file, and reports "Login expired". launchd's own job lost the socket on every KeepAlive retry and the doctor still reported the session ready because it only asked whether any server answered. - Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start the server in the foreground when nothing owns the socket, exit 0 when an Aqua-born server does, and otherwise stop the foreign server, wait for the socket, and exec the server at once. - Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or remote-client-bridge ancestry matched on argv[0] and whole arguments). - Render the agent as the login shell exec'ing the guard with KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded job's successful-exit semaphore, and report a session served outside the Aqua login session as fixable so --fix retakes it through launchd; the reload waits for an Aqua-born owner rather than any running server. - Correct the doctor and docs: the launch shell provides environment parity, the launchd domain provides keychain access. - Pin the guard's decision table and the doctor's verdicts against real marker-carrying processes, and record the dated audit-session evidence. * no-mistakes(review): Verify Aqua ownership through launchd domains * no-mistakes(document): Document macOS lsof ownership requirement
…uid#2881) * fix(bin): prefer a live no-mistakes run over a terminal one A worktree can bind to more than one recorded no-mistakes run at once. The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an exact-equal commit and a worktree-is-an-ancestor match, but never stated which wins when both bind, so the tie fell to whichever candidate the caller reached first. Observed on a live fleet: a crashed validation daemon left a FAILED run at the worktree's own commit while the live run that replaced it validated a descendant commit on the same branch. Bare `axi status` answers with the most-recently-touched run - the corpse - and it bound by the equal-commit rule, so every recomputation reported `failed` for a task whose real run was healthy. The same label had also read `failed` earlier while the work was genuinely stalled, so the signal was wrong in both directions. State the live-over-terminal policy in the matching rule's own contract, where the equal-commit and ancestor rules already live, and add fm_nm_run_status_class as the one classifier that decides liveness from a recorded status word. fm-crew-state.sh applies it on both selection paths: the runs listing now scans past a terminal row for a live one, and a terminal `axi status` answer is provisional until the listing has been asked whether this worktree also has a live run. Same-liveness-class candidates keep the listing's newest-first precedence, and a status word the classifier cannot place keeps the caller's own ordering rather than displacing a known result, so a single-run task and a task whose runs are all terminal are unchanged. Regression coverage reproduces the proven case (terminal run at the worktree's exact commit plus a live run descending from it) and its runs-list twin; both fail under the old tie-break. Two companion cases pin the no-widening half - two terminal rows still resolve newest-first, and a terminal run with no live sibling keeps its full run-step detail - and both pass before and after the change. * no-mistakes(review): accept unfetched live sibling anchored at exact worktree head * docs(bin): name both ledger reads behind the runs-limit setting The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described the runs ledger as scanned only by the cross-branch fallback, but the live-over-terminal fix also consults it as the live-sibling probe behind a terminal axi status answer. Point the comment at docs/configuration.md as the setting's owner instead of restating a second copy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…chenguid#4009) * fix(procevent): make the ordinary stop signal actually stop a runner The owner guard that shipped in kunchenguid#3904 half-reaps. Against a poll child that handles the ordinary stop signal and keeps waiting, the guard signals the group, loses the runner leader to its own signal, then reads that success as a leaderless group and exits without escalating. It destroys the only proof of ownership that would have authorised the forced signal, so the survivor becomes unreachable by retire, reconcile, sweep-home and the guard alike. A guard that turns a leaking-but-identifiable generation into a permanently unreachable one is worse than no guard at all. Two defects, and they hid each other: - The escalation re-derived ownership from the leader. `runner_group_signal` now takes a `proved` mode, passed only by the escalation inside the stop that already proved and signalled that exact generation moments earlier. A leader dying to our own signal is the ordinary outcome, not fresh ambiguity. - Every stop held the per-source lock across its wait while the runner's own exit cleanup waited unboundedly for that same lock. That circular wait was broken only by the forced signal, so the forced signal silently became the normal path - and, by keeping the leader alive through the whole window, it masked the escalation defect above. The runner's exit cleanup now refuses that lock instead of waiting for it, which is what its existing `return 0` already said it did. Fixing the lock alone would have turned every stop of a signal-proof child into a refusal that leaves it running, so both land together and the tests pin that. Measured on macOS with a stand-in poll child that traps TERM, INT and HUP: the guard left it running past 70s and now clears the group within the lease plus one check; retiring a healthy runner fell from ~2.8s with a forced group signal every time to ~0.6s on the ordinary signal alone. Unchanged and stated deliberately: a leader lost to anything other than the stop's own signal still leaves a group that retire, reconcile, sweep-home and the guard all refuse, permanently - and that source stops listening without saying so. Whether such a group may ever be signalled is an open decision and is not answered here. * fix(review): Fix proved escalation race and stop regression assertions * fix(review): Preserve proved escalation through transient identity failures * fix(review): Simplify proved escalation and correct guard timing documentation * fix(document): Clarify process-event stop ownership and cleanup limits * fix(document): Clarify process-event stop ownership and fixture comments * revert(skills): restore the leaderless-ambiguity limit to the loaded skill An automatic documentation step in this branch's validation edited .agents/skills/process-event-sources/SKILL.md, which no instruction in this change asked it to touch. That file is not documentation about the code: it is the agent-loaded instruction surface, what an agent reads to know what it is permitted to do. The step deleted this line: - leaderless PID/PGID-reuse ambiguity preserves the claim without signalling or replacement, as owned by the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); and folded it, with its neighbour, into a generic "registration and ownership transitions, stop authority, and claim reclamation follow the operating contract". That deleted line states a PROHIBITION - that such a group is preserved WITHOUT SIGNALLING - and it is the exact limit an open captain decision currently rests on. Folded into a pointer, an agent reading the skill to learn what it may do would have to chase a second document to discover it may not signal. A prohibition that requires a second lookup is not a prohibition. The effect was to weaken, in the instructions themselves, the boundary that keeps one home from signalling another's process group - while the question of whether that boundary should move at all is still open. This is a deliberate revert, not an oversight, and it restores the file exactly to its pre-branch state. The full statement also survives in docs/configuration.md; that does not rescue it, because the agent handling a process-event wake loads the skill and not the documentation. * revert(procevent): restore the open-question marking beside the escalation The same automatic documentation step that edited the loaded skill also removed this from the comment above runner_group_signal: A leaderless group nobody in this call ever proved remains refused too, for every caller. That untouched refusal is what makes a crashed leader's group permanent, and relaxing it is a separate open question, not something this path assumes. and replaced it with a pointer to docs/configuration.md. This one fails differently from the skill deletion, which is why it is restored separately. There, a prohibition was moved out of the reader's path, and a missing prohibition gets violated. Here the prohibition survives in code - the unproved path still refuses - and what was removed is the fact that the limit is UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as settled design, and settled design gets relied on, extended, and eventually relaxed by someone confident they understand why it is there. That question is open right now. The rule this branch's four instances produce, stated once here because this is the point of decision: an unresolved question must be marked unresolved AT THE POINT OF DECISION, not only where the contract is documented. A reader who does not know something is open will treat it as closed, and that default is stronger than any pointer overcomes. The pointer added by that step is kept alongside; this restores what it replaced rather than reverting it. * docs(verification): restore the measured guard bound and its reason The document step's rewrite of this record dropped the concrete figure while keeping the surrounding measurements. What went missing was the bound itself - lease plus two consecutive failed checks plus the stop's grace, roughly 630 seconds at the shipped 600-second lease and 15-second check - together with the reason there are two checks rather than one: a single unreadable read must not be enough to kill a live runner. The mechanism survived elsewhere and the reason survived in docs/configuration.md, so nothing was lost from the repository. The concreteness was, and that is what this restores. A number recorded without why it is that number is the one a later reader shortens; the reason is the whole safety argument for the debounce, and the debounce is what stops the reaper killing a live runner on one bad read. * fix(document): Replace stale stop-authority summaries with owner pointers * test(procevent): make the guard-bound case able to fail for its own reason An automated reviewer observed that this case allowed sixty seconds for a bound of roughly eight, so it could not go red for the reason it names: it would have passed a guard that took fifty-five seconds. That is correct, and it is the same family as the defect the case exists to defend against - a check that is green because it cannot fail, rather than because the thing it guards is working. The deadline is now derived from the bound itself - the lease, plus the two consecutive failed checks the guard debounces on, plus the stop's own ordinary and forced signal windows - rather than from a flat wall-clock number, and the shortened lease and check the fixtures run under have a single definition so a derived deadline cannot silently diverge from the settings the guard is given. The doubling that remains is a load allowance and is documented as one; widening it to make a slow guard pass would convert the assertion back into decoration. Proven by mutation rather than by argument. Against the repaired case: correct code ok guard debounces on 20 misses instead of 2 not ok - "still holding the group after 16s, against a documented bound of 8s" proved escalation removed (the original defect) not ok - same code restored ok The previous sixty-second version passes every one of those mutations. The reviewer's other claim, that the guard can survive past the announced bound when an owner disappears immediately after a check, was measured and does not hold against what this branch announces. Sweeping the phase deliberately at 0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed checks plus stop grace, which is up to 8s at those settings. The mechanism the reviewer describes is real and is the announced mechanism; the bound it was measured against is a phrasing this branch no longer carries. * revert(scope): return the instruction surfaces to their base state This delivery is being split. It carries the two proven process fixes alone; the instruction text travels separately, through a run that removes the documentation step rather than refusing it at its gate. Two surfaces are therefore returned to exactly what the base branch has, so this delivery neither adds to them nor removes from them: .agents/skills/process-event-sources/SKILL.md - identical to base again. Three bullets an automatic documentation step had folded into a pointer, including that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT SIGNALLING and that there is one identity-matched owner per canonical source across homes sharing one store. The header comment block of bin/fm-procevent.sh, which is what the script prints as its own help. Seven lines were removed from it: that a live owner is never displaced, that only a claim whose stale owner and independently absent process group prove its whole generation gone is reclaimed, that a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, and that reconcile signals only a live identity-matched runner group and otherwise keeps the claim without starting a replacement. The help output is now byte-identical to base. Neither removal was requested by any instruction in this change, and both were made to text that predates it. Returning them is scoping, not a third restoration: nothing is being added to those files here. * fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor * test(procevent): make the post-TERM cases report what they saw when they fail On the failure path only, these cases now print what they actually saw: the identity recorded at claim time, the identity readable at that moment, the size of the signals file, the leader's state and wchan, every live member of the runner's process group with its own state and wchan, the elapsed time since the stop began, and what retire said. None of it runs when a case passes. WHY THIS IS KEPT, stated accurately rather than by its original reason. It was written to make an unexplained CI failure verifiable. That failure is now explained - it was a fixture deadlock, diagnosed and repaired in the preceding commit - so that justification has expired and is not the reason given here. The reason it stays is smaller and independent of that failure: it is already written, it is small, it sits in the file whose assertion this change reworked, and an assertion that could not say why it failed cost most of a morning to diagnose from the outside. The next failure will not be this one. WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already observed many times, not proof that anything is fixed. Only a failure carrying the evidence above establishes a cause. * fix(document): Clarify process-event fixture diagnostic rationale * fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor * fix(procevent): bound owner-guard cleanup at one check interval, not two A THIRD WAY, not a capitulation to the reviewer and not a refusal of it. The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the number of observations the guard makes before it acts. It asked for a single read because that was the only route it could see to an acceptable bound. There was another route, and this change takes it: the bound is reached and both reads are kept. TIGHTENED - the SPACING of the guard's two reads, not their number. The owner watchdog now sleeps half the configured check interval and still requires two consecutive failing reads, so the pair completes inside one check interval instead of costing two. Worst-case detection falls from the lease term plus TWO check intervals to the lease term plus ONE. At the shipped 600s lease and 15s interval the stated bound falls from ~635s to ~620s. PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root identity read, and either can fail transiently on a live, healthy home. Acting on the first failure would let one isolated unreadable read kill a live service. Requiring a second, independent read is what makes that impossible, and it is a protection rather than padding. Nothing was traded away to reach the bound. Both properties are now guarded by their own case, and each was proven by MUTATION rather than asserted: - putting a full interval back between the two reads fails the bound case: "still running 17.0s after the last owner activity, against a documented bound of 15s"; - acting on one failed read fails the new debounce case: "one unreadable lease read ended a runner whose home was still alive" - while the bound case then passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one is exactly why these are two cases: one elapsed-time case would have registered the removal of the protection as an improvement. MEASURED, sampling the phase between the guard's check clock and the lease clock across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned listener whose home stopped refreshing its lease: lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after The 4s configuration is the informative one: the gap is about one check interval, which is precisely the term that was removed. A previously unstated term of the bound surfaced while measuring: the lease age is compared in whole seconds, so a configured lease of N is honoured until that age reads N+1. It is now part of the documented bound and of the regression's derivation instead of being absorbed into a fudge factor. The bound regression derives its deadline from the documented bound instead of a flat number, and PINS the phase between the guard's check clock and the lease clock rather than sampling it, because with a sampled phase a guard spending two intervals passes about half the time on a lucky alignment. Its load slack is additive and stays under half a check interval, so an extra whole interval cannot hide inside it. The two flat deadlines that were there before (40s and 20s) and the doubling allowance on the derived one are gone; that looseness was the reviewer's third complaint. The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the forced one. It is a ceiling paid only by a group that outlives the signal it was sent, not a delay every stop pays - a healthy runner's whole retire measures 0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is unreachable by any implementation, since signalling a process and giving it any chance to exit takes non-zero time; detection now meets it and the stop runs inside its own ceiling, and the contract says so rather than glossing it. NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs start rather than after they report. On the previous head, "Behavior portable serial 1" and "Behavior portable serial 4" were both CANCELLED at the job ceiling, independently of this finding. A new head triggers fresh runs, so those two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly and sometimes is not; the shard-packing repair for it is open separately. Do not reread a lucky pass here as a resolution. Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh was NOT updated even though the two new cases add ~19s of wall clock. docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI timing artifacts of green runs, and that repair is the open request doing it; a hand-edited estimate here would collide with it and silently repack the shards. This suite runs in portable serial shard 3, which was green in the last run. Verification: tests/fm-procevent.test.sh green, plus tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was removed" case flaked in 4 of 7 local full runs; an isolated 20-trial reproduction measured it at 13/20 unclean before this change and 11/20 after, so it is issue 4080 and is not aggravated here. * fix(procevent): repair our decimal-interval regression and enforce the timing phase REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added by the previous commit read a zero-prefixed interval as octal: 010 halved to 4 instead of 5, and 08 was not a number at all, so the owner guard died before reporting ready and the runner failed closed and never listened. The validator accepts those values and `[` compares them as decimal, so this broke a configuration that worked before. Introduced by this delivery, found in review, repaired here. Forcing base ten before the arithmetic is the whole runtime fix. Proven by driving it rather than by reading the source: a new case starts a real listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and 5s. Removing the normalisation turns that case red with "a zero-prefixed decimal interval (08) prevented the listener from starting". THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced nothing. Review was right that this is not enough: once startup reaches about two seconds the expiry lands in a different part of the interval and the case silently stops rejecting a two-interval guard while still reporting success. A bound that cannot fail for the reason it names is the defect this whole delivery exists to correct, so it must not ship inside the fix for it. Now the lease is synchronised to the guard's own FIRST observed lease read, every later real read is recorded, and the case REFUSES unless one recorded read proves the required phase: it read the synchronised reference, it was still fresh, and it began late enough that two further full intervals could not finish before the deadline. An unestablished precondition refuses; it does not proceed on trust. The derived deadline, the two-read debounce and the additive slack are unchanged, and the slack invariant is now asserted rather than left to a comment. Review also found the deadline was only ever checked while the group was still alive, so a sampler descheduled past it would see the group gone and certify success. The observed completion time is now checked too. PROVEN BY MUTATION, each one run against this code: - remove the decimal normalisation -> the interval case fails on 08; - a full interval between the two reads -> "the guard exceeded its bound: group still running 17.1s ... against a documented bound of 15s"; - a full interval WITH startup forced to ~2.5s, which is exactly the condition the old construction pin could not survive -> still red, same message; - the same ~2.5s startup with the correct guard -> still passes, 13.0s against the 15s bound, so the delay alone does not break the case; - phase evidence made unavailable -> "could not establish the required pre-expiry guard-read phase", a refusal rather than a pass, even though the group stopped quickly; - act on one failed read -> the debounce case fails and the bound case passes FASTER, 9.5s against 12.7s, which is why these remain separate cases. Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean. * fix(document): Correct process-event timing and debounce comments
…kunchenguid#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…guid#4119) * fix(bin): stop false watcher-down alarms on long Claude turns A healthy Stop auto-arm rewake or open claim already explains a mid-turn beacon that has aged past grace, because turn-end will re-arm. Keep the supervision-off banner for a missing, failed, or exhausted generation. * no-mistakes(review): Bind Claude rewakes to active recovery generation * no-mistakes(document): Document Claude long-turn supervision exception * no-mistakes(lint): Fix empty ShellCheck assignment * no-mistakes(ci): Fixed all reported CI failures: quoted the hyphenated recovery-delivery value to satisfy ShellCheck SC2100, and updated the session-lock auto-arm fixture to emit the recovery marker and watcher beacon now required for a valid rewake. Verified fm-session-lock-ancestry, fm-test-run, stale-banner, Claude auto-arm, targeted lint, ShellCheck, and workflow lint checks pass
* fix(herdr): classify a gone session's endpoint as recoverable A task whose Herdr endpoint could not be read was classified `unreadable`, which blocks recovery by design. The commonest reason that read fails is that the recorded session's server is not running at all - a host reboot, a server exit, a session never restored - and that is authoritative absence for every pane in that session, not an ambiguous answer about one of them. Tasks in that state had no sanctioned way back. The recovery-grade read now settles an uninterpretable pane read with the session server's own `.server.running` state: positively stopped reads `missing`, while a running server, or a server state that cannot itself be read, still reads `unreadable`. Resting the verdict on that field rather than on the `server_not_running` error code is what keeps it working across Herdr 0.8.x and 0.9.0, since the field is present on both and the code is not. Only that one boundary is widened. The husk classifier under it stays strict, so duplicate prevention, rollback, and teardown - the paths that can destroy something - keep refusing on exactly the reads they refused on before. Separately, a relaunch refused outright when the endpoint's shell had drifted out of the recorded worktree. An agent's own exit routinely leaves its shell somewhere else, so that refusal stranded tasks whose work was sitting untouched on disk. The shell is now told once to return, and only a shell that will not go refuses; the replacement still never starts outside the copy holding the work. Herdr 0.8.x is not installed on this host, so protocol-20 coverage is structural plus the adapter fixture exercising both response shapes, and is recorded as such rather than as a live result. Fixes kunchenguid#4091. * no-mistakes(review): Restrict drift recovery to Herdr endpoints * no-mistakes(review): Correct Herdr recovery verification coverage * no-mistakes(document): Document Herdr endpoint recovery boundaries
kunchenguid#4033) * fix(bin): keep an escalated undelivered handoff wake retryable A remote backlog handoff holds its outbox until the backlog receipt and the receiver wake are both confirmed, and retries the wake under the same pending-reply correlation on every resume. When that wake's remote transport was lost, the correlation stayed undelivered in delivery_unknown and the watcher's next pending-reply tick escalated it. Both the reuse predicate and the known-undelivered reset refused an escalated record, so the resume refused to resend the wake forever and every later handoff to that mate jammed behind the outbox. Treat an escalated record with no confirmed delivery as the undelivered correlation it is: fm_pending_reply_corr_reusable accepts it for its own task and fm_pending_reply_reset_known_undelivered returns it to awaiting_report for the idempotent remote resend, while a delivered record is still never reset and a missed-report escalation keeps its meaning. The published delivery-unknown decision stays open until the record resolves, so a repeat loss neither re-notifies nor strands it. Reproduce the deadlock end to end in the remote handoff test (lost wake transport, watcher escalation, resume) and pin the predicate contract in the pending-reply suite; the fm-send fixture that pinned the refusal now uses a genuinely stale delivered escalation. * no-mistakes(review): Decouple durable outboxes from best-effort wake retries * no-mistakes(review): Align handoff documentation with durable receipt release policy * no-mistakes(review): Handle unrecordable wake state as dropped * no-mistakes(review): Prevent stale wake markers blocking handoffs * no-mistakes(review): Prevent stale delivered markers suppressing new wakes * no-mistakes(document): Clarify retry escalation decision lifecycle * no-mistakes(document): Document pending receiver wake retries
* docs: bound the mandatory captain address to the chat channel AGENTS.md's opening address rule said "address the user as captain at least once in every response" and never said what a response is. The artefact exclusion two lines below governed only the optional nautical seasoning, not the mandatory address. An agent that reads this file without being the first mate - a pipeline corrector agent running inside a copy of this repo - therefore read the obligation as applying everywhere and the exclusion as applying only to flavour, and opened its delivery message with "Captain,". That reading was correct. Patch the existing owner rather than adding a rule elsewhere: - bound the obligation to chat messages sent to the captain; - state the artefact exclusion once, explicitly binding every agent that reads this file whether or not it is the first mate, and naming commit messages, PR and issue descriptions, briefs, code and comments; - fold the seasoning under the same bound instead of carrying a second, narrower copy of the exclusion. The obligation itself is unchanged: the captain is still addressed in every chat message. AGENTS.md goes from 603 to 602 lines: the redundant "never send a response with zero direct address" clause and the duplicated seasoning exclusion pay for the new bound. The two cross-references that paraphrased the unbounded wording (bin/fm-parent-channel-lib.sh's header and docs/secondmate-parent-channel.md's problem statement) now match the owner; neither restates the rule. * fix(review): Limit address exclusions to artifacts while preserving public replies * fix(document): Consolidate captain address guidance
…nguid#4131) * fix(herdr): close persisted-focused tabs when no live client is attached The teardown active-tab guard treated Herdr's last-focused pointer as a live viewer, so detached sessions could not close panes on that tab. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): allow detached seeded-tab prune after live-client gate Projection create still restored the persisted focused tab after a successful prune, so a detached last-focused seeded tab still quarantined the spawn. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): probe live client after seeded prune only when that tab was focused The extra title-clear read after every prune shifted canned CLI fixtures and failed projection create. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Tighten Herdr active-tab close guard * no-mistakes(review): Guard Herdr mutations with fresh target focus * no-mistakes(document): Document Herdr live-viewer teardown guard --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…guid#4169) * fix(bin): escalate decision-owned wakes once as the decision The away-mode daemon treated a needs-decision: queued payload as an unknown wake, so suppression markers never committed and the same open decision re-escalated on every poll. Classify that payload through the existing signal path so it escalates once, labelled as the decision, and an unchanged repeat is suppressed on the same terms as any other signal. Fixes kunchenguid#4096 * no-mistakes(review): Escalate captain-held decision-owned rows once as the decision * no-mistakes(review): Self-handle captain-held decision-owned rows instead of escalating them * no-mistakes(document): Name away daemon as needs-decision payload reader
* test(herdr): pin leftover-shell vs live-idle via agent get Herdr 0.9.0 already distinguishes a Pi that exits to a surviving pane shell from a sibling live idle occupant. Pin that pair through agent get and the recovery classifier so a lagged pane-get status cannot silently reclaim the leftover shell as alive. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Document Herdr leftover-shell liveness regression * no-mistakes(ci): Fixed Lint failure SC2034 by replacing the unused wait-loop variable with `_`. Verified with the pinned project lint command, Bash syntax check, and git diff check --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…nchenguid#4151) * ci: rebalance the portable parallel lanes on measured runner durations Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at that cap on every request raised on 2026-09-10 while lane 2 finished in about 3.5 minutes, so no request could go green. CONTRACT CLASS: RESTORE. The workflow already promises two duration-balanced lanes and the shard documentation already claims a measured wall; this re-establishes both against what the lanes now cost, and changes no lane count, no cap, and no scope of what runs. The counter-argument, so nobody has to take that on trust: two pieces here are genuinely new rather than restored, and either could be argued to make this a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured durations where it previously handed every parallel script the serial default weight and returned an alphabetical order; and `--check-coverage` gains three reported fields. I classify the change RESTORE because both exist only to make the already-promised property checkable, but they are named here rather than folded into the restoration. === PART 1: THE TOTAL, AND HOW IT WAS OBTAINED === This section stands on its own. It establishes what the parallel set costs. It derives no packing; Part 2 does that, from this number. THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts. Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split. HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its duration did not exist as a recorded figure anywhere and no timing artifact was expected for it. It turned out to be recoverable from the real lane without estimating, by two routes, across six CI runs on 2026-09-10 (34459949083, 34460760299, 34462530836, 34462758357, 34466966385, 34470382458): - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all 11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The upload step is if: always(), so the cancellation did not suppress it. This is one full, untruncated lane-1 measurement. - The five other lane-1 jobs were cancelled mid-suite, but each logs every script that had already finished as an FM_TEST_END duration_ms= marker. Those per-script records are complete measurements of completed scripts; only the script in flight at cancellation is lost, and it differs by run. Lane 2 completed in all six runs, so its scripts come from the six uploaded fm-test-timing-portable-parallel-2 artifacts. Every one of the 24 scripts therefore carries at least one untruncated measurement: 20 of them measured in all six runs, two in three or four runs, and two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run. Each hint is the SLOWEST value that script reached, so the total is an upper envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED LANE, and no lower bound was ever extrapolated into a total. THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the real portable parallel lane on a GitHub ubuntu-latest runner, produced by the lane's own CI job. It transfers because it is not a proxy for the lane; it is the lane. Nothing in the total came from this machine or from any harness of mine. That mattered, and here is what it would have cost. A same-day macOS cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against 31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore do not scale the lane, they REORDER it, so a packing derived from them would have balanced the wrong thing while looking clean. WHAT IT REPLACES, which is the root cause. The lanes were packed from the 2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers. That record answers whether the candidates are isolation-safe, not how long a SERIAL CI lane runs, so it was structurally incapable of representing lane wall clock even when it was fresh. It was also never refreshed while the set grew about 3.2x. Both the wrong instrument and the staleness are fixed here: the hints now come from the lane itself and carry their run ids and date. === PART 2: THE SPLIT DERIVED FROM THAT TOTAL === Longest-processing-time assignment over those hints gives 414269 ms and 414299 ms, 30 ms apart, against 624299/204269 before. tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which installs the Pi package, so ci.yml needs no step changes. === PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE === Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before any variance is applied, and the cap covers the whole job rather than the suite. worst lane, script time 414299 ms 6.90 min job overhead, measured on the real lane ~18 s (see below) expected healthy job ~432300 ms 7.21 min x1.29 on the script time, plus overhead ~552400 ms 9.21 min cap 600000 ms 10.00 min room left after the multiplication ~47.6 s 7.9% of cap The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as supplied; it is not this lane's own figure. This lane family does have its own, and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms, a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about 2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather than the tighter one this lane actually shows, and note that the hints are already per-script maxima, so 1.29x on top is conservative twice over. THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping. The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a 192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install. The cap is unchanged, the lane count is unchanged, and nothing in the serial lane, its shard count, its guard or its hint table is touched. === PART 4: THE RECORDED FACT === The workflow comment no longer restates the shard wall as a literal, which is how "~1 min of serial sum" survived a 10x change without announcing it. It now points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms, parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the current number is computed on demand. The shard documentation carries the dated run ids, which route it was taken by, and the local cross-check that shows why local numbers are not admissible as hints. Two regressions pin what rotted: lane membership must be stored longest-measured-first, and the lanes must be fully hinted and packed within 5% of each other. Both were run against the old composition and both fail on it (420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they replace named a specific script by hand and had itself gone stale. === PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE === tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole set, so it is the floor of any two-lane split: no repacking can put a lane below it. After this rebalance the cap is about 1.45x the healthy lane where the sibling serial lane keeps roughly 2x. Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT bounds the serial lane. parallel_unhinted is reported, not enforced, which is what let this drift for three weeks unnoticed. * fix(review): Restrict parallel scheduling hints to portable parallel lanes * fix(document): Clarify parallel lane scheduling and timing evidence
… PR (kunchenguid#4148) pr_for_task fell back to scraping the whole status log with tail -1, so any PR URL a worker ever mentioned in prose - including a scout citing someone else's PR - became the task's delivered PR in the parent-channel terminal report. Recorded meta pr= is now the only authoritative source, the fallback scrape accepts only a preferred terminal line in a mode's ready-signal shape (done: PR <url> or done: PR <url> checks green), and a scout never carries pr= at all.
…claims instead of counting a dead drop as started (kunchenguid#4212) * fix(procevent): stop a dead runner owning a source and reconcile reporting it The captain answered ten calls on a bearings board, the board accepted them, and nothing collected them. He had to answer all ten again in chat. A surface that presents as armed while being a dead drop is worse than one that visibly fails, because the answers looked recorded. Two independent defects, reproduced together in an isolated home where reconcile reports started=1 on every run while ownership never moves and no runner ever attaches. 1. reconcile counted a launch it never verified. detach_runner is fire-and-forget and discards the child's stderr, so a runner that died before it could claim was counted exactly like one that is listening. Launches are now confirmed - the source observed owned, or its runner record moved - before being reported as started; the rest are reported as failed= with a non-zero exit. The runner-record clause is what keeps a fast-completing source from being reported as a failure when it finished between two polls. One bounded window covers a whole cycle's launches, so a home full of broken sources costs the same wait as one. 2. A claim whose whole generation is provably gone could be refused forever. Reclaiming it ran cleanups over that dead generation's own leftovers, and any failure vetoed the claim - permanently, because none of those conditions clears on its own. Every one of those leftovers is keyed by the dead generation's claim token and a replacement always claims a fresh one, so none can collide with what replaces it. fm_procevent_claim_capture_reservation_reclaim_locked already said this for the reservation record; the staging file and the shape check on the registry directory recorded to hold it now take the same rule. Removing the claim record itself stays a hard precondition: two owners is the one outcome worse than none. Two smaller repairs to the same "registered is not listening" confusion: - `list` reported OWNER=none for a source nothing can claim. A reused PID whose process group survives reaches that state through the stale branch rather than the leaderless one, so it read as an idle source waiting to be started - the reassuring answer this surface gave while a board collected nothing. It now reports the orphaned state it shares. - reconcile relaunched into that same unclaimable state on every cycle, spawning a runner that could only die on the claim. docs/configuration.md already promised it preserves such a claim without starting a replacement; the code now does that and reports it as uncertain. This is NOT a third instance of today's two lock-identity defects (4e1bf9aa and its replayed predecessor). Those were wrong liveness predicates: a reused PID read as a live holder, then an exec'd holder read as dead. Here the predicate is right - the code correctly proves the owner dead and refuses the claim anyway, on a condition unrelated to liveness. Regression coverage, each failing on the parent commit for its own reason: - tests/fm-procevent.test.sh: a source that cannot start is reported as failed rather than started; a dead generation whose leftovers cannot be tidied no longer keeps owning its source (the parent reports a start while nothing ever runs); the existing reused-PID fixture now also asserts the orphaned listing and that no doomed relaunch is reported. - tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the keyed-answer intake through the runner end to end - durable capture, the wake, and the closed task carrying the captain's selection. This one passes on the parent, because that chain was never what broke. fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50, fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has two failures identical on the parent commit (EACCES on package install in this sandbox) and unrelated to this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG * no-mistakes(review): confirm reconcile launches on durable launch stamps * no-mistakes(review): announce stranded sources and refuse bad confirm windows * no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs * no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim * no-mistakes(review): nonce launch-failed keys, refuse bad window at arm * no-mistakes(review): state only observed launch outcome, shorten episode nonce * no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake * no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record * no-mistakes(lint): restructure SC2015 chain into explicit if-block * test(watch-triage): fix two timing-exposed defects the pipeline found Both surfaced in the no-mistakes test step on this branch, each failing one full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to retry past. 1. The new launch-failed delivery test assumed an already-surfaced key never wakes the watcher again. That is false: a fresh watcher legitimately re-surfaces any unacknowledged queue row through its downtime-recovery path ("check: rearm-resurface"), so the assertion failed whenever a re-arm landed between its two checks. The pipeline's own fix tolerated any wake lacking the repeated key's headline; this tightens it to exactly one tolerated reason, by its exact line, with a failure message that names the expectation so a reworded path reads as "the tolerated recovery path changed" rather than as a mystery - and so nobody restores the strict silence check. The positive assertion (a fresh-suffix key is delivered under its own headline) is unchanged. 2. seed_captured_procevent_result retired its source in the gap between the runner publishing its wake and releasing its claim, so retire read the exiting runner's ownership as uncertain and refused ("cannot confirm runner identity"). The fixture and retire path pre-date this branch; the confirm window returns reconcile closer to the moment of capture, which made the gap easier to hit. The fixture now waits, bounded, for the claim release the publish promises, with the reason at the wait. Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50, fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13, fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh exit 0. First attempt, no retries. * no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gistration (kunchenguid#4191) * fix(herdr): verify agent registrations at process level before trusting them Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle) after the Pi process has exited to a plain shell whenever a nested interactive shell sits under the pane's top shell, which is the crew shape `treehouse get` leaves behind. The pane classifier trusted that registration alone, so `fm-control.sh <id> relaunch`, `fm-spawn.sh --relaunch`, and the crew-state recovery read all treated a shell-only pane as a live agent and refused recovery for as long as the record lived. The Herdr adapter now reads `pane process-info` plus the real process table through a shared harness-process classifier (bin/fm-agent-process-lib.sh, moved verbatim out of the tmux adapter so both backends mean the same thing by agent, shell, and other) before a registered agent counts as live. A registration over a shell-only pane is the new explicit `stale-agent` pane state, which the recovery-grade read maps to `dead`; husk detection, reclaim, presentation recovery, and session cleanup keep refusing it, so recovery reuses the pane and nothing gains close authority. A working record is verified the same way before the native busy verdict reports busy, so the recovery classifier never reports a shell-only pane as working. An unreadable process view reads unknown, trusting neither the registration nor its absence. Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves exit and relaunch through the control plane, and the portable suites pin the classifier over real processes. Fixes kunchenguid#4115. Duplicates: kunchenguid#3639, kunchenguid#3487, kunchenguid#2908, kunchenguid#3545. * no-mistakes(review): settle transient prompt helpers before trusting herdr process state * no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round * no-mistakes(review): untrack stray .codegraph/.gitignore * no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor * no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted. * no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable * no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory * no-mistakes(fix): update remote herdr fixture to the real pane process-info shape The shared remote-secondmate herdr fixture still returned the old flat process-info body ({"result":{"process":{"name":...}}}). The process-level liveness classifier added for kunchenguid#4115 requires the real {"result":{"type":"pane_process_info","process_info":{...foreground_processes}}} shape and treated the old body as unreadable, so an already-launched remote endpoint's agent-state read failed and any relaunch attempt against it died with "remote endpoint state is unreadable; refusing duplicate launch" instead of reaching the state it was actually exercising (tests/fm-remote-secondmate-parent-binding.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh). * no-mistakes(review): test: add empty-foreground regression test for herdr flake fix * no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points
* Bind GitHub merges to a live green head and require an away-task grant. A GitHub merge now re-reads the pull request and passes --match-head-commit, so a red or moved head cannot land the way GitLab already refused. While an away record exists, only yolo or a named grant may merge, so hold-for-return cannot ship an ungated PR. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Harden away merge authorization and grant parsing * no-mistakes(review): Restrict fallback outcomes to proved GitHub merges * no-mistakes(document): Refresh merge safety documentation * no-mistakes(ci): Fixed all three CI failures by updating legacy GitHub merge fixtures for live-head verification/direct gh merges and removing a process-event runner cleanup race. Verified fm-pr-check-security, fm-captain-hold-lifecycle, and fm-watch-triage pass locally; shell syntax and git diff checks also pass * no-mistakes(document): Document attended red-check exception --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…m epoch (kunchenguid#4221) The --claude guard's re-block budget charged the auto-arm ledger epoch, not the re-block: `budget_account_current_epoch` advanced the session count only when `state/.claude-autoarm-epoch` named a different generation than the previous accounting. The epoch advances only inside the auto-arm hook's generation claim, so a hook kept inert before that claim - a session lock held by a live harness outside its ancestry, a hook that never fires, or an identity or write failure ahead of `fm_autoarm_claim_next` - left the ledger frozen at its last outcome and the count frozen with it. Reproduced in a fixture: twelve consecutive Stops re-blocked with the count at 0 and the attended fail-open never fired, leaving only Claude's silent 8-block override, the blind end the bounded alarm exists to prevent. The budget now charges a re-block against an epoch the previous re-block already charged, while still charging each epoch at most once per Stop so the wait loop's repeated observations of one fresh terminal outcome and the same invocation's block decision cannot double count. The advancing-epoch progression is unchanged: three re-blocks, then one attended fail-open for a verified failure episode, and a frozen epoch now follows the same shape. Budget exhaustion without a verified failure still blocks, by the existing contract, and positive watcher recovery still clears the whole episode. Regression coverage drives the real auto-arm hook against a foreign session lock holder, asserts the ledger itself stays frozen, and fails before the fix in both the verified and unverified shapes; the existing unverified budget test now proves its budget actually ran out.
…enguid#4242) * feat(herdr): attach a real foreground viewer so the live-client teardown cases can be driven PR kunchenguid#4131 gated the Herdr active-tab close refusal on a live foreground client instead of the persisted `.focused` pointer, but only its two detached scenarios could be validated live. Every pseudo-terminal the runner built started at a zero-sized window grid, so Herdr registered no foreground client and `terminal title clear` kept answering `no_foreground_client`, leaving the four attached-client scenarios untested. That was a harness limit, not a product one. Add `fm-herdr-lab.sh viewer start|stop <session>`, backed by `bin/fm-herdr-lab-viewer.py`. The launcher sets the pty window size on the master fd BEFORE the fork, so the TUI cannot read the grid until it is already non-zero, and scrubs the inherited `HERDR_*` variables so Herdr's nested-viewer refusal does not fire when the helper runs inside one of its own panes. Attach and detach are both confirmed against the session's own foreground-client reason rather than assumed from a signal. The viewer inherits the lab's isolation contract: it attaches only to a session carrying this lab's ownership tripwire, never to `default`, and it signals only the processes it recorded, so a client someone else attached is never touched. Teardown now refuses while an owned viewer is still attached. Turn the reproduction into the regression with `tests/fm-herdr-attached-viewer-live-e2e.test.sh`, which drives kunchenguid#4131's scenarios 3, 4, 5, and 7 live against real Herdr and asserts the close refusal fires. Scenarios 4 and 5 need a focus change at one exact product boundary, so a PATH shim performs the real `tab focus` when the close helper issues its planning `pane get`. Removing either half of the recipe from the launcher makes the guard fail with the same `no_foreground_client` symptom kunchenguid#4131 reported. * test(herdr): fail loudly when an attached-viewer fixture cannot be created The fixture helpers run inside command substitutions, where fail() exits only the subshell and leaves the script running with empty ids. Return non-zero instead and carry the message at each call site. * fix(herdr): stop the viewer launcher's kill timer from raising on an exited child The SIGALRM escalation called os.kill unguarded, so a viewer that exited during the grace window turned an ordinary shutdown into a traceback inside the signal handler. * docs: list the lab viewer's pty engine in the bin toolbelt * no-mistakes(review): Harden Herdr viewer ownership and live CI coverage * no-mistakes(review): Validate viewer startup timeout and process ownership * no-mistakes(document): Document Herdr viewer safety contracts * no-mistakes(review): Fix viewer timeout to two seconds * no-mistakes(review): Cancel timed-out viewers and fix PTY grid * no-mistakes(review): Serialize viewer transitions and verify process parentage * no-mistakes(review): Harden viewer ownership locks and deduplicate CI * no-mistakes(review): Release interrupted locks and preserve viewer escalation * no-mistakes(review): Remove viewer locks and cancel interrupted launches * no-mistakes(review): Close viewer launch signal races * no-mistakes(document): Document attached Herdr viewer regression
…unchenguid#3766) * fix(bootstrap): allow nonvisual work without Lavish * no-mistakes(review): Gate scout brief Lavish line on bootstrap version floor
…kunchenguid#3825) * fix(tests): isolate fixture Git configuration from host preferences Ignore global and system Git configuration in the shared test library, which all four fixture helper entry points source. Keep local config, command-line overrides and explicitly supplied test config usable without changing the caller's environment or real project signing preferences. Exercise global and system signing inputs through all four helpers, real fixture and child commits, explicit signing overrides, unchanged input files, and signing refusal outside fixture subprocesses. Verification evidence for issue kunchenguid#3770: On pristine upstream f09de8a, all 12 reported suites failed and each logged "No secret key" using a private GIT_CONFIG_GLOBAL containing commit.gpgsign=true and gpg.format=openpgp, GIT_CONFIG_NOSYSTEM=1, and an empty private GNUPGHOME (GIT_CONFIG_COUNT and GIT_CONFIG_PARAMETERS unset). With this change, all 12 pass in the identical environment through bin/fm-test-run.sh --per-script-timeout-secs 900: fm-backlog-atomicity, fm-bootstrap-network-parallel, fm-bootstrap, fm-crew-state, fm-fleet-sync, fm-gate-refuse, fm-grok-harness, fm-session-start, fm-sessionstart-nudge, fm-tangle-guard, fm-test-run, and fm-update (all tests/<name>.test.sh). The new fm-test-fixtures regression failed before the library change and passes after it. Canonical bin/fm-lint.sh passes. Additional verification exposed fm-teardown's herdr-preflight-missing-adapter assertion on both this branch and an unchanged f09de8a archive with signing neutralized. That pre-existing failure needs separate disposition; it is not repaired or skipped here. The separately owned Muse and composer fixture defects remain untouched. Fixes kunchenguid#3770 * no-mistakes(review): Complete fixture Git isolation and scope config assertions * no-mistakes(review): Share Git isolation across standalone fixture entry points * no-mistakes(review): Map git-config helper changes to lib.sh dependents * no-mistakes(review): Select fixture-isolation regression on runner change; halve config matrix * no-mistakes(review): Scope fixture-isolation regression selection to the runner alone * no-mistakes(document): Give fixture Git isolation helper its owning header * no-mistakes(document): Record fixture Git-isolation coverage in fixtures suite header * no-mistakes(review): Fix linked-worktree fixtures and remove redundant Git isolation * no-mistakes(document): Correct stale runner-selection documentation * no-mistakes(document): Clarify family antecedent in isolation-proof runner evidence
… in auto mode (kunchenguid#4239) * feat(spawn): add config/claude-permission-mode to launch Claude workers in auto mode Every Claude worker launched with --dangerously-skip-permissions, and a captain who refuses bypass mode had no way to select Claude Code's classifier-reviewed auto mode instead. A new one-token local config, config/claude-permission-mode, selects the permission flag for every Claude launch: absent or `bypass` keeps today's launch byte-for-byte, `auto` swaps in --permission-mode auto, and any other value refuses the spawn before any endpoint, worktree, or record exists and names the accepted values. fm-spawn resolves the file on every spawn and relaunch, threads the flag through the Claude launch template for crewmates, scouts, and secondmates alike, and records claude_permission_mode=auto in the task meta only under auto so the default meta stays unchanged; a relaunch re-resolves rather than preserving the line. The file is a captain-wide safety preference, so it joins the inherited local material pushed into secondmate homes. The Claude adapter reference records the verified auto launch shape on Claude Code 2.1.269 and that it never meets the once-per-machine bypass confirmation dialog; docs/configuration.md owns the schema. * no-mistakes(review): drop unread claude_permission_mode meta line and its assertions
… untouched (kunchenguid#4243) * fix(teardown): refuse to return a Treehouse pool slot reassigned to another task A pool slot is reused across tasks, so a finished task's worktree= line can name a slot a different, live task now holds. Teardown already refused when a second task record named the same live path, but that scan cannot prove the record it is tearing down is the current owner: the task that took the slot next may leave no record the scan can reach - its own worker may have exited and its record been cleaned up, or it may live in a home this machine does not register. Teardown then killed every process under the path, hard-reset it and returned it, and its unlanded-work refusal never fired because it was inspecting a directory that no longer belonged to the task being torn down (observed 2026-09-07). Treehouse's own state file cannot answer the ownership question. It records a slot's owner as a live process lease (owner_pid plus owner_started_at, with `treehouse status` reporting in-use from the processes actually running under the path), which names no task and is released by the very event that makes a record stale - the worker exiting. An unleased slot therefore reads identical whether it is still this task's or has since been handed on, and a slot whose new holder has also exited but left uncommitted work reads as free. So the identity source is Firstmate's own claim, not Treehouse's lease. fm-spawn writes that claim - the task id - into the slot at the moment it takes it, under the same project lock that allocates the slot, and fm-teardown drops it only after the slot is genuinely returned. It lives at <pool>/<slot>/.fm-slot-owner, a sibling of the repo checkout rather than a file inside it, so claiming a slot can never dirty the copy the landed-work checks inspect. A claim naming another task, or one that cannot be read, refuses; --force does not lift either refusal, because --force authorizes discarding this task's unlanded work, never another task's live work. A slot that cannot be claimed refuses the spawn instead. An absent claim proceeds on exactly the record-scan protection it had before: slots taken before claims existed, and slots already returned, carry none, and refusing those would strand every task in flight across this change on no evidence at all. The refusal is deliberately all-or-nothing rather than partially completing the task's own cleanup. state/<id>.meta is the only durable record naming the worktree and endpoint, so removing it would destroy the evidence needed to reconcile which record is wrong, and its removal is one step with the backlog transition. Nothing is stranded: clearing the stale worktree= line leaves a record with no slot to release, which then tears down normally, and the refusal names that remedy. Repairing the previous claimant's stale worktree= line at spawn time is left for separate work. It would have the new owner write another task's record - the same class of cross-task mutation this bug is - and would need that record's own meta lock; with the claim in place teardown refuses on evidence rather than depending on the stale pointer having been scrubbed. For the same reason the relaunch path writes no claim: it holds no allocation lock, and a record whose worktree= is already stale would stamp the wrong task's claim onto a live sibling's slot. The regression reproduces the reuse sequence with only one discoverable record, including a clean, fully landed ship copy torn down without --force - the shape of the real incident, which the previous code returned to the pool - and fails against the previous code; the existing two-record, cross-home, own-slot and no-claim cases still pass unchanged. This builds ON upstream b028e8b (kunchenguid#3837), which is already in this branch's base (origin/main 40c50ea) and owns the record-exclusivity scan. Nothing here replaces that scan; the claim is the positive proof it cannot supply. Claude-Session: https://claude.ai/code/session_01JTBmuqKugaPUj7k9TXQwFS * no-mistakes(review): teardown leaves reassigned slot; spawn abort drops claim * no-mistakes(review): narrow Treehouse lease evidence; gate abort claim release on lock * no-mistakes(review): pin spawn-side slot claim; narrow abort-release header * no-mistakes(document): docs: point slot-claim rationale at fm-wake-lib owner
…nguid#4247) The reviewer treats Captain's intent as acceptance criteria, so a widened ask there drives over-built work; the spec should carry only what the ask requires.
* feat(bearings): name the Underway rows and order Charted Next newest filed first The fleet board's Underway rows led with the run status alone, so a scan told the captain where a pipeline stood but never which task the row was, and Charted Next rendered in backlog order rather than by when work was filed. The snapshot now projects the durable task name onto every in_flight row - from this home's backlog title, and from a secondmate home's own ledger for an active child - and the durable filed date onto every gate. The board's Underway row leads with that name and keeps the run status on its second line, and Charted Next renders newest filed first, with rows carrying no comparable date keeping their payload order after every dated row. The payload validator requires an explicit name marker on every Underway row and refuses a filed value that is not an ISO date, so the board can never sort on garbage or invent a label. * no-mistakes(review): Fix Bearings labels, bounds, and filed validation * no-mistakes(review): Fix Bearings identifiers and eligible queue bounds * no-mistakes(document): Document Bearings labels and newest-first bounds * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 56 to 59 Bearings tests. Verified the suite under /bin/bash 3.2: all 59 tests pass. git diff --check also passes
…nchenguid#4248) * fix(bearings): report the away-return catch-up instead of refusing A captain returning from away and asking for bearings got zero bytes and an error: fm-bearings-snapshot.sh ran the away-return guard with `|| exit $?` before reading any fleet state, so the mere existence of the catch-up gate killed every bearings mode (and /ahoy with them). Bearings now consults that guard rather than obeying it. fm-afk-return.sh separates its two refusal branches by exit status, so an ACTIVE away window still refuses exactly as before - the right answer there is to run the return first - while return catch-up (exit 4) lets collection and projection proceed and is disclosed as one action-free `(return-catchup)` gate row, following the existing `(main-inventory)` precedent. It stays out of decisions_open: these blockers are firstmate-actionable, not the captain's own call, and the per-task blockers already project as their own Underway rows. The guard's refusal text also stops promising a blocker list it cannot produce: a gate retained for a lifecycle reason alone now names that retention reason, and bearings carries the same reason in the gate row's title. Reporting is not ordinary work. AGENTS.md already scopes the return hold to work rather than reporting, so only the /afk and bearings skills needed the correction. * no-mistakes(document): Refresh away-return Bearings verification * no-mistakes(review): Reserve catch-up gate outside Bearings truncation * no-mistakes(review): Preserve filed dates in catch-up gate output * no-mistakes(document): Document reserved catch-up gate projection
…forked code-root copy (kunchenguid#4223) * fix(backlog): address the home's backlog from any directory and detect a forked code-root copy A home outside the code root forks its queue: the tracked .tasks.toml names data/backlog.md relative to tasks-axi's working directory, so a bare tasks-axi call from the code root writes the code root's data/ while session start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file over its target and rename(2) replaces a symlink: add, start, hold, and done from the code root each turn the link back into a regular file. The archive path is resolved against the working directory too, even with --file. bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any directory, using the lifecycle transitions' existing addressing (run from the data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It keeps relative --to/--*-file arguments meaning the caller's paths, and refuses a caller --file, an unresolvable home, and a symlinked home backlog. The fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which relied on cwd discovery, now go through it with an explicit FM_HOME and a cleared data override, so they keep addressing exactly $FM_HOME/data and an ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command names it instead of bare tasks-axi. Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only: when the home's data directory is not the code root's, a code-root data/backlog.md or data/done-archive.md that is not the home's own file is reported as a fork, with the merge procedure in bootstrap-diagnostics. * test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready The completion hint now points at the home-addressed command instead of a bare tasks-axi call, so the dependency-cleared follow-up assertion checks for that command. * no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh * no-mistakes(document): drop bare tasks-axi example from cd-guard doc * no-mistakes(lint): replace ls -A decoy listing with find for SC2012 * no-mistakes: apply CI fixes * revert: keep the compliance gate unchanged; the synchronize race is filed separately
* fix(spawn): pre-register Claude workspace trust for secondmate homes A claude --secondmate launch skipped workspace-trust registration entirely, so a standalone-clone secondmate home (an explicit ~/fm-homes/<id> path) had no store entry and its pane wedged on the "Is this a project you trust?" dialog before it read its charter. The step was gated on the task kind rather than on the harness, so the spawn's fail-closed guard had nothing to run against and reported a launch that could never start work. fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a whole firstmate instance, produced either as a leased worktree or as a standalone clone, so the linked-worktree test cannot decide it and the seed is the evidence instead: the .fm-secondmate-home marker must be a regular file this user owns naming exactly the id being spawned, the home must hold AGENTS.md and bin/, and each operational directory must resolve inside the home. That is the set fm-home-seed.sh writes and fm-spawn.sh's own home validation re-checks, so nothing wider than a home a secondmate spawn would launch into can earn home-level trust. The worktree path is unchanged, and still refuses a home. fm-spawn.sh now runs the registration for every claude launch and keeps refusing the spawn when it fails, rather than launching an agent that would wedge. * no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run When the base branch advances, GitHub cancels a pull request's in-flight run and re-triggers it. The cancelled run stays in statusCheckRollup beside the passing re-run, so the rollup can hold several runs of one check name at the same head while GitHub itself reports the pull request CLEAN. github_checks_not_green judged every run independently, so that superseded failure refused a genuinely mergeable pull request and pushed the operator toward a needless --allow-red. Group the rollup by the reported name and judge each check by its current run. Supersession is proven, never assumed: a name leaves the red set only when every one of its non-green runs is strictly older than one of its green runs, dated by the forge's own settled timestamp - a check run's completedAt once its status is COMPLETED, or a status context's createdAt - and only in the whole-second UTC form GitHub emits, which is the one spelling that orders correctly as plain text. A run with no such timestamp is never superseded, so a still-running, queued or undated run keeps its check red, and a name with no green run at all stays red. An unnamed entry is grouped alone so two unrelated unnamed checks are never treated as one. Every comparison is one-directional: it can only clear a failure a later success provably replaced, and never clears a check whose current run failed, is pending, or is missing. No other guard moves - the pull request must still be open, undrafted, mergeable, conflict-free and head-bound, and --allow-red still waives exactly its named check with every other check green. Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer SUCCESS for one check name and was refused; it now verifies, while kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse. * no-mistakes(review): Use check-run start times for safe supersession * no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266) * fix(merge): persist the merge authority on poll-detected merge outcomes The merge ledger tags a merge with the authority that permitted it while the away-posture record existed, but only the direct attended merge in bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge poll detected after the fact, published an untagged row, so exactly the merges no agent watched were the least auditable. bin/fm-merge-authority-lib.sh now owns that answer, read from the same structured sources the merge gate already used: the task's recorded yolo posture and the away-posture record's mechanical grant list, never prose. bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer; bin/fm-watch.sh only records it on the row its poll publishes, so reading the authority never becomes a second path to a merge. An unresolved answer records an untagged row rather than dropping the outcome or inventing an authority. * no-mistakes(review): Persist canonical merge authority for queued poll outcomes * no-mistakes(review): Harden merge authority persistence against lifecycle races * no-mistakes(review): Serialize poll authority publication with teardown * no-mistakes(document): Clarify persisted merge authority lifecycle * no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281) The 2026-09-12 Actions starvation incident found firstmate CI with no concurrency deduplication, so every superseded PR head kept its full 13-job fan-out, and four jobs with no timeout at all. Add per-PR supersession keyed on the PR number for pull_request events and on the unique run id for push events, cancelling only pull_request runs, so a new PR head replaces its own in-flight CI while every main push keeps its own group and is never cancelled. Add hang tripwires to the four previously unbounded jobs: 25 minutes for lint (measured at 14-16 minutes) and 5 minutes each for the coverage guard, the timing aggregate, and the repo invariants. Measured lane bounds are unchanged. tests/fm-ci-workflow.test.sh resolves the workflow's concurrency expressions against simulated pull_request and push contexts and holds every job's finite timeout.
…henguid#4288) Every other make_hold_home caller in this file skips when tasks-axi is absent; this test was the one unguarded call, so hosts without tasks-axi hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027) * fix(bin): bound each backlog row read so one wedged backend cannot blind a session start bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog backend once per item through fm_backlog_row_show, and that read was unbounded. A single wedged `tasks-axi show` therefore consumed the whole FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue, supervision instructions, fleet state, and context sections ever printed, leaving the fleet unsupervised with no live watcher. The harm was a blind startup, not a slow one. Bound the read with the existing shared timeout primitive (bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item it could not read and the loop continues to the next one. The first bound hit also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one bound rather than one per item and still names every item it skipped, which is what keeps the digest whole on a home carrying a large fleet. The bound holds regardless of any particular tasks-axi install, so it does not depend on the 0.2.5 `show` hang being resolved separately. * fix(bin): set the wedged-backend latch where it survives, and prove it The latch added with the read bound was inert. fm_backlog_row_show runs inside a command substitution in both of its status-capturing callers, so the subshell read the inherited value correctly but its write died with the subshell. Every item still paid a full bound and reported `exceeded`, never `skipped`, which left the large-fleet case the latch existed to cover completely uncovered. Move the write to the two callers that capture the read's status and own the surviving shell, and leave fm_backlog_row_show reading the latch only. Correct the comments that claimed an ownership the function never had. The test that was supposed to cover this asserted only that the second read finished under a generous ceiling, which is true whether or not the latch works. Assert instead that a latched read is strictly faster than one bound and that it reports its own item as skipped, so an inert latch fails the test. * test: cover every item the wedged-backend latch skips The latch assertion exercised a single skipped item, so "every skipped item is still named" was inferred rather than tested. Probe three items instead and assert each skipped one names itself and costs less than a bound. Verified as a real guard by removing both latch writes: the suite then fails on the first skipped item instead of passing. * no-mistakes(review): distinguish backlog read-bound hits from absent rows * no-mistakes(review): preserve read-bound status through the captain verify gates * no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows * no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites * no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS * no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure * no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285) * fix(merge): serialize the away-authority check with a synchronous merge bin/fm-pr-merge.sh read the away-posture record for merge authority (the per-task merge grant and the yolo/away-grant decision) and handed the merge to the forge afterwards. An archive at the captain's return or a grant revoked by a replacement record could land in between, so a merge could proceed on away authority that no longer held. The away record now carries a cross-subsystem lock, built on the existing bounded lock primitive rather than a new lock format: the record-mutating subcommands hold it across their mutation, and the merge holds it across both its authority read and the forge command. Because a queued or auto merge returns before the pull request lands, and would therefore outlive the lock, an away merge is now refused whenever it could land asynchronously: a requested --auto, a base branch whose merge-queue state does not prove an immediate merge, and GitLab's asynchronous flags and configuration. What remains permitted while away is the synchronous merge that lands inside the lock. This closes the common away-record/merge race against a live lock owner. It does not make the merge atomic in every case, and two narrow races are accepted and documented at their sites rather than hidden, both confused-agent-grade in the sense bin/fm-lease-lib.sh already uses: - A merge-queue rule change or a PR base change in the window between the queue-free preflight and the forge call can still enqueue the merge, which can then land after its grant lapses. - Killing the lock-owning shell while its gh or glab child is still running lets stale-owner recovery reclaim the lock and the record be archived or replaced, after which the orphaned child can complete the merge on lapsed authority. Closing either one needs landing verification or an ownership handoff, which is deliberately out of scope here. No existing gate is relaxed. The lock is taken after the live green-at-head verify and the captain-hold check, the in-lock authority read is unchanged, and a lock that cannot be taken refuses the merge rather than proceeding unlocked. The away grant stays a structured field; no prose is parsed. * no-mistakes(review): Fix GitHub rollup fixture base branch * no-mistakes(document): Document atomic away-authority merge locking * no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200) * feat(agy): verify Antigravity CLI as third worker/scout adapter Detection by anchored ancestry in fm-harness.sh (no marker of its own); bootstrap harness and effort validation; launch template with model and effort mapping plus reachable-catalog model validation; rendered-tail busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh; control mechanics with crewmate/scout-only refusal; tmux liveness naming; router entry with concise adapter reference; dated verification record; portable regression plus opt-in live drift guard. Verified live on agy 1.2.0: supervised spawn, durable steering, same-copy relaunch, and exit, with Herdr-native busy agreement. * no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature * no-mistakes(review): pre-register agy workspace trust, make readiness gate strict * no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching * no-mistakes(document): Document agy adapter in stale harness enumerations * no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound * no-mistakes(document): Fix stale test-shard snapshots after agy lane additions * no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental) * no-mistakes(test): Give agy typed sends a longer submit-confirm budget * no-mistakes(document): Document agy send budget, trust gate, and control coverage * no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
…uid#4337) * feat(afk): add quiet supervision mode for a present captain Adds a first-class quiet supervision mode alongside /afk for kunchenguid#2356: the same away-mode daemon, injection, busy/composer guards, classification policy, and reliability properties, but the captain staying present and chatting no longer exits it - only an explicit /quiet off does. state/.afk's first line now declares its mode (away, the default, or quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader, falling back to away for missing/empty/unreadable/unrecognized content (including the legacy bare-epoch-timestamp format written before mode existed) so nothing regresses. fm_afk_flag_write() preserves the on-disk mode on a bare refresh (no explicit mode given) rather than defaulting to away, which is what keeps the daemon's own redundant terminal-side re-write from silently resetting a captain's quiet mode back to away underneath them. New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing /afk for every shared mechanism, per the one-owner rule. AGENTS.md gains the state/.afk table entry and section 8's exit-trigger line. bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and bin/fm-guard.sh's stale-watcher banner all become mode-aware so a quiet-mode captain is never misdirected to /afk in captain-facing text. Closes kunchenguid#2356 * no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag * no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode --------- Co-authored-by: NewAiCoder <claude@theinbtw.com> Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#3578) * fix(bin): let verified harness ancestry outrank retained markers (#3) * fix(bin): let a structural harness ancestor outrank a retained marker bin/fm-harness.sh treated a verified environment marker as unconditionally authoritative, so a Codex session started from an environment that had retained CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned supervision protocol to a Codex primary, and every turn end was blocked for missing Claude recovery. The defect is the precedence boundary, not any one harness. codex, opencode, kimi, and muse publish no identity marker at all, so with markers winning outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering was a point patch on the same class of problem, and the launch-time marker clearing only ever covered sessions fm-spawn started. Markers and ancestry are now separate evidence layers that detect_own arbitrates: - no ancestry match, or no marker: the single available layer answers, unchanged; - same harness family: the marker's finer verdict stands, so a launch-selected pi-signed is not flattened to pi by an ancestry walk that can only see the shared launcher name; - different harness with a structural (command-name) ancestor: ancestry wins, because only ancestry proves who owns the process tree; - different harness with only a bare-interpreter script-path match: the marker wins, since a harness-shaped path in some node process's arguments is weaker evidence than a harness publishing its own identity. The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude worker nested under cursor either. Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so a real harness process can be asked what the walk makes of it. tests/fm-harness-precedence.test.sh is the portable regression, built from real renamed processes with no harness installed. Every case drives the two layers apart and asserts each alone as well as the combination, so no case can pass vacuously; it also pins Codex's real two-process install topology, since the fix depends on the native binary being what a tool subprocess meets first. The opt-in drift guard gains the matching live half: each installed harness's real running process must still be identified by the ancestry walk, and it fails naming the harness and version when a release changes that name. Documentation follows the corrected contract in the script header, the harness-adapters detection section, the codex, opencode, kimi, and cursor references, and a dated verification record. * fix(tests): drop the unused argument pass-through in the shim-topology helper bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and forwarded "$@", but every call site that varies the environment or passes the ancestry subcommand invokes the shim entry point directly, so the helper is only ever called with no arguments (ShellCheck SC2120/SC2119). Behavior is unchanged: with no arguments "$@" expanded to nothing. * fix(bin): examine the top of the process chain instead of assuming init harness_ancestry stopped as soon as the next pid was 1, on the assumption that pid 1 is always init and can never be a harness. Inside a PID namespace that assumption inverts: the harness itself is pid 1, so the walk never examined the one process that proves who owns the tree, reported no ancestry at all, and handed the verdict straight back to a retained marker. A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry precedence boundary in place. The same probe now resolves codex and renders the Codex foreground checkpoint. A host's real pid 1 (init, systemd, launchd) matches no harness name, so examining it costs one ps call and can introduce no false positive; the walk still stops once that top process has been read, and a non-numeric or zero ppid still ends it. tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that reports every process as bash with ppid 1 and pid 1 as the harness. The case asserts the marker still answers alone when pid 1 is host-shaped, so it cannot pass vacuously, and it fails against the previous stop condition. * docs(verification): record the real-Codex retained-marker evidence The existing record proved the precedence boundary with the portable regression and recorded each installed harness's process name behind the ancestry walk, but it had no evidence from a real Codex process actually holding a retained Claude marker, which is the failure the boundary exists for. Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`, with the exact command and the decisive verdict and rendered protocol on each side, and records the second boundary that shape exposed: the walk must examine the top of the process chain, because inside a PID namespace the harness is pid 1. Refreshes the portable regression's observed output for the case it gained. * no-mistakes(review): blind ancestry in marker-pinned harness tests * no-mistakes(review): blind ancestry in the Pi guard-routing test * no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims * no-mistakes(review): model the spawn-and-wait Codex shim topology * no-mistakes(document): correct stale muse marker-clearing detection claims * no-mistakes: apply CI fixes * fix(bin): examine the top of the chain in the lock and nudge walks too The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two other harness-ancestry walks, on the exact topology the branch verified against a real Codex process. bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could not find that harness at all and did not recognize its own session lock. bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a lock pid of 1, so the same session was told to run session start again on every turn. Both walks now compare the top process before stopping, matching the shape used in bin/fm-harness.sh. For the lock walk this is safe because fm_harness_process_matches rejects a host's real pid 1. For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged `kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent rather than acting on init. Each walk gains one regression case. The lock case drives a deterministic process table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace, because the builtin `kill -0` gate cannot be reached through a fake ps, and it first proves the same fixture nudges with no lock present; it skips explicitly where unprivileged namespaces are unavailable. * no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard * fix(bin): verify the live harness guard at the strength the guarantee needs The marker-versus-ancestry boundary this branch ships is a strength claim: detect_own hands an args-strength verdict straight back to a retained foreign marker, so a harness is only protected where the ancestry walk reaches it at comm strength. The installed-harness drift guard probed the pane process alone. Under an interpreter shim the pane process IS the shim, whose own script path is args strength, while the native binary that carries comm strength is its child. The guard therefore observed args for Codex, passed, and would have kept passing if a release stopped spawning that native child at all, while real sessions silently regressed to the original bug. fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane process and every descendant of it, the vantage a tool subprocess actually occupies. The guard now requires comm strength somewhere in that set and requires every vantage to name the same harness. This supersedes the preceding commit's in-guard leaf walk, which reached the same vantage but left the logic inside the test file, where CI could not pin it and nothing else could reuse it. A harness-dependent check needs both halves: `tests/fm-harness-precedence.test.sh` now carries a portable case proving the subtree probe reaches a strength the top-of-session probe cannot, mutation checked twice, once against the pre-change script and once by disabling descendant enumeration. The subtree walk also avoids depending on tty and process-group semantics that differ between Linux and macOS. Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code 2.1.257 reports [comm claude]. * no-mistakes(review): narrow drift guard to the upward vantage path * no-mistakes(review): judge only comm-strength vantages in drift guard * no-mistakes(document): drop duplicated rationale in detection precedence evidence * no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion * no-mistakes(document): drop branch-relative phrasing in detection precedence evidence * no-mistakes(review): guard remaining empty positional expansions in fm-harness * no-mistakes(document): scope cursor marker-ordering claim to the marker layer * no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties * no-mistakes(document): Document comm-strength descent tie-break --------- * no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests * no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript * no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs * no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…rkers (kunchenguid#3944) Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved) reads only the canonical git-root project entry in ~/.claude.json, which its own worktree-to-primary-checkout canonicalization means is never the task worktree fm-claude-trust.sh registered. The trust dialog kept working previously only because its check has an ancestor-walk fallback that happens to reach the worktree entry; the external-imports check has no such fallback. Verified by disassembling the installed claude binary and reproducing in an isolated three-way tmux launch: identical flags registered only at the worktree key still showed the external-imports dialog, and registering them at the primary checkout key suppressed both dialogs. fm-claude-trust.sh now registers all three flags on both the worktree entry and the primary-checkout entry in one atomic write, and refuses when the <project> argument is not itself a primary checkout (its own write target would then be wrong). Extends the harness-adapters Claude reference and the trust test suite. Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355) The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves state/.watcher-down behind in an acked:* state after a downtime episode is handled. health_snapshot's presence check reported that as an open gap on every later return, so a handled episode kept surfacing as a false GAP forever.
…kunchenguid#4361) * fix(update): rebind fm-procevent-when watches after a self-update A self-update fast-forwards bin/ in place, changing an armed watch's action executable bytes with no tampering involved. The watch's trust binding was hashed at arm time, so the very next fire was refused as not matching the registered binding and the watch died silently. Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the trust binding for every watch whose action executable lives under FM_ROOT, using the same spec/trust validation as an ordinary fire, and leaves any watch whose action lives outside FM_ROOT untouched. Wire it into fm-update.sh right after a successful fast-forward, for both the primary home and any local secondmate home that advances. * no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check * no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence * no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016 * no-mistakes(review): Reload trust binding from disk before firing to reach live pollers * no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race * no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows --------- Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…e-head merge gates, mail plane Brings the fork up to upstream waypoint b182d0f ("fix(bin): rebind fm-procevent-when trust bindings after a self-update (kunchenguid#4361)"), 47 upstream commits over 225 files, +32120/-2373, on top of batch 18 (891dc51). The range carries the durable AFK posture lifecycle (kunchenguid#4048) and quiet supervision mode (kunchenguid#4337), the agy adapter (kunchenguid#4200), live-head merge gates and away task grants (kunchenguid#4199, kunchenguid#4285, kunchenguid#4266), the IMAP/SMTP mail plane (kunchenguid#3765), config/claude-permission-mode (kunchenguid#4239), the Claude turn-end re-block budget bound against a frozen auto-arm epoch (kunchenguid#4221), live-over-terminal no-mistakes run preference (kunchenguid#2881), bounded backlog row reads (kunchenguid#4027), superseded-PR CI cancellation (kunchenguid#4281), process-event runner reaping (kunchenguid#3904, kunchenguid#4212), Herdr fixes, and the trust-binding rebind after self-update (kunchenguid#4361). 35 files conflicted. Each is resolved under this fork's sync conflict rule in one of three classes: 1 union, 2 take upstream (fork mechanism superseded), 3 take upstream and re-append fork-only sentences whose mechanism still exists in code. CLASS 1 - UNION .gitignore: upstream's blank line and `.tools/`, then the fork's `.fm-lint-parity.*` comment and entry. bin/fm-supervise-daemon.sh housekeeping: local state=$1 now due f key task win marker age last max_defer oldest pause_secs marker_epoch until bounded_until pause_reason liveness_age bin/fm-supervise-daemon.sh handle_wake (auto-merged, not a marked conflict): the fork's own `needs-decision:*)` arm (c7e0d41, batch 12) sat after upstream's new `signal:*|needs-decision:*)` arm (5300a4f, kunchenguid#4169), which classifies the same rows the same way, so the fork arm was unreachable (shellcheck SC2221/SC2222). The fork arm is removed; upstream's arm is unchanged. bin/fm-teardown.sh remote secondmate teardown: rm -f -- "$STATE/$ID.turn-ended" "$STATE/$ID.progress" \ "$STATE/$ID.liveness.sh" "$STATE/$ID.liveness-trust" bin/fm-test-run.sh family lists: live-harness-optin keeps upstream's fm-agy-signals-live-e2e.test.sh and the fork's fm-nm-status-shape-live-e2e.test.sh; the afk family reads fm-afk-contract.test.sh|fm-afk-inject-e2e.test.sh|fm-afk-return.test.sh|fm-afk-shell-pane-refusal.test.sh) bin/fm-spawn.sh: the effort case keeps the fork's literal `default` and upstream's native-only `ultra`: ''|default|low|medium|high|xhigh|max|ultra) ;; *) echo "error: --effort must be one of default, low, medium, high, xhigh, max, ultra" >&2; exit 1 ;; and the header carries upstream's ultra sentence, then the fork's literal-"default" sentence. bin/fm-config-inherit-lib.sh: upstream's claude-permission-mode paragraph, then the fork's launch-env-allowlist sentence. FM_INHERITABLE_CONFIG already lists both, so both are really inherited. bin/fm-pr-merge.sh: every fork guard and every upstream refusal survives whole, and every refusal runs before upstream's record-before-forge-call block. Order: 1. one argument loop: --attended-override, --allow-red, --require-ancestor, then `--) shift; break ;;`; 2. the GitLab --allow-red refusal, then the --require-ancestor SHA check; 3. upstream's reject_* calls, auto-merge detection and fm_backlog_directory_present; 4. the fork's FM_MERGE_GUARD posture case; 5. upstream's META, wake-lib source, lease, meta check and spawn-gen read; 6. the fork's meta_field, recorded waypoint read and guard functions; 7. the fork guard window, then upstream's control lock onward unchanged. The guard window reads: # The merge guards run before the per-task control lock below is taken. A # check can run for minutes, and the watcher's merge poll and every other # per-task control operation wait on that lock, so holding it here would # stall them for the whole check. trap 'fm_static_guard_cleanup' EXIT if [ -n "$REQUIRED_ANCESTOR" ]; then required_ancestor_assert || exit 1 fi ...upstream-history guard, static merge guard, merge_guard_record... fm_static_guard_cleanup and upstream's block before the forge call is byte-identical: away_status=0 require_current_away_authority || away_status=$? [ "$away_status" -eq 0 ] || exit "$away_status" require_recorded_pr_identity || exit 1 record_pr_metadata || exit 1 require_released_captain_hold || exit 1 Three fork lines did not survive, each because it would break an upstream guarantee: - the fork's early `record_pr_metadata || exit 1` ran before require_recorded_pr_identity; bin/fm-pr-check.sh re-records pr= freely, so it would have defeated upstream's PR-identity rebind refusal; - the fork's `trap 'fm_static_guard_cleanup' EXIT` for the rest of the script would have replaced upstream's `trap merge_control_cleanup EXIT` and leaked the control, meta and away locks; it now covers only the guard window and upstream's trap is unchanged; - the fork's top-of-file fm-wake-lib.sh source is dropped because upstream now sources the same library before its first use. Upstream's record_pr_metadata definition is back at its upstream position. The only upstream line changed in the file is Usage: # Usage: fm-pr-merge.sh <task-id> <pr-url> [--attended-override] [--allow-red <check-name>] [--require-ancestor <sha>] [-- <extra forge merge args>] The guards run before upstream's per-task control lock (not after it) because the static check has a 900s budget and that lock is waited on without bound by the watcher's new PR-poll lock, fm-control, fm-teardown and fm-captain-hold. A fork guard refusal therefore no longer re-arms the merge poll at merge time; bin/fm-pr-check.sh arms it at PR-ready time. bin/fm-watch.sh: upstream's PR_POLL_CONTROL_LOCK= and pr_poll_control_release() first, then the fork's watcher_close_has_nothing_to_recover() (PR 47). bin/fm-nm-run-lib.sh: upstream's fm_nm_run_status_class first, then the fork's fm_nm_head_allows_inflight and fm_nm_head_binds_run. bin/fm-crew-state.sh keeps upstream's comment (live-over-terminal exception and the sibling-run paragraph) followed by the fork's bounded-and-checked list-call paragraph. tests/lib.sh: upstream's process-event runner reaping block and FM_TEST_STUB_MAX_BLOCK_SECONDS, then the fork's guarded fixture removal block. fm_test_cleanup reaps first, then removes through fm_test_rmtree. tests/fm-crew-state.test.sh reset_fakes: export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_HERDR_PROCESS FM_FAKE_HERDR_SHELL_PID FM_FAKE_CI_LOGS export FM_FAKE_DAEMON_DOWN FM_FAKE_NM_HANG FM_FAKE_NM_CALLS FM_CREW_STATE_NM_TIMEOUT tests/fm-captain-hold-lifecycle.test.sh: `. lib.sh || exit 1` plus upstream's fm-timeout-lib.sh source; both invocations, upstream's test_board_answer_reaches_the_keyed_answer_intake first. tests/fm-brief.test.sh: upstream's scout Lavish floor test, then the fork's three report-contract tests. bin/fm-brief.sh hunk 2 keeps the fork's six-field scout report list and replaces only its Lavish sentence with upstream's $LAVISH_LINE. docs/configuration.md: the dispatch paragraph carries upstream's `ultra` sentence, then the fork's PROFILE and SPAWN `default` sentences; the env list carries upstream's FM_MAIL_CHECK_BUDGET, FM_MAIL_POLL_MAX_WAKES and FM_MAIL_TIMEOUT, then the fork's FM_MERGE_GUARD, FM_STATIC_CHECK_TIMEOUT, FM_MAIN_GUARD_BUDGET and FM_STATIC_GUARD_FETCH_TIMEOUT. docs/scripts.md: the fork's fm-teardown.sh row ("require completed scout and ship completion reports") and upstream's fm-harness.sh row (native-only `ultra`). docs/captain-hold-lifecycle.md: the fork's regression paragraphs, then upstream's extended projection line. .agents/skills/firstmate-coding-guidelines/SKILL.md: the fork's three Bash 3.2 empty-array bullets, then upstream's rewritten lint bullet. .agents/skills/harness-adapters/SKILL.md, docs/tmux-backend.md, docs/architecture.md hunk 2: upstream's wording with the fork's Rovo adapter kept: Muse, Gemini, Rovo, and AGY are verified only for crewmate and scout work, never a secondmate or primary. It classifies recognized Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Gemini, Muse, Rovo, and AGY process identities as `alive`, ... Gemini stays in the tmux-backend sentence because upstream's own bin/fm-agent-process-lib.sh classifies it; upstream's process-name vocabulary sentence follows. CLASS 3 - UPSTREAM, THEN FORK SENTENCES WHOSE MECHANISM EXISTS `grep -c FIRSTMATE_OP_END bin/fm-operational-input.sh bin/fm-supervise-daemon.sh` is 2 and 0 on the merged tree, and fm_operational_terminator_kind plus the daemon's message_is_injection are present, so the double-ended envelope sentences stay. .agents/skills/afk/SKILL.md: upstream's return paragraph and bullets, then - A message that carries the envelope's **trailing terminator** (U+2063 followed by `FIRSTMATE_OP_END: v1 <kind>`) but no leading prefix -> a daemon escalation whose front was cut in delivery: stay away, and recover the full text from `state/.supervise-daemon.log`, which the digest names. Do not treat it as the captain, and do not act on the surviving fragment alone; it is a tail, not a body. after the Re-invoking bullet, and after upstream's /quiet line The exit bias applies to genuinely ambiguous input, not to a cut escalation: ... `bin/fm-supervise-daemon.sh`'s `message_is_injection` is the executable form of this whole predicate. Hunk 2 takes upstream's portability line (claude, codex, opencode, grok, and kimi), then "The marker is repeated at the end because ...". AGENTS.md: upstream's `state/.afk-contract` bullet, then - The daemon's injection envelope is marked at both ends, so a message that begins mid-sentence and ends with the U+2063 `FIRSTMATE_OP_END:` terminator is a daemon escalation whose front was cut in delivery, never the captain talking; stay away and recover the full text from the daemon log the digest points at. The state-layout line drops `.claude-autoarm-absent` and "guard-recorded non-participation" (class 2 below), matching upstream. docs/architecture.md hunk 1: upstream's rewritten daemon sentence, then the fork's "That envelope is marked at both ends ..." sentence. bin/fm-brief.sh hunk 1: upstream's `paused:` wording including `until <YYYY-MM-DDTHH:MMZ>`, with the fork's park-moment rule (the paused-status mechanism exists) inserted before "Use `blocked:`": Whatever the wait, append it at the MOMENT you park: starting a gate or a full test run and waiting for it, waiting on a validation pipeline step, or any other wait you expect to clear without firstmate doing anything. Append a \`working: ...\` line when it resumes. tests/fm-backend-herdr.test.sh, tests/fm-remote-doctor.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh: upstream's new fixture sourcing lines, keeping the fork's `|| exit 1` on the first source line. CLASS 2 - TAKE UPSTREAM, FORK MECHANISM SUPERSEDED bin/fm-turnend-guard.sh, docs/turnend-guard.md, docs/supervision-protocols/claude.md: upstream's files byte-for-byte. The fork's PR 33 (deb6698) non-participation recording is replaced by upstream's budget_account_current_epoch re-block charging (kunchenguid#4221). tests/fm-turnend-guard.test.sh is upstream's file plus the fork's `|| exit 1` and PR 41's test_hook_x_mode_reason_keeps_stop_owned_contract in place of test_hook_x_mode_reason_sources_cadence (invocation list too). Remnants removed: `.claude-autoarm-absent` from bin/fm-wake-lib.sh's recovery reset, the fm-session-lock-lib.sh row wording in docs/scripts.md, the stood-down fail-open section and its header fact in tests/fm-claude-stop-autoarm-live-e2e.test.sh, and the matching sentence in docs/verification/supervision.md. That live test keeps its vendor-fact section (both Stop hooks run on a refused Stop), which drives only synthetic lab hooks and still describes upstream's contract. bin/fm-afk-launch.sh and tests/fm-afk-launch.test.sh: upstream's files. The fork's fm_afk_launch_resolve_captain_pane (3e5c89f, #23) is superseded by upstream's posture-record entry. Upstream still never splits or co-tenants the captain's pane on Herdr: it launches the daemon in a dedicated `--no-focus` workspace, and the daemon keeps #23's startup refusals with a durable alarm and the per-digest agent proof. bin/fm-wake-lib.sh: both fork fm_lock_section_enter/leave "$FM_WAKE_QUEUE_LOCK" lines (2386360, #7) are removed. Upstream split fm_wake_append into a wrapper that takes the queue lock and fm_wake_append_locked, so the fork's lines would re-acquire a lock the caller already holds. tests/fm-test-run.test.sh: the fork's install_runner (d6190d3) stays at all eight sites and also copies upstream's tests/git-config-helpers.sh: install_runner() { # <fixture-bin-dir> local bin=$1 cp "$RUNNER" "$bin/fm-test-run.sh" cp "$ROOT/bin/fm-test-env-lib.sh" "$bin/fm-test-env-lib.sh" cp "$ROOT/tests/git-config-helpers.sh" "$bin/../tests/" chmod +x "$bin/fm-test-run.sh" } tests/fm-watcher-lock.test.sh - CLASS CHANGED from 2 to 1: the fork's two-adapter X-mode case (Claude asserts the Stop-owned repair and never the cadence source; Grok asserts the arm line and the source line) passes alongside upstream's rewrite once upstream's blind PATH is prepended, so both guard calls read `PATH="$blind:$PATH" CLAUDECODE=... GROK_AGENT=... FM_HOME=...`. MERGE CONSEQUENCES FIXED IN THIS MERGE (each reproduced first) bin/fm-check-lib.sh: upstream's new bin/fm-mail-check.sh rolls a failed re-arm back through fm_custom_check_registered, a name the fork's generalization (f9a59bb, #1) replaced with fm_task_script_registered <kind>. On the merged tree the call was "command not found" and the rollback deleted a check that was still bound; upstream's tree keeps it. A one-line wrapper under upstream's name restores upstream's behaviour, pinned by tests/fm-liveness-source.test.sh test_failed_rearm_keeps_a_still_bound_check (fails without the wrapper, passes with it). bin/fm-captain-hold.sh tasks_gated_by: upstream's bounded row reads (kunchenguid#4027) changed task_show to hand its row back in TASK_SHOW_OUTPUT instead of printing it. The fork's freed-work reminder read `show=$(task_show ...)`, got empty text and never printed (tests/fm-captain-hold-lifecycle.test.sh "answering did not remind the caller to re-check the freed work's real preconditions"). It now reads the variable inside the same substitution, keeping the reminder advisory. tests/fm-captain-hold-lifecycle.test.sh, tests/fm-teardown-endpoint-safety.test.sh: three new upstream cases tear down a ship task without `--force` and without the completion report the fork's teardown requires (745a468, #6), so teardown refused before the step each case is about. Each seeds a placeholder report first, as batch 10's 840ff3a did for the same class; tests/fm-teardown.test.sh keeps owning the report contract. tests/fm-watch-triage.test.sh: upstream's new away-record captain-hold case (kunchenguid#4048) acknowledged its silent phase-A watcher stop with the strict ack_stopped_cycle. The fork's PR 47 (watcher_close_has_nothing_to_recover) opens no recovery episode for a stop that delivered no wake, so there was nothing to acknowledge ("could not acknowledge the intentional phase-A stop"). That call now uses the fork's ack_stopped_cycle_if_any, as every other phase-A stop in the file already does. tests/fm-static-guard.test.sh, tests/fm-merge-waypoint-guard.test.sh: the fork's merge-guard suites drove the old merge path. Upstream's GitHub merge now reads `gh pr view --json ...statusCheckRollup` before merging and merges with `--match-head-commit`, so the gh mocks answer that read and the assertions check the exact merge line from a gh call log. The no-pull-ref case seeds pr_head= in meta because the guards now run before the merge re-records the PR. tests/fm-send-agent-pane-live-e2e.test.sh: its failure message pointed at fm_backend_tmux_classify_process_name, which upstream moved (kunchenguid#4191); it now names bin/fm-agent-process-lib.sh's fm_agent_process_classify_name. tests/fm-test-fixtures.test.sh: upstream's new Git-config isolation fixture (eb0ea3a, kunchenguid#3825) copies bin/fm-test-run.sh and bin/fm-timeout-lib.sh into a scratch runner, but the fork's runner (8f41910, #41) refuses to start without bin/fm-test-env-lib.sh beside it ("fm-test-run: missing bin/fm-test-env-lib.sh beside this runner"). The fixture now copies that library too, as d6190d3 did for tests/fm-test-run.test.sh. Source guards: the fork's tests/fm-test-fixture-cleanup.test.sh (#29) pins that every tests/ file sourcing a tests/-local helper refuses to run from outside tests/, the `|| exit 1` tests/lib.sh's header calls load-bearing. Fifteen test files new in this range source lib.sh without it. Run from outside tests/, tests/fm-agy-signals-live-e2e.test.sh lost its opt-in gate along with lib.sh and launched a real agy prompt ("fm-agy-signals-live-e2e.test.sh ran real checks from outside tests/ instead of refusing"). The first helper source line of each now carries `|| exit 1`, as batch 10's 806826c did for the files that batch brought in: tests/fm-afk-contract.test.sh, tests/fm-agy-harness.test.sh, tests/fm-agy-signals-live-e2e.test.sh, tests/fm-backend-herdr-agent-exit-shell-e2e.test.sh, tests/fm-backlog-read-bound.test.sh, tests/fm-ci-workflow.test.sh, tests/fm-harness-precedence.test.sh, tests/fm-herdr-attached-viewer-live-e2e.test.sh, tests/fm-herdr-pi-stale-registration-live-e2e.test.sh, tests/fm-mail-check.test.sh, tests/fm-mail.test.sh, tests/fm-pi-codex-native.test.sh, tests/fm-remote-herdr-guard.test.sh, tests/fm-send-agy-confirm.test.sh, tests/fm-tasks-axi.test.sh. The sweep then passes over 194 dependent files. VERIFICATION Silent-splice sweep over the 108 files both sides changed since 891dc51: every non-blank line upstream added is present except the deliberate widenings above (harness-adapters, fm-pr-merge Usage, the two fm-spawn effort lines, the daemon local, the teardown rm, the fm-test-run afk family, the tmux-backend sentence) and upstream's eight git-config-helpers cp lines folded into install_runner. Of the 117 upstream-only files, 102 equal upstream and the 15 above differ only by the `|| exit 1` source guard. Every function body both sides changed is upstream's body plus additive fork lines. Fork-added calls into functions upstream changed or removed were swept too, which is how the task_show and moved-classifier consequences were found; the full portable-serial lane found the last two. `git diff b182d0f -- .github/workflows/` is empty, and all 93 fork commits remain in ancestry.
…aude launch settings prose
…f this sync batch) made bin/backends/herdr.sh's fm_backend_herdr_pane_process_state cross-check a registered herdr agent against the pane's actual foreground OS process, not just the self-reported `herdr pane report-agent` call. tests/fm-afk-inject-herdr-e2e.test.sh's supervisor pane ran its loop script directly under a bare `bash` interpreter, whose process name classifies as "shell" (fm_agent_process_classify_name in bin/fm-agent-process-lib.sh), so despite the script faithfully self-registering as an agent, fm_backend_herdr_agent_state now reads it as stale-agent/dead and the daemon refuses to start — causing "Behavior tests (Herdr)" to fail (FM_TEST_SUMMARY failed=1; the printf "Broken pipe" lines in the CI log are pre-existing benign noise from fm_backend_herdr_events_capable's grep-early-exit pattern and are unrelated to the failure). The fix (found already staged uncommitted in the worktree from a prior investigation session, which I verified rather than re-deriving) launches the supervisor loop script under a symlink named `claude-shim` pointing at the real bash binary instead of bare `bash`, so the OS process name matches the `*claude*` agent pattern and classifies as "agent". This mirrors the identical, already-proven AGENT_SHIM pattern used by the analogous tmux fixture tests/fm-afk-inject-e2e.test.sh. It's a symlink (not a copy) because a copied binary fails macOS arm64 code-signing. I confirmed the fix is minimal (test-file only, 14 insertions/1 deletion, no production code touched) and ran tests/fm-afk-inject-herdr-e2e.test.sh directly against a real local herdr install (0.8.2) — all 4 scenarios (A-D) now pass: "all real-herdr afk injection e2e tests passed". No other files are modified in the worktree
…-lib.sh deferred HUP/INT/TERM signals only AFTER attempting `fm_lock_acquire_wait`. Lock acquisition is not atomic — it calls `fm_current_pid`, which forks a subshell via command substitution on any shell lacking `BASHPID` (confirmed in the CI/test environment). A HUP/INT/TERM landing in that window hit the pre-deferral disposition and could make the watcher exit while abandoning a lock it never finished tracking, leaving `.wake-queue.lock` stranded on disk. This was the exact scenario CI's `tests/fm-watcher-signal-safety.test.sh` exercises (deterministically pins the watcher in a contended critical section, sends TERM, and asserts the wake-queue lock is not left behind) — it was the only failing assertion ("Behavior portable serial 2" / family=watcher-wake-lock). Fix (found already staged uncommitted in the worktree from a prior investigation in this same session; I verified it rather than re-deriving it): reorder `fm_lock_section_enter` to call `fm_signal_defer_begin` BEFORE `fm_lock_acquire_wait` instead of after, so the whole (non-atomic) acquisition sequence runs under deferral. On acquire failure it now explicitly calls `fm_signal_defer_end` (unwinding deferral and re-raising any pending signal) before returning 1, preserving the existing caller failure-path contract. This is a single-function reorder, no new subsystems — and it matches a pre-existing comment elsewhere in the same file (`fm_lock_try_acquire`) that explicitly warns a different fix for this same bug class (a recursion depth bound) was tried and reverted, confirming the reorder is the validated direction. Verification performed: - `tests/fm-watcher-signal-safety.test.sh` (the actual failing script) now passes all 5 assertions, including "the signalled watcher abandoned the wake-queue lock mid-section" — reproduced clean across 5 repeated runs (this was flagged as an intermittent timing race, so flakiness was specifically checked). - The other 4 scripts in the same `watcher-wake-lock` CI family (fm-inactive-reconcile, fm-watch-arm, fm-watch-recovery-loop, fm-supervision-events) all still pass — no regression from reordering the shared `fm_lock_section_enter` primitive used by ~15 call sites. - `bash -n bin/fm-wake-lib.sh` — syntax OK. Only bin/fm-wake-lib.sh is touched (14 insertions/1 deletion net, matching the diff already in the worktree); no other files were modified, bin/fm-lint.sh was not run (lint is a separate phase), and no full-suite run was performed — only the targeted family relevant to the failing check, per the batch's test-scope constraints. The fix is left in the working tree uncommitted for the outer no-mistakes executor to carry forward, since this ci phase does not own commit/push/PR steps
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Captain's standing policy for this fork (2026-09-05, verbatim): "the main goal for firstmate is to become sync with upstream. once we reach that state, we avoid introducing more changes and stay aligned with the upstream. until then you can keep adding work that helps aligning with the upstream easier and reaches the intended state quickly."
Captain's conflict rule (2026-09-05, verbatim): "if there is something upstream and our code both worked on and fixed but approaches differ, take the upstream change, however, if our change is better than upstream, file a PR. for other items, take the upstream changes."
Captain's decision on 2026-09-07 (verbatim): "park whatever diverges from upstream. our goal was to sync with upstream for firstmate. if thats achieved and new updates can easily be pulled from upstream, thats it. no further work on firstmate." Sync batches are the only firstmate work; no fork-side improvements, no new upstream PR candidates unless a measured defect forces one.
Captain's cadence rule: a sync batch every fortnight or every ten accumulated upstream commits, whichever comes first. Batch 18 landed 2026-09-08 at waypoint 891dc51; 47 upstream commits have accumulated since, so the ten-commit trigger fired.
The ask this task serves: upstream sync batch 19 - bring the fork's
origin/main(now 232cc7c: batch 18 landed as #52, waypoint 891dc51) up to upstream's current tip, waypointb182d0f908b78d08c7ccb8dce3775bdca8c5d657("fix(bin): rebind fm-procevent-when trust bindings after a self-update (kunchenguid#4361)"). Forty-seven upstream commits, 225 files, +32120/-2373, including: the durable AFK posture lifecycle (kunchenguid#4048) and quiet supervision mode (kunchenguid#4337); the Antigravity CLI (agy) adapter (kunchenguid#4200); live-head merge gates and away task grants (kunchenguid#4199) with away-authority serialization for synchronous merges (kunchenguid#4285) and persisted merge authority for poll-detected outcomes (kunchenguid#4266); the IMAP/SMTP mail plane (kunchenguid#3765); config/claude-permission-mode (kunchenguid#4239); the Claude turn-end re-block budget bound against a frozen auto-arm epoch (kunchenguid#4221); live-over-terminal no-mistakes run preference (kunchenguid#2881); bounded backlog row reads (kunchenguid#4027); superseded-PR CI cancellation and job bounds (kunchenguid#4281); process-event runner reaping and launch-storm prevention (kunchenguid#3904, kunchenguid#4212); Herdr liveness, teardown and remote fixes; and the trust-binding rebind after self-update (kunchenguid#4361).Test evidence for this batch: the full portable-serial lane already ran once (178 scripts; results recorded in the completion report, 28 host-caused failures listed there). The test step must run only the targeted suites named in the brief post-merge checks (fm-turnend-guard, fm-claude-stop-autoarm, fm-watch-arm, fm-watch-recovery-loop, fm-watch-triage, fm-wake-queue, fm-pr-merge plus the require-ancestor and static-guard suites, fm-crew-state, fm-nm-*, fm-afk-launch, fm-afk-return, fm-brief, fm-captain-hold-lifecycle, fm-test-run, fm-spawn, fm-backend-herdr, fm-remote-doctor) and must NOT run bin/fm-lint.sh (the lint step owns it) or the whole suite. Known host-only failures listed in the brief are reported, not fixed.
What Changed
bin/fm-afk-contract.sh,bin/fm-agent-process-lib.sh, expandedbin/fm-afk-launch.sh/bin/fm-afk-return.sh, new.agents/skills/quiet/SKILL.md), the Antigravity CLI (agy) worker/scout adapter (newbin/fm-agy-trust.sh,.agents/skills/harness-adapters/references/harness/agy.md,docs/verification/agy.md), and the IMAP/SMTP mail plane (newbin/fm-mail.sh,bin/fm-mail.py,bin/fm-mail-check.sh).bin/fm-merge-authority-lib.sh,bin/fm-landed-lib.sh,bin/fm-remote-herdr-guard.sh,bin/fm-remote-herdr-owner-lib.sh, expandedbin/fm-pr-merge.sh,bin/fm-captain-hold.sh,bin/fm-crew-state.sh), plus the post-self-update trust-binding rebind forfm-procevent-when.sh.bin/fm-procevent.sh/bin/fm-procevent-lib.sh), bounds backlog row reads and CI job/superseded-PR handling (.github/workflows/ci.yml), and updates Herdr backend/lab tooling (newbin/fm-herdr-lab-viewer.py) alongside a large expansion of the test suite (new suites for afk-contract, agy, mail, herdr-guard, backlog-read-bound, harness-precedence, tasks-axi, and more) and docs to match.Risk Assessment
🚨 High: All four dispatched deep-verification passes have now completed. The batch's own novel 'merge consequence' fixes (fm-check-lib.sh wrapper, fm-captain-hold.sh reminder, ack_stopped_cycle_if_any swap, static-guard test mocks) and the fully-rewritten authorization/lock ordering in bin/fm-pr-merge.sh both check out as correct, with only minor follow-up-grade issues (a metadata-ordering quirk, a vacuous test assertion, dead code, log noise). The material risk is a security review that found two concretely reachable authorization bypasses in newly-introduced-to-this-fork, security-critical code: the away-authority merge gate fails open when its state file is absent, so any same-uid crewmate process can delete/replace it to collapse 'away, ungranted' into 'attended' and let a merge proceed without the required grant; and the away record's confirm path accepts a fresh merge grant with no captain-presence evidence, which now chains with this same batch's new IMAP/SMTP mail plane so that an inbound email could plausibly prompt-inject the away-supervision session into minting an unauthorized grant. A third finding shows the post-self-update trust-binding rebind (kunchenguid#4361, this batch's own tip commit and headline fix) can itself be defeated via a symlink, undermining the protection it claims to add. All three are inherited verbatim from upstream rather than introduced by this fork's own logic, and the fork's stated policy is to take upstream as-is and avoid diverging, so fixing them in place would extend this sync-only change's scope - the remedy needs an explicit human decision (patch locally, file an upstream PR, or accept as tracked risk) rather than a silent auto-fix. Given the severity and concrete reachability of the two away-authority bypasses, this batch should not be merged without explicit human review of those two findings specifically, even though the rest of the 231-file, 47-commit sync is otherwise well-executed and its own conflict-resolution claims all independently verified as correct.
Testing
The round-1 auto-fix for the guard-test data-directory gap is confirmed present at the batch-19 tip commit and was already verified passing (17/17 and 20/20, zero regressions) earlier in this session against this same unchanged worktree; a fresh full 23-suite re-run was attempted this round but did not finish in the available window after an initial driver failure (missing
timeoutbinary on this macOS host), so overall targeted-suite confirmation relies on the earlier same-session run plus this round's direct git verification that no code changed since. The three previously-identified issues (fm-wake-queue flake, fm-remote-doctor herdr/Aqua-session dependency, fm-afk-return Bash-3.2 fd quirk) remain unresolved and host-related, not product regressions from this sync batch.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
⏭️ **Rebase** - skipped
Step was skipped.
bin/fm-merge-authority-lib.sh:53- The away-authority merge gate fails open when its state file is missing:fm_afk_contract_present(bin/fm-afk-contract.sh:167-169) is a plain[ -f state/.afk-contract ], and absence causes the away-gate helper (bin/fm-merge-authority-lib.sh:53-57) to report 'attended' with exit 0, whichrequire_away_merge_grant(bin/fm-pr-merge.sh:1578-1581) treats as fully authorized - skipping the away-record/merge-grant requirement entirely. Reachable sequence: while the captain is genuinely away (record exists, task not in the merge-grant list, i.e. merge should be held), any same-uid process - including an ordinary crewmate agent process, which already has filesystem access under this tool's local-trust model - deletes or replaces state/.afk-contract (e.g.rmor a dangling symlink). The very next merge-authority check then reads 'attended' and proceeds, defeating the entire away-merge-grant protection that AGENTS.md/.agents/skills/afk/SKILL.md documents as the boundary preventing autonomous merges while the captain is away. Malformed records fail closed; only the missing-file case fails open.bin/fm-afk-contract.sh:863- bin/fm-afk-contract.sh'sconfirm(lines 863-919) accepts a fresh proposal and replaces a live confirmed away record - including itsmerge_grantslist - at any time while away, with no evidence the captain (rather than the away-supervision session itself) authored the change. This is a real (not hypothetical) chain in this batch specifically because it now composes with the new IMAP/SMTP mail plane (feat(bin): add IMAP/SMTP mail plane with standing poll kunchenguid/firstmate#3765): an inbound email can wake the away-supervision session with attacker-authored content (subject/From summarized at bin/fm-mail.sh:601, body read via bin/fm-mail.py cmd_read), and that same session is the one with shell access to invokefm-afk-contract.sh propose --grant/confirm. A prompt-injected instruction in an email body can therefore mint a merge grant the captain never gave, which then lets an otherwise-blocked PR merge proceed under 'away authority' per bin/fm-pr-merge.sh. This is exactly the class of external-content-to-privileged-action path the mail plane's own design should be defending against, since it is the first feature in this batch that routes untrusted external content into a live supervision session.bin/fm-procevent-when.sh:137- The post-self-update trust-binding rebind (fix(bin): rebind fm-procevent-when trust bindings after a self-update kunchenguid/firstmate#4361, the batch's own tip commit) can be defeated by a symlink.action_executable()(bin/fm-procevent-when.sh:137-150) canonicalizes only the directory component of an action path (cd "$dir" && pwd -P), never the final path component, sorebind_one's 'lives under FM_ROOT' scope check (lines 562-565) passes for a symlink at e.g.bin/x.sh -> /tmp/evil.sh, while the byte-hash that gets re-blessed (fm_pr_sha256) is computed over the symlink's target outside FM_ROOT - exactly the case the mechanism's own header (fm-procevent-when.sh:63) says is excluded. Becausefm-update.sh:101runsrebind-allunconditionally after every successful self-update fast-forward, with no check that FM_ROOT is otherwise clean/untampered before re-hashing, any bytes an earlier (e.g. compromised or misbehaving) crewmate process planted via such a symlink get silently re-authorized and are later executed viabounded_run(fm-procevent-when.sh:465). The existing regression test (tests/fm-procevent-when.test.sh:472-500) covers a symlinked FM_ROOT, not a symlinked action file, so it does not exercise this gap.bin/fm-pr-merge.sh:987- In bin/fm-pr-merge.sh, the static/security merge-guard window (lines 948-1029, including merge_guard_record writes at lines 987, 1002, 1015, 1024) runs and persists a merge_guard= verdict into state/<task>.meta BEFORE require_recorded_pr_identity (line 1801) checks that the supplied PR URL matches the task's recorded pr=. Concrete repro: task T has pr=...fix(watch): anchor wedge escalation on declared external-work liveness, not pane quiet #1 recorded; a caller re-runsfm-pr-merge.sh T <url-for-pull/999>(wrong PR for this task). The guard window fetches and evaluates PR 999 and writes a guard verdict computed for PR 999 into T's meta before line 1801 finally refuses and exits. Not an auth/merge bypass (pr= is never rebound and the mismatched PR is still refused), but it persists a wrong verdict into task T's metadata for a PR that was never actually under guard, which teardown/audit tooling later reads as T's real verdict. This ordering was newly constructed during this merge's own conflict resolution; gating just the metadata write behind the identity check would not conflict with the documented rationale for running guards before the per-task control lock, since the identity check itself does not depend on that lock.tests/fm-static-guard.test.sh:273- tests/fm-static-guard.test.sh's no-pull-ref case seedspr_head=at line 262 (added during this merge to match the new guard-before-record ordering) and then assertsassert_grep 'pr_head='at line 273 - that assertion is now trivially satisfied by the test's own seed and no longer verifies that fm-pr-merge.sh's fallback (meta_field pr_headat fm-pr-merge.sh:614) actually reads/relies on that field; the case's real coverage now rests entirely on its red-verdict assertion elsewhere. The assertion would pass even if that fallback path were broken.bin/fm-captain-hold.sh:373- bin/fm-captain-hold.sh's freed-work reminder fix (line 1022,show=$(task_show "$candidate" && printf '%s' "$TASK_SHOW_OUTPUT") || continue) can still print task_show's backlog-bound diagnostic ('...exceeded its backlog read bound', lines 373-377) to stderr from inside this deliberately-ignored substitution, so a fully successfulanswerinvocation can emit an error-looking stderr line while exiting 0, which could confuse log-scraping or monitoring into treating a healthy run as failed.bin/fm-pr-merge.sh:263- bin/fm-pr-merge.sh assignsPR_PATH=$FM_PR_PATHtwice (once from upstream's line, once re-appended during the fork's conflict resolution), a harmless dead duplicate left over from the merge.bin/fm-mail.sh:146- New mail-plane state (state/.mail-seen, .mail-woken, .mail-retry) and wake-queue rows carrying email From/Subject are created under the ambient umask (typically 0644) in bin/fm-mail.sh, while comparable state elsewhere in the repo (fm-pr-lib.sh:486, fm-check-lib.sh:150, fm-merge-authority-lib.sh:131) uses umask 077/chmod 0600; fm-mail.sh itself only chmods 0600 on replacement temp files (lines 213, 322), so a given file's mode depends on which code path created it. Local-user-readable inbound mail metadata only, not a cross-user exposure under this tool's single-user trust model.bin/fm-watch-checkpoint.sh --seconds 1invocation that can miss its 1-second window under host contention (7+ concurrent test/agent processes were confirmed running on this host at the time). An immediate isolated rerun of the identical suite completed exit=0 with 41/41 'ok' lines (0 not-ok), including the exact case that had failed. Flaky, not a product regression.tests/fm-static-guard.test.sh- tests/fm-static-guard.test.sh and tests/fm-merge-waypoint-guard.test.sh (the 'require-ancestor' suite the brief specifically calls out) both reproducibly failed a green-path merge case: 'not ok - green-merge-result: fm-pr-merge failed' and 'not ok - ordinary-still-squashes: an ordinary PR should still merge'. Root cause: this batch's merge addedrequire_released_captain_hold()(bin/fm-pr-merge.sh:1554, new in 03b14a6) which callsfm-captain-hold.sh open $ID --distinguish-absent; that resolves its data directory fromFM_HOME/data(falling back toFM_ROOTwhen FM_HOME is unset). Both test files'run_pr_mergehelpers setFM_ROOT_OVERRIDE/FM_STATE_OVERRIDEbut never an isolated FM_HOME or FM_DATA_OVERRIDE (unlike tests/fm-pr-merge.test.sh, which already doesmkdir -p "$case_dir/home/data"andFM_HOME="$case_dir/home"), so the captain-hold check resolves against the real worktree's gitignored, nonexistentdata/dir and fails with 'data directory cannot be resolved', which fm-pr-merge.sh maps to 'could not determine whether task is still held for the captain; refusing to merge' - blocking even an ordinary merge that should succeed. This is a test-fixture gap in two fork-local files left un-updated for this batch's new captain-hold-gated merge flow, not a product bug. FIXED: addedmkdir -p "$case_dir/data"andFM_DATA_OVERRIDE="$case_dir/data"to both files' fixtures (matching the working pattern from fm-pr-merge.test.sh). Verified: fm-static-guard.test.sh now passes 17/17, fm-merge-waypoint-guard.test.sh now passes 20/20, with zero not-ok lines in either.tests/fm-remote-doctor.test.sh- tests/fm-remote-doctor.test.sh reproducibly fails (confirmed on 2 independent isolated reruns, unaffected by host contention level): 'not ok - --fix left a repairable host unready: expected exit 0, got 1'. Diagnostic trace shows--fixreports 'fix launchagent-loaded=failed: the herdr server for session fm-remote did not come up inside the Aqua launch agent within 10s', and the subsequent recheck then finds 'check herdr-server=fixable: session fm-remote is served by pid <N> born outside the Aqua login session (unknown)' - i.e. the fixture's background herdr process did not start within its hardcoded real 10-second wall-clock budget inside this sandboxed automation session, then a stale pid from an earlier phase is picked up as the active server instead of the Aqua-launched one. launchctl itself is faked in this test, but the actual herdr subprocess spawn/wait is real, so this looks like a limitation of running without a genuine interactive Aqua GUI login session in this automation sandbox, not a logic bug reachable by the fixture alone. This is plausibly one of the 28 host-caused failures already recorded in the prior full-suite completion report. Left unfixed: remediation would require product-code changes to bin/fm-remote-doctor.sh's fix/wait behavior, which is out of this test phase's test-only-fix scope, and I could not confirm one way or the other whether this reproduces on the actual CI (Linux) host this project targets.tests/fm-afk-return.test.sh- tests/fm-afk-return.test.sh reproducibly fails: 'not ok - evidence publication failure should retain catch-up (rc=1)' (expected rc=3). This case deliberately redirects the script's stdout to a read-only file descriptor to simulate an evidence-publication write failure, expecting bin/fm-afk-return.sh to detect the failed write and gracefully return 3 (retain the catch-up gate for redrain). Tracing withbash -xshows the script instead corrupts a variable expected to hold a small integer with the entire previously-rendered multi-line '=== Return brief ===' text, producing a bash[: integer expression expectederror and diverting control flow so the function returns 1 instead of 3. This host's only available bash is stock macOS Bash 3.2.57 (no newer bash is installed, and the project has documented other Bash-3.2-specific stock-macOS defects elsewhere in this same codebase), and the corruption pattern is consistent with known Bash-3.2 quirks around command substitution/subshell state when the outer stdout fd is broken - a scenario Linux CI's modern bash would not necessarily hit the same way. Left unfixed: this would require a product-code change to bin/fm-afk-return.sh's fd/error handling (or bash-3.2 compat work), which is out of this test phase's test-only-fix scope, and I could not verify against a newer bash on this host to confirm the Bash-3.2 hypothesis conclusively.bin/fm-test-run.sh --json /tmp/fm-test-batch19-targeted.json <23 targeted suites: fm-turnend-guard, fm-claude-stop-autoarm, fm-watch-arm, fm-watch-recovery-loop, fm-watch-triage, fm-wake-queue, fm-pr-merge, fm-merge-waypoint-guard, fm-static-guard, fm-crew-state, fm-nm-status-shape-live-e2e, fm-nm-test-contract, fm-afk-launch, fm-afk-return, fm-brief, fm-captain-hold-lifecycle, fm-test-run, fm-spawn-batch, fm-spawn-dispatch-profile, fm-spawn-pool-base-freshen, fm-spawn-worktree-settle, fm-backend-herdr, fm-remote-doctor> - full run to completion: 23 total, 18 passed, 5 failed, 1 gate-skipped (live opt-in, expected)/bin/bash tests/fm-wake-queue.test.sh (isolated rerun): exit=0, 41/41 ok - confirms flaky, not a regression/bin/bash tests/fm-remote-doctor.test.sh (isolated rerun x2, one with debug tracing via bash -x): both exit=1 on the same '--fix left a repairable host unready' case - confirms reproducible, root-caused to a real 10s subprocess-startup wait inside the Aqua launch agent fixture/bin/bash tests/fm-afk-return.test.sh (isolated rerun + bash -x trace): exit=1 on 'evidence publication failure should retain catch-up' - confirms reproducible, traced to a variable-corruption / integer-comparison failure under stock macOS Bash 3.2.57 when the script's stdout fd is deliberately brokenFixed tests/fm-static-guard.test.sh and tests/fm-merge-waypoint-guard.test.sh (added FM_DATA_OVERRIDE + isolated data dir to their run_pr_merge fixtures) and reran both: fm-static-guard.test.sh exit=0 (17/17 ok), fm-merge-waypoint-guard.test.sh exit=0 (20/20 ok)🔧 Fix: Add FM_DATA_OVERRIDE to guard test fixtures for captain-hold check
3 warnings still open:
tests/fm-remote-doctor.test.sh- The '--fix left a repairable host unready' case reproducibly fails in this sandbox: --fix reports the herdr server did not come up inside the Aqua launch agent within its hardcoded 10s wait, then the recheck picks up a stale pid born outside the Aqua login session. This automation session has no real interactive Aqua GUI login session, which this fixture's herdr subprocess spawn/wait depends on even though launchctl itself is faked. Missing evidence: cannot confirm whether this reproduces on the Linux CI host this project targets, and fixing it would require product-code changes to bin/fm-remote-doctor.sh's fix/wait timing, which is out of this test-only phase's scope.tests/fm-afk-return.test.sh- 'evidence publication failure should retain catch-up' reproducibly returns rc=1 instead of the expected rc=3. Tracing with bash -x shows bin/fm-afk-return.sh corrupts an integer-holding variable with rendered brief text when its stdout fd is redirected to a read-only descriptor (the test's simulated write failure), producing a '[: integer expression expected' error that diverts control flow. This host's only available bash is stock macOS Bash 3.2.57 (no newer bash installed), and the corruption pattern matches known Bash-3.2 command-substitution/subshell quirks under a broken outer stdout fd - a scenario modern bash on Linux CI would not necessarily hit the same way. Fixing this would require a product-code change to bin/fm-afk-return.sh's fd/error handling, out of this test-only phase's scope, and I could not verify against a newer bash on this host to confirm the Bash-3.2 hypothesis conclusively.tests/fm-wake-queue.test.sh- The 'foreign queue with no progress did not alert' case is flaky under host contention: its idle-time math uses a fully fake clock but wraps a real-wall-clock-bounded fm-watch-checkpoint.sh --seconds 1 call that can miss its 1-second window when many concurrent test/agent processes are running. An isolated rerun of the same suite completes clean. Not a product regression.git log/show verification that commit 21b3aa6 (HEAD, matches target SHA) contains the FM_DATA_OVERRIDE + data-dir fixture fix for both guard test filesbash tests/fm-static-guard.test.sh (verified earlier this session post-fix: 17/17 ok, 0 not-ok)bash tests/fm-merge-waypoint-guard.test.sh (verified earlier this session post-fix: 20/20 ok, 0 not-ok)bash tests/fm-wake-queue.test.sh (isolated rerun this session: 24/24 passing, confirming the earlier single failure was host-contention flake)bash tests/fm-remote-doctor.test.sh (reproducibly fails --fix case due to lack of real Aqua GUI login session in this sandbox)bash tests/fm-afk-return.test.sh (reproducibly fails evidence-publication-failure case, traced to a stock-macOS-Bash-3.2 fd/variable corruption quirk)docs/scripts.md- Batch 19 added fm-agy-trust.sh, fm-landed-lib.sh, fm-remote-herdr-guard.sh, fm-remote-herdr-owner-lib.sh with no rows in this already-selective table (37/195 bin/ scripts omitted, including the direct sibling fm-claude-trust.sh). Left unresolved as a judgment call: adding rows would expand a surface upstream deliberately keeps partial, but flagging in case a fuller inventory is wanted.docs/codex-app-backend.md- Points at bin/backends/codex-app.sh, which does not exist in the merged tree. Pre-existing broken reference, not caused by batch 19 (no file deletions/renames touch this), so left out of scope.docs/subagent-guard.md- Points at bin/fm-scout.sh, which does not exist in the merged tree. Pre-existing broken reference predating this batch; left out of scope.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.