Merge upstream Firstmate through 4768e98d (tail round 1) - #277
Merged
Merged
Conversation
…nguid#4037) * test(lib): set fixture mtimes through one portable epoch helper On macOS the visible symptom was ONE red case in the turn-end guard suite. The actual damage was TWO cases that had quietly stopped testing their subject. The red one was the harmless half - people read one red case as one broken thing, and here that intuition is wrong. `touch -d @<epoch>` is a GNU extension; BSD touch rejects it outright and leaves the file at its current mtime. So on macOS the three away-mode beacon cases never aged their beacon at all. The 400s case exists to pin that 400s is stale under the flat 300s default but fresh under the poll-derived grace (660s at FM_POLL=600). Deleting the grace it guards (FM_POLL=60, so max(300,120)=300) and re-running proves what it was worth on this platform: pre-fix input (beacon left at now): ok - passes with the feature DELETED post-fix input (beacon 400s old): not ok - expected exit 0, got 2 It was green while measuring nothing, and could not have caught a regression in the grace it names. Only the 700s case broke loudly. `touch -t [[CC]YY]MMDDhhmm[.SS]` is POSIX and both platforms accept it, so the only host-specific step left is formatting the epoch into that stamp, which date(1) spells two incompatible ways. fm_touch_epoch in tests/lib.sh owns that probe once and fails loudly rather than leaving an unset timestamp behind - the failure mode that caused this. Verified on BSD touch/date here and on GNU coreutils 9.7 in a container. Three real sites, and one consistency change - not four fixes. The stale destination lock in tests/fm-remote-backlog-handoff.test.sh was never a defect: its `uname = Darwin` branch made the `touch -d` line unreachable on macOS, and BSD touch accepts that space-separated form anyway. `touch -t` takes the date directly on both platforms, so the branch goes rather than standing as a second copy of the same platform assumption. Known limit: the third case (away mode off) is only HALF recovered here. It now receives the input its name claims, but it is still insensitive after this fix - its verdict is identical with a 0s and a 400s beacon, because the fixture records a daemon lock and no watcher lock, and with away mode off the daemon lock proves nothing. Not fixed here; tracked separately, with the requirement that any fix be shown to FAIL when the protection is removed. FULL SUITE ON macOS: 189 scripts, four red, none of them this change. The turn-end guard and remote-handoff suites are clean. Attribution was established by running the four failures at the base commit and at this head on an idle machine, because base-idle against head-under-load moves two variables at once: script head/loaded base/idle head/idle verdict fm-calm-pi-extension red red red pre-existing fm-backlog-atomicity red red red pre-existing fm-procevent red red red pre-existing fm-startup-network red green green cause unestablished, load-sensitive under a full run Reported, not fixed. fm-calm-pi-extension deserves its own note: it FAILS because Chrome is absent instead of declaring the capability it needs and standing aside, so its verdict is about the machine rather than its subject - the same family as the defect above, with the red at least announcing itself. Neighbouring class, reported not changed: `file_mode()` - a verbatim `uname = Darwin ? stat -f %Lp : stat -c %a` - is copy-pasted across at least five test scripts plus a `reread_mode` variant, and epoch-mtime reads are open-coded as `date -r … || stat -c %Y` in three more; same one-owner shape as the defect above. `git init` without `-b main` depends on the host's init.defaultBranch in several scripts (branch-name case, tracked elsewhere). timeout, sha256sum and sed -i uses are all correctly guarded where checked. Observed while building the check rather than the fix: the first watcher I wrote to wait for the suite matched its own command line, so it was waiting on its own existence and could never fire. Same shape as the cases above - machinery answering confidently about something other than its subject, by including itself in the evidence it was meant to judge. The file sentinel it was replaced with cannot be produced by the observer that reads it. * fix(review): Pin fixture timestamps to UTC across DST transitions * fix(document): Clarify shared fixture suite coverage
…ision (kunchenguid#4038) * fix(pi): preserve native Codex effort and guarded supervision * test(pi): identify native compatibility guard versions * no-mistakes(review): share native-main follow rule between build and picker * no-mistakes(document): document native progress marker and ultra effort owners * no-mistakes(ci): Failing check "Behavior portable serial 1" was caused by this PR. The shard ran the default-on live guard tests/fm-pi-branch-responsiveness-live-e2e.test.sh, whose idle arm loads .pi/extensions/fm-branch-supervision.ts into a scratch project with a fixed list of copied libs. This PR added `import { registerFirstmateTool } from "./lib/fm-native-contract.ts"` to that extension and updated every other loading fixture's copy list, but missed this guard. Pi 0.85.1 therefore refused to load the extension ("Cannot find module './lib/fm-native-contract.ts'"), never drew its TUI, the test failed with "Pi 0.85.1 never drew its TUI in the idle arm", and the job hit its 20-minute cap. Fix (one line): added fm-native-contract to the lib copy loop in tests/fm-pi-branch-responsiveness-live-e2e.test.sh. Swept all other suites referencing fm-branch-supervision.ts / fm-primary-pi-watch.ts; the remaining ones without the new lib only hash, path-reference, or string-match the files and do not load them into Pi, so no further fixture changes are needed. Verification: reproduced the mechanism against the installed Pi 0.85.1 by building the lab copy with the old lib list (load error as above) and the fixed list (loads cleanly). tmux is not installed on this machine, so the live guard itself gate-skips locally ("skip: live: tmux absent") and could not be run end to end here; CI (which has tmux and Pi) will exercise it. shellcheck is clean on the edited file --------- Co-authored-by: Talon Stark <talonstark@gmail.com>
…guid#4041) * fix(herdr): step around a stale client the running server refuses A remote host can carry a self-updated herdr in ~/.local/bin beside a package-managed one, and the fixed remote-job PATH resolves ~/.local/bin first. After the server upgraded to 0.9.0 (protocol 22) the stale 0.8.2 client (protocol 20) was answered with protocol_mismatch on every command, which the read classifiers folded into `unreadable`: the live remote secondmate read unknown, every doorbell into it failed, and both the spawn and relaunch recovery paths refused, so the defect trapped itself. The adapter's session-scoped CLI wrapper now recognizes that refusal, reads status per session from each distinct herdr on PATH, adopts the first one the running server reports compatible, retries once, and keeps it for the process. The happy path makes no extra call and no other failure reselects. An endpoint that still reads unreadable names the refused client, both protocols, and the fix on stderr; the remote state read, fm-crew-state, and the launch refusal carry that reason, and fm-remote-doctor reports the selected client and rebinds the launch agent to it. Regression coverage: fake two-client hosts in the herdr unit suite, the doctor suite, the crew-state remote arm, and the real host-local control script in the remote lifecycle e2e; the real-herdr smoke refreshes the status shape the selection reads. * no-mistakes(review): Reselect Herdr client after every protocol mismatch * no-mistakes(review): Remove unrequired Herdr diagnostics and launch-agent rebinding * no-mistakes(review): Scope cached Herdr clients to their selected session * no-mistakes(review): Restrict herdr client selection to reactive CLI calls * no-mistakes(document): Document session-scoped Herdr client reselection * no-mistakes(document): Clarify Herdr client selection documentation * no-mistakes(ci): Updated the trusted fm-remote-doctor.sh SHA-256 in bin/fm-remote-entrypoint.sh after the PR changed the doctor, restoring git-unavailable bootstrap authentication. Verified tests/fm-on.test.sh, tests/fm-backend-herdr.test.sh, bin/fm-lint.sh, and git diff --check all pass
* feat(afk): record the away posture and its lifecycle (phase 1) Away mode becomes a posture of the one supervision session, recorded in state/.afk-contract by the new bin/fm-afk-contract.sh: the one owner of the record schema, the mandate-clause grammar and compiler, refusal naming the missing part, the read-back rendering, the entry announcement (hold-for-return only, no phone channel), and the archive at return. This release records clauses and does not execute them; the announcement and return brief say so. bin/fm-afk-launch.sh gains propose and confirm, confirms the record before any daemon launch, refuses to launch the daemon on Pi and pi-signed, and archives the record last on stop. bin/fm-afk-return.sh snapshots supervisor health before shutdown, renders the return brief (health, mandate, waiting on the captain, could not fix, handled, cost) from the archived record, the outcome store, the held set, and the status logs, and shrinks the blocker gate to what the away session could not fix. While the record exists the watcher and the daemon never recheck an item held for the captain. Declared external waits get a four-hour default cadence and honor `until <UTC ISO 8601>` on the paused line, in both postures, bounded by FM_PAUSE_UNTIL_MAX_SECS. The /afk skill, AGENTS.md's layout and away-mode stub, the session-start digest, and the architecture, Pi branch, configuration, and scripts docs describe the record. The Pi/Herdr e2e now proves the no-daemon posture on a real Pi primary; its verification record carries the 2026-09-08 run. * no-mistakes(review): Fix AFK confirmation, grammar, waits, and return gating * no-mistakes(review): Harden AFK authority and posture lifecycle * no-mistakes(review): Preserve AFK history and tighten authority grammar * refactor(afk): record clause fields with no natural-language parser By the captain's mandate the away-posture record keeps no static parser that tries to understand natural language. A mandate clause is now given as explicit fields (--action, --object, --when, optional --stop) that bin/fm-afk-contract.sh records verbatim. The structural check asserts only that the action, object, and precondition fields are present and that the action is a listed verb; whether a precondition holds is the supervision session's judgment at execution time in a later phase. The never-set stays as a forbidden-concept safety scan: fields mentioning credentials, passwords, logins, legal or financial acceptance, payments, invoices, one-time codes, or an attended prompt are refused, matched at token prefixes after punctuation normalization so compound and plural spellings are caught. The red-check grammar, class-word rejection, unconditional-word detection, clause-reference resolution, and condition aliases are removed. --words-file keeps the captain's words verbatim, trailing newline included. The skill, docs, launcher help, and tests describe the field form. * no-mistakes(review): Preserve AFK words and tighten safety refusals * no-mistakes(review): Preserve clause bytes and honor declared waits * no-mistakes(review): Harden deny-list and gate unreadable outcomes * no-mistakes(review): Demote never-set scan and clarify authority * no-mistakes(review): Gate return on unreadable held and status data * no-mistakes(review): Validate posture archives and enforce Pi detection * fix(afk): make the never-set a non-refusing flag and keep return fail-safe Per the captain's decision the never-set scan is a coarse best-effort flag, never a refusal and never the gate: a clause naming a listed concept is still recorded with a flag the read-back, announcement, and return brief show, and the scan matches listed terms exactly or with a plain inflection at punctuation-delimited token boundaries, so unrelated names such as ping-service or tokenize-worker are never flagged and joined compounds remain a documented miss. Authoritative never-set and forbidden-action enforcement is the supervision session's judgment at execution time in phase 4. A replacement copies the superseded record through a temporary name and renames it atomically so a failed copy leaves no partial archive, the record owner gains validate and flags subcommands, and the return keeps catch-up gated when a superseded archive cannot be read. * no-mistakes(review): Harden AFK record validation and return reconciliation * no-mistakes(review): Harden AFK record validation and simplify commands * no-mistakes(review): Harden mandate validation and retain missing records * no-mistakes(review): Refuse blank explicit mandate stops * no-mistakes(review): Recover restored posture epoch before return * no-mistakes(review): Prevent return brief status symlink reads * no-mistakes(document): Refresh AFK posture documentation * no-mistakes(ci): Fixed both CI failures: updated lint telemetry for the new fourth source directive, quoted the hyphenated fixture value, and removed unreachable test cleanup. Verified with tests/fm-lint.test.sh, targeted CI-mode ShellCheck, bin/fm-lint.sh, bash syntax checks, and git diff checks * no-mistakes(ci): Bound structured pause deadlines by FM_PAUSE_RESURFACE_SECS in watcher and daemon housekeeping, added distinct bounded-horizon reasons, regression coverage for near, passed, and wrong-year deadlines, and updated documentation. Verified targeted behavior tests, full daemon tests, ShellCheck source-following lint, syntax, and diff checks
) Opt-in IMAP/SMTP mail plane (fm-mail.sh / fm-mail-check.sh). Absent FM_MAIL_* stays off. Speaking as Kun's firstmate: this is merged. Thank you @feilipu — really appreciate you taking the time on this.
* fix(remote): start the fm-remote Herdr agent through a login shell Launchd was exec-ing herdr directly, so the Aqua agent inherited a background session without login-keychain access. Start it via /bin/zsh -lc exec so panes keep login env and can refresh OAuth tokens after reboot. * fix(remote): start fm-remote Herdr via the account login shell Resolve UserShell from Directory Services and invoke it with separate -l and -c so bash, fish, and zsh all get login-keychain access. Fall back to SHELL, then /bin/zsh, then /bin/sh without failing the render. * no-mistakes(review): Fix launch-agent shell fallback resolution * no-mistakes(review): Preserve and escape Directory Services shell paths * no-mistakes(document): Document login-shell LaunchAgent behavior * no-mistakes(ci): Updated the trusted fm-remote-doctor SHA-256 identity in bin/fm-remote-entrypoint.sh. Verified with tests/fm-on.test.sh, tests/fm-remote-doctor.test.sh, bash syntax checks, and git diff --check * no-mistakes(ci): Resolved the login shell exactly once per doctor invocation and threaded it through plist rendering, installed/loaded contract validation, repair reporting, and post-repair checks. Added a regression test proving repeated repair remains healthy and performs no reload when a hypothetical second Directory Services lookup would differ. Updated the trusted doctor hash. Verified doctor, fm-on, remote-entrypoint, lint tests, ShellCheck, syntax, and diff checks * no-mistakes(ci): Made Darwin shell resolution hermetic with executable injection and a 2-second Directory Services timeout. Updated tests to inject shells by default, isolate dscl-specific cases, parse plists semantically, and verify stalled dscl fallback. Updated the trusted doctor hash. Doctor, fm-on, entrypoint, syntax, hash, and diff checks pass * no-mistakes(ci): Raised portable serial CI timeout from 20 to 30 minutes, refreshed the specified timing hints, added missing hints, and recomputed shard documentation. Verified coverage, runner behavior tests, workflow lint tests, shell syntax, requested timing maxima, and diff checks
…enguid#3710) * fix(bearings): keep captain-approved deliveries in Recently Landed A closed task is never held: tasks-axi clears the held flag when a task closes and keeps hold-kind and the hold reason as the record of the call that was made. Recently Landed excluded every Done row whose hold-kind was captain, so the marker it treated as "closed while still waiting on the captain" was in fact the proof that the captain had approved the work. Every merge routed through a captain decision disappeared from the list of what shipped, including under --all-landed. The selector now asks whether the closed row delivered something. Recently Landed is merged PRs, completed scouts, and finished local-only merges, so a row carrying one of those artifacts belongs there whoever approved it. A captain question closes with an answer and no artifact of its own, and that is what still stays out, so an answered question is never rendered as shipped work. The same rule was written twice - the bearings projection selects this home's Done rows and the fleet snapshot selects each secondmate home's Done rows into the roll-up the same section merges in - which is why one defect hid deliveries in every home. Both now share bin/fm-landed-lib.sh. * fix(review): Normalize landed evidence and exclude answered captain questions * fix(review): Normalize captain delivery evidence across relocated data * fix(review): Record authoritative delivery provenance with legacy fallback * fix(review): Harden delivery provenance across forced and pruned completions * fix(review): Replace premature merge closure with existing release contract * fix(review): Document provenance-based Recently Landed selection * fix(document): Align documentation with completion provenance * fix(lint): Fix targeted ShellCheck warnings * fix(ci): order the pinned tasks-axi install before its stock-Bash consumers In `.github/workflows/ci.yml` the pinned tasks-axi install now precedes both stock-Bash consumers, and the Bearings expectation is updated from 49 to 50 tests. Verified with macOS Bash 3.2: snapshot 16/16, Bearings 50/50, public-followup 1/1. Full repository lint and all three workflow validations pass, and `git diff --check` is clean. * fix(review): Make completion provenance unambiguous * fix(review): Make completion verdict authoritative over quoted provenance * fix(review): Preserve retained artifacts through resumed captain closes * fix(review): Unify completion provenance ordering across writer and reader * fix(review): Preserve artifacts across failed captain closes * fix(review): Refresh v1 assertions; provenance authority remains unresolved * fix(review): Remove unreliable provenance while preserving landed deliveries * fix(review): Reject stale home summaries visibly * fix(review): Restore retained deliverable recording * fix(review): Match landed artifacts and restore retention documentation * fix(review): Disambiguate captain calls and restore landed artifact matching * fix(review): Persist retained report and PR artifacts * fix(review): Preserve staged artifacts before captain answers * fix(review): Avoid wedging answers on unsupported report paths * fix(review): Exclude unreleased captain-held pull requests * fix(review): Exclude held local-only answers from landed * fix(review): Preserve retained scout reports across snapshot rendering * fix(document): Align landed lifecycle documentation with release semantics * fix(review): Enforce landed artifact-kind ownership * fix(review): Infer canonical task kinds in snapshots * fix(review): Require captain-hold release before merges * fix(review): Qualify merge lifecycle regression evidence * fix(review): Serialize captain holds with merge operations * fix(review): Document merge cleanup residuals honestly * fix(test): Replace vacuous Bearings regression with behavioral cases * fix(document): Align Bearings verification and merge lifecycle documentation * fix(review): Serialize merges and exclude captain calls from landed * fix(review): Harden merge identity and landed selection * fix(document): Clarify landed selector compatibility filtering * fix(bin): keep merge entrypoints usable on records without an incarnation The merge identity guard refused any task record with no spawn_gen field. That field identifies one exact incarnation, so comparing it across the wait for the merge lock is what catches a task relaunched while the merge was queued. Requiring it to be present is a different rule, and it refused every record written before the field existed: a legacy task could no longer be merged at all, and five behaviour suites refused before reaching the check they were written to exercise. The comparison only needs to notice a change. An absent field is now read as an empty incarnation and compared like any other value, so a record that gains, loses, or alters one is still refused, while a record that simply predates the field merges. An ambiguous or unreadable field stays an error, because a record that cannot name one incarnation cannot be compared. The missing-record message each entrypoint had before the guard is restored, so a genuinely absent record still says so in its own words. The role partition now precedes reading the record. Refusing the supervision branch is a statement about the actor, not about the task, so it cannot depend on a record the wrong actor may not have. A backlog file that does not exist meant "no longer an open captain call". For a caller that asked to tell absence apart it now means absent, so a board card whose home carries no backlog stays visible instead of being dropped as resolved. Fixture repositories pin their initial branch instead of inheriting init.defaultBranch, which resolved to main on a developer machine and master on a runner, so a fixture naming main failed only in CI. * fix(review): read local-only note from body; surface pending-close failures * fix(review): keep kindless local-only landings in Recently Landed * fix(review): bind local-only note scan to the tasks-axi note line * fix(review): Guard unavailable captain-hold authority records * fix(document): Document unreadable authority predicate outcome * fix(bin): read an absent backlog as absence, not an unreadable record The merge gate refused every task whose home carries no backlog file. A backlog that does not exist holds no captain call, so nothing can be held and the merge is safe; only a backlog that exists and cannot be read may hide a live hold. Those two states were collapsed into one refusal, which stopped merges in any home that keeps no backlog. The predicate now reports a missing backlog file as absence, alongside a row the backlog does not carry. A record that exists but cannot be read still leaves by the existing cannot-tell path, which both merge entrypoints already refuse, so the restrictive direction is unchanged. That leaves no way to reach the separate unavailable-record result, so the result and the two branches that handled it are removed rather than left describing an outcome that can no longer occur. The lifecycle documentation loses the same claim. Regressions cover both directions in each entrypoint: a home with a task record and no backlog merges, and a backlog present but unreadable refuses without reaching the forge. * fix(review): Fail closed unreadable backend configuration * fix(tests): pin the bare origin's initial branch in the remote seed fixture The fixture created its bare origin with no initial branch, so that repository's HEAD followed init.defaultBranch while the source repository pushed the branch fm_git_init_commit pins. On a host that still defaults to master the two disagreed: the bare origin's HEAD named a branch the push never created, cloning it warned that the remote HEAD referred to a nonexistent ref and checked out nothing, and the seed assertion for the cloned README failed. A machine whose default is already main paired the two by accident and hid it, which is why the fixture passed locally and failed on the runner. Pinning the bare origin to the same branch removes the dependency on the ambient default from both sides. Verified under both conditions: with init.defaultBranch set to master, and set to main, the suite passes 26 of 26. * fix(review): Fail closed unreadable user backend configuration * fix(bin): republish the home summary as v1 and record two load-bearing rules The published home-summary schema had moved to v3, which routed every secondmate home still emitting the earlier version to the stale branch: their landed rows, open decisions and holds all came back empty and their state read as unknown until each home was updated. The payload never justified that. Its field set, field order, truncations and the landed array construction are byte-identical to v1, so only which rows the selector places in landed differs, and a v1 consumer reads that the same way. Republishing as v1 removes the rollout regression and, with it, the tolerance machinery that existed only to soften the bump: the stale-schema predicate, its two collection branches, the flag and its provenance branch, the omitted surface that can no longer be reached, and the fixtures and assertions that covered them. Two rules that a scope review proposed removing are kept, each now carrying the reason it exists, because both were measured to be load-bearing: The artifact-kind ownership clause is what keeps an explicit scout that recorded no report out of Recently Landed. Without it such a row has none of the three artifacts, satisfies the compatibility fallback and renders as shipped work with an empty artifact. The kind fallback is needed because tasks-axi omits the kind metadata entirely when a title begins with a canonical keyword. Without it a scout titled "SCOUT ..." reports no kind, its recorded report stops counting as a delivery, and it drops out of the section this selector exists to repair. * fix(bin): move the scout guard note onto the rule and drop two dead pieces The LOAD-BEARING note sat on an unreachable branch. Measured in both directions: removing that branch together with the kind-is-not-scout guards lets an explicit reportless scout into Recently Landed and fails tests/fm-captain-hold-lifecycle.test.sh, while removing the branch alone leaves that suite passing at 49 assertions. The guards carry the rule, so the note now sits on them and the unreachable branch is gone. A note pointing a later reader at the wrong line is the hazard this change corrects elsewhere. summary_file_has_schema lost its only caller when the stale-schema machinery was removed, so it goes with it. * fix(review): Fix legacy report artifacts and canonical keyword boundaries * fix(review): Update pinned Bearings test count to 56 * fix(document): Clarify landed summary compatibility documentation * fix(review): Preserve unreadable backend configuration errors * fix(review): Honor backend resolution errors at existing call sites * fix(test): Stabilize remote collector tests under host load * fix(document): Document backend resolution failure contracts * fix(lint): Suppress intentional deferred probe expansion warnings * fix(ci): Captain, quoted the two literal test IDs in tests/fm-backlog-atomicity.test.sh to fix SC2100 without changing behavior. Both warnings reproduced before the fix; the targeted fm-lint.sh run now passes with ShellCheck 0.11.0. Bash syntax and git diff --check also pass * fix(ci): Fixed the resolver’s two configuration-parent checks to return 2 for inaccessible directories while preserving genuine absence. Added two behavioral tests; RED/GREEN and both requested mutation proofs confirmed. All 10 focused checks, targeted lint, syntax, and whitespace checks passed. Broader merge suite stopped after 10 passing cases under host load. Declined portable checks and merge-authority code remain unchanged
…nguid#4090) * fix(remote): let the Aqua launch agent own the fm-remote Herdr session A herdr server keeps the macOS audit session of whatever started it, and only the Aqua login session (gui/<uid>) can read the login keychain without a prompt. Herdr's SSH remote attach starts the fm-remote server as its own child when it finds none, wins the socket at boot because sshd accepts connections before the login session exists, and every claude pane under that server then gets `security` exit 36, falls back to a stale plaintext credentials file, and reports "Login expired". launchd's own job lost the socket on every KeepAlive retry and the doctor still reported the session ready because it only asked whether any server answered. - Add bin/fm-remote-herdr-guard.sh, the launch agent's exec target: start the server in the foreground when nothing owns the socket, exit 0 when an Aqua-born server does, and otherwise stop the foreign server, wait for the socket, and exec the server at once. - Add bin/fm-remote-herdr-owner-lib.sh, the single owner of socket-owner discovery (lsof; pgrep cannot see herdr's argv on macOS) and the birth markers (SSH_*, XPC_SERVICE_NAME, FM_REMOTE_JOB_ACTIVE, sshd or remote-client-bridge ancestry matched on argv[0] and whole arguments). - Render the agent as the login shell exec'ing the guard with KeepAlive={SuccessfulExit=false} and ThrottleInterval=10, check the loaded job's successful-exit semaphore, and report a session served outside the Aqua login session as fixable so --fix retakes it through launchd; the reload waits for an Aqua-born owner rather than any running server. - Correct the doctor and docs: the launch shell provides environment parity, the launchd domain provides keychain access. - Pin the guard's decision table and the doctor's verdicts against real marker-carrying processes, and record the dated audit-session evidence. * no-mistakes(review): Verify Aqua ownership through launchd domains * no-mistakes(document): Document macOS lsof ownership requirement
…uid#2881) * fix(bin): prefer a live no-mistakes run over a terminal one A worktree can bind to more than one recorded no-mistakes run at once. The branch-and-code-identity rule in bin/fm-nm-run-lib.sh accepts both an exact-equal commit and a worktree-is-an-ancestor match, but never stated which wins when both bind, so the tie fell to whichever candidate the caller reached first. Observed on a live fleet: a crashed validation daemon left a FAILED run at the worktree's own commit while the live run that replaced it validated a descendant commit on the same branch. Bare `axi status` answers with the most-recently-touched run - the corpse - and it bound by the equal-commit rule, so every recomputation reported `failed` for a task whose real run was healthy. The same label had also read `failed` earlier while the work was genuinely stalled, so the signal was wrong in both directions. State the live-over-terminal policy in the matching rule's own contract, where the equal-commit and ancestor rules already live, and add fm_nm_run_status_class as the one classifier that decides liveness from a recorded status word. fm-crew-state.sh applies it on both selection paths: the runs listing now scans past a terminal row for a live one, and a terminal `axi status` answer is provisional until the listing has been asked whether this worktree also has a live run. Same-liveness-class candidates keep the listing's newest-first precedence, and a status word the classifier cannot place keeps the caller's own ordering rather than displacing a known result, so a single-run task and a task whose runs are all terminal are unchanged. Regression coverage reproduces the proven case (terminal run at the worktree's exact commit plus a live run descending from it) and its runs-list twin; both fail under the old tie-break. Two companion cases pin the no-widening half - two terminal rows still resolve newest-first, and a terminal run with no live sibling keeps its full run-step detail - and both pass before and after the change. * no-mistakes(review): accept unfetched live sibling anchored at exact worktree head * docs(bin): name both ledger reads behind the runs-limit setting The FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh still described the runs ledger as scanned only by the cross-branch fallback, but the live-over-terminal fix also consults it as the live-sibling probe behind a terminal axi status answer. Point the comment at docs/configuration.md as the setting's owner instead of restating a second copy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…chenguid#4009) * fix(procevent): make the ordinary stop signal actually stop a runner The owner guard that shipped in kunchenguid#3904 half-reaps. Against a poll child that handles the ordinary stop signal and keeps waiting, the guard signals the group, loses the runner leader to its own signal, then reads that success as a leaderless group and exits without escalating. It destroys the only proof of ownership that would have authorised the forced signal, so the survivor becomes unreachable by retire, reconcile, sweep-home and the guard alike. A guard that turns a leaking-but-identifiable generation into a permanently unreachable one is worse than no guard at all. Two defects, and they hid each other: - The escalation re-derived ownership from the leader. `runner_group_signal` now takes a `proved` mode, passed only by the escalation inside the stop that already proved and signalled that exact generation moments earlier. A leader dying to our own signal is the ordinary outcome, not fresh ambiguity. - Every stop held the per-source lock across its wait while the runner's own exit cleanup waited unboundedly for that same lock. That circular wait was broken only by the forced signal, so the forced signal silently became the normal path - and, by keeping the leader alive through the whole window, it masked the escalation defect above. The runner's exit cleanup now refuses that lock instead of waiting for it, which is what its existing `return 0` already said it did. Fixing the lock alone would have turned every stop of a signal-proof child into a refusal that leaves it running, so both land together and the tests pin that. Measured on macOS with a stand-in poll child that traps TERM, INT and HUP: the guard left it running past 70s and now clears the group within the lease plus one check; retiring a healthy runner fell from ~2.8s with a forced group signal every time to ~0.6s on the ordinary signal alone. Unchanged and stated deliberately: a leader lost to anything other than the stop's own signal still leaves a group that retire, reconcile, sweep-home and the guard all refuse, permanently - and that source stops listening without saying so. Whether such a group may ever be signalled is an open decision and is not answered here. * fix(review): Fix proved escalation race and stop regression assertions * fix(review): Preserve proved escalation through transient identity failures * fix(review): Simplify proved escalation and correct guard timing documentation * fix(document): Clarify process-event stop ownership and cleanup limits * fix(document): Clarify process-event stop ownership and fixture comments * revert(skills): restore the leaderless-ambiguity limit to the loaded skill An automatic documentation step in this branch's validation edited .agents/skills/process-event-sources/SKILL.md, which no instruction in this change asked it to touch. That file is not documentation about the code: it is the agent-loaded instruction surface, what an agent reads to know what it is permitted to do. The step deleted this line: - leaderless PID/PGID-reuse ambiguity preserves the claim without signalling or replacement, as owned by the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); and folded it, with its neighbour, into a generic "registration and ownership transitions, stop authority, and claim reclamation follow the operating contract". That deleted line states a PROHIBITION - that such a group is preserved WITHOUT SIGNALLING - and it is the exact limit an open captain decision currently rests on. Folded into a pointer, an agent reading the skill to learn what it may do would have to chase a second document to discover it may not signal. A prohibition that requires a second lookup is not a prohibition. The effect was to weaken, in the instructions themselves, the boundary that keeps one home from signalling another's process group - while the question of whether that boundary should move at all is still open. This is a deliberate revert, not an oversight, and it restores the file exactly to its pre-branch state. The full statement also survives in docs/configuration.md; that does not rescue it, because the agent handling a process-event wake loads the skill and not the documentation. * revert(procevent): restore the open-question marking beside the escalation The same automatic documentation step that edited the loaded skill also removed this from the comment above runner_group_signal: A leaderless group nobody in this call ever proved remains refused too, for every caller. That untouched refusal is what makes a crashed leader's group permanent, and relaxing it is a separate open question, not something this path assumes. and replaced it with a pointer to docs/configuration.md. This one fails differently from the skill deletion, which is why it is restored separately. There, a prohibition was moved out of the reader's path, and a missing prohibition gets violated. Here the prohibition survives in code - the unproved path still refuses - and what was removed is the fact that the limit is UNDECIDED. A prohibition that has quietly lost its "this is still open" reads as settled design, and settled design gets relied on, extended, and eventually relaxed by someone confident they understand why it is there. That question is open right now. The rule this branch's four instances produce, stated once here because this is the point of decision: an unresolved question must be marked unresolved AT THE POINT OF DECISION, not only where the contract is documented. A reader who does not know something is open will treat it as closed, and that default is stronger than any pointer overcomes. The pointer added by that step is kept alongside; this restores what it replaced rather than reverting it. * docs(verification): restore the measured guard bound and its reason The document step's rewrite of this record dropped the concrete figure while keeping the surrounding measurements. What went missing was the bound itself - lease plus two consecutive failed checks plus the stop's grace, roughly 630 seconds at the shipped 600-second lease and 15-second check - together with the reason there are two checks rather than one: a single unreadable read must not be enough to kill a live runner. The mechanism survived elsewhere and the reason survived in docs/configuration.md, so nothing was lost from the repository. The concreteness was, and that is what this restores. A number recorded without why it is that number is the one a later reader shortens; the reason is the whole safety argument for the debounce, and the debounce is what stops the reaper killing a live runner on one bad read. * fix(document): Replace stale stop-authority summaries with owner pointers * test(procevent): make the guard-bound case able to fail for its own reason An automated reviewer observed that this case allowed sixty seconds for a bound of roughly eight, so it could not go red for the reason it names: it would have passed a guard that took fifty-five seconds. That is correct, and it is the same family as the defect the case exists to defend against - a check that is green because it cannot fail, rather than because the thing it guards is working. The deadline is now derived from the bound itself - the lease, plus the two consecutive failed checks the guard debounces on, plus the stop's own ordinary and forced signal windows - rather than from a flat wall-clock number, and the shortened lease and check the fixtures run under have a single definition so a derived deadline cannot silently diverge from the settings the guard is given. The doubling that remains is a load allowance and is documented as one; widening it to make a slow guard pass would convert the assertion back into decoration. Proven by mutation rather than by argument. Against the repaired case: correct code ok guard debounces on 20 misses instead of 2 not ok - "still holding the group after 16s, against a documented bound of 8s" proved escalation removed (the original defect) not ok - same code restored ok The previous sixty-second version passes every one of those mutations. The reviewer's other claim, that the guard can survive past the announced bound when an owner disappears immediately after a check, was measured and does not hold against what this branch announces. Sweeping the phase deliberately at 0.0, 0.2, 0.4, 0.6 and 0.8 of a check interval gave 7.21s, 7.31s, 6.75s, 6.49s and 6.31s, worst 7.31s, against the announced lease plus two consecutive failed checks plus stop grace, which is up to 8s at those settings. The mechanism the reviewer describes is real and is the announced mechanism; the bound it was measured against is a phrasing this branch no longer carries. * revert(scope): return the instruction surfaces to their base state This delivery is being split. It carries the two proven process fixes alone; the instruction text travels separately, through a run that removes the documentation step rather than refusing it at its gate. Two surfaces are therefore returned to exactly what the base branch has, so this delivery neither adds to them nor removes from them: .agents/skills/process-event-sources/SKILL.md - identical to base again. Three bullets an automatic documentation step had folded into a pointer, including that leaderless PID/PGID-reuse ambiguity preserves the claim WITHOUT SIGNALLING and that there is one identity-matched owner per canonical source across homes sharing one store. The header comment block of bin/fm-procevent.sh, which is what the script prints as its own help. Seven lines were removed from it: that a live owner is never displaced, that only a claim whose stale owner and independently absent process group prove its whole generation gone is reclaimed, that a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, and that reconcile signals only a live identity-matched runner group and otherwise keeps the claim without starting a replacement. The help output is now byte-identical to base. Neither removal was requested by any instruction in this change, and both were made to text that predates it. Returning them is scoping, not a third restoration: nothing is being added to those files here. * fix(ci): Captain, live CI revealed a fixture deadlock: it suspended the runner before startup released its lock. Added a public-list synchronization barrier in tests/fm-procevent.test.sh. Forced-delay reproduction detected the deadlock before the fix; all four cases passed afterward. Targeted lint, Bash syntax, and whitespace checks passed. Greptile’s watchdog requirement conflicts with the recorded R2 decision; runtime behavior and documentation remain unchanged. Full CI rerun belongs to the outer executor * test(procevent): make the post-TERM cases report what they saw when they fail On the failure path only, these cases now print what they actually saw: the identity recorded at claim time, the identity readable at that moment, the size of the signals file, the leader's state and wchan, every live member of the runner's process group with its own state and wchan, the elapsed time since the stop began, and what retire said. None of it runs when a case passes. WHY THIS IS KEPT, stated accurately rather than by its original reason. It was written to make an unexplained CI failure verifiable. That failure is now explained - it was a fixture deadlock, diagnosed and repaired in the preceding commit - so that justification has expired and is not the reason given here. The reason it stays is smaller and independent of that failure: it is already written, it is small, it sits in the file whose assertion this change reworked, and an assertion that could not say why it failed cost most of a morning to diagnose from the outside. The next failure will not be this one. WHAT A PASSING RUN WOULD NOT MEAN: a pass is a sample of behaviour already observed many times, not proof that anything is fixed. Only a failure carrying the evidence above establishes a cause. * fix(document): Clarify process-event fixture diagnostic rationale * fix(ci): Captain, fixed two cleanup races in tests/fm-procevent.test.sh: removed premature child completion and waited for runner exit before retiring the restart fixture. Controlled Linux reproductions demonstrated failure before and success after. The full Linux process-event suite, six focused macOS checks, targeted ShellCheck, Bash syntax, and whitespace checks passed. Runtime behavior, guard debounce, and documentation remain unchanged. CI rerun belongs to the outer executor * fix(procevent): bound owner-guard cleanup at one check interval, not two A THIRD WAY, not a capitulation to the reviewer and not a refusal of it. The automated reviewer's grievance was the LOOSENESS OF THE BOUND, not the number of observations the guard makes before it acts. It asked for a single read because that was the only route it could see to an acceptable bound. There was another route, and this change takes it: the bound is reached and both reads are kept. TIGHTENED - the SPACING of the guard's two reads, not their number. The owner watchdog now sleeps half the configured check interval and still requires two consecutive failing reads, so the pair completes inside one check interval instead of costing two. Worst-case detection falls from the lease term plus TWO check intervals to the lease term plus ONE. At the shipped 600s lease and 15s interval the stated bound falls from ~635s to ~620s. PRESERVED - the second read. bin/fm-procevent.sh's two-consecutive-miss rule is untouched. WHY IT PROTECTS: the guard's inputs are a lease read and a state-root identity read, and either can fail transiently on a live, healthy home. Acting on the first failure would let one isolated unreadable read kill a live service. Requiring a second, independent read is what makes that impossible, and it is a protection rather than padding. Nothing was traded away to reach the bound. Both properties are now guarded by their own case, and each was proven by MUTATION rather than asserted: - putting a full interval back between the two reads fails the bound case: "still running 17.0s after the last owner activity, against a documented bound of 15s"; - acting on one failed read fails the new debounce case: "one unreadable lease read ended a runner whose home was still alive" - while the bound case then passes FASTER, 9.9s against 13.1s. The unsafe variant being the quicker one is exactly why these are two cases: one elapsed-time case would have registered the removal of the protection as an improvement. MEASURED, sampling the phase between the guard's check clock and the lease clock across eight runs per variant, on macOS (Darwin 25.5.0). Reaping an orphaned listener whose home stopped refreshing its lease: lease 2s / interval 1s: 4.41-5.29s before, 3.48-4.65s after lease 2s / interval 4s: 7.69-8.12s before, 5.94-6.13s after The 4s configuration is the informative one: the gap is about one check interval, which is precisely the term that was removed. A previously unstated term of the bound surfaced while measuring: the lease age is compared in whole seconds, so a configured lease of N is honoured until that age reads N+1. It is now part of the documented bound and of the regression's derivation instead of being absorbed into a fudge factor. The bound regression derives its deadline from the documented bound instead of a flat number, and PINS the phase between the guard's check clock and the lease clock rather than sampling it, because with a sampled phase a guard spending two intervals passes about half the time on a lucky alignment. Its load slack is additive and stays under half a check interval, so an extra whole interval cannot hide inside it. The two flat deadlines that were there before (40s and 20s) and the doubling allowance on the derived one are gone; that looseness was the reviewer's third complaint. The stop's own grace is untouched: 2s for the ordinary signal, then 2s for the forced one. It is a ceiling paid only by a group that outlives the signal it was sent, not a delay every stop pays - a healthy runner's whole retire measures 0.40-0.66s on this host. The reviewer's literal "lease plus one tick" is unreachable by any implementation, since signalling a process and giving it any chance to exit takes non-zero time; detection now meets it and the stop runs inside its own ceiling, and the contract says so rather than glossing it. NECESSARY BUT NOT SUFFICIENT, and written BEFORE this head's integration runs start rather than after they report. On the previous head, "Behavior portable serial 1" and "Behavior portable serial 4" were both CANCELLED at the job ceiling, independently of this finding. A new head triggers fresh runs, so those two lanes MAY complete this time. IF THEY DO, THAT IS NOT EVIDENCE THE CEILING DEFECT IS FIXED. It is one more sample of a lane that has been cut repeatedly and sometimes is not; the shard-packing repair for it is open separately. Do not reread a lucky pass here as a resolution. Relatedly, and deliberately: the per-script duration hint in bin/fm-test-run.sh was NOT updated even though the two new cases add ~19s of wall clock. docs/fm-test-portable-shards.md says those hints are replaced wholesale from CI timing artifacts of green runs, and that repair is the open request doing it; a hand-edited estimate here would collide with it and silently repack the shards. This suite runs in portable serial shard 3, which was green in the last run. Verification: tests/fm-procevent.test.sh green, plus tests/fm-captain-hold-lifecycle.test.sh, the test-coverage guard, and bin/fm-lint.sh. The unrelated "reconcile stops a runner whose registration was removed" case flaked in 4 of 7 local full runs; an isolated 20-trial reproduction measured it at 13/20 unclean before this change and 11/20 after, so it is issue 4080 and is not aggravated here. * fix(procevent): repair our decimal-interval regression and enforce the timing phase REPAIRED BEFORE PUBLICATION, AND IT WAS OURS. The half-interval arithmetic added by the previous commit read a zero-prefixed interval as octal: 010 halved to 4 instead of 5, and 08 was not a number at all, so the owner guard died before reporting ready and the runner failed closed and never listened. The validator accepts those values and `[` compares them as decimal, so this broke a configuration that worked before. Introduced by this delivery, found in review, repaired here. Forcing base ten before the arithmetic is the whole runtime fix. Proven by driving it rather than by reading the source: a new case starts a real listener at 08 and at 010 and observes the guard's actual sleep argument - 4s and 5s. Removing the normalisation turns that case red with "a zero-prefixed decimal interval (08) prevented the listener from starting". THE TIMING PHASE IS NOW OBSERVED AND ENFORCED, NOT ASSUMED. The bound case pinned its phase by CONSTRUCTION, from an assumed startup time, and enforced nothing. Review was right that this is not enough: once startup reaches about two seconds the expiry lands in a different part of the interval and the case silently stops rejecting a two-interval guard while still reporting success. A bound that cannot fail for the reason it names is the defect this whole delivery exists to correct, so it must not ship inside the fix for it. Now the lease is synchronised to the guard's own FIRST observed lease read, every later real read is recorded, and the case REFUSES unless one recorded read proves the required phase: it read the synchronised reference, it was still fresh, and it began late enough that two further full intervals could not finish before the deadline. An unestablished precondition refuses; it does not proceed on trust. The derived deadline, the two-read debounce and the additive slack are unchanged, and the slack invariant is now asserted rather than left to a comment. Review also found the deadline was only ever checked while the group was still alive, so a sampler descheduled past it would see the group gone and certify success. The observed completion time is now checked too. PROVEN BY MUTATION, each one run against this code: - remove the decimal normalisation -> the interval case fails on 08; - a full interval between the two reads -> "the guard exceeded its bound: group still running 17.1s ... against a documented bound of 15s"; - a full interval WITH startup forced to ~2.5s, which is exactly the condition the old construction pin could not survive -> still red, same message; - the same ~2.5s startup with the correct guard -> still passes, 13.0s against the 15s bound, so the delay alone does not break the case; - phase evidence made unavailable -> "could not establish the required pre-expiry guard-read phase", a refusal rather than a pass, even though the group stopped quickly; - act on one failed read -> the debounce case fails and the bound case passes FASTER, 9.5s against 12.7s, which is why these remain separate cases. Verification: full tests/fm-procevent.test.sh green, and bin/fm-lint.sh clean. * fix(document): Correct process-event timing and debounce comments
…kunchenguid#3417) * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * fix(backlog): honor configured task adapters * no-mistakes(review): Harden backend purity lint against prefixed Beads calls * no-mistakes(document): Document configured backend lifecycle transitions * fix(backlog): preserve markdown exemptions * no-mistakes(review): Enforce backend purity for explicit lint paths * no-mistakes(document): Update lifecycle backend documentation * no-mistakes(lint): Remove redundant backend lint pattern * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(document): Document environment-selected backlog adapters * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * no-mistakes(review): Reject partially quoted direct Beads commands * no-mistakes(document): Align lifecycle documentation with configured adapters * test(backlog): keep structural cases markdown-only * fix(lint): catch dollar-quoted beads commands * no-mistakes(review): Drop unused authorized-data-dir param; canonicalize fm-lint ROOT with pwd -P * no-mistakes(document): Document fm-lint backend-purity check in header and CONTRIBUTING * no-mistakes(ci): Two failing checks, one code-caused and one attestation-only. 1. Greptile Review (P1: ANSI-C quoting bypasses lint) — REAL DEFECT, FIXED. The backend-purity normalizer in bin/fm-lint.sh (invokes_bd awk function) stripped $'...' quote delimiters without decoding ANSI-C escapes, so a core script containing $'\x62\x64' close fm-example executes as `bd close` while the lint accepted it. Fix: the normalizer now tracks whether a single-quote context came from $' (ansi flag) and decodes ANSI-C escapes while building the command word: \xHH (1-2 hex), \uHHHH / \UHHHHHHHH (4/8 hex), \NNN (1-3 octal), \cX control characters, simple escapes (a/b/e/E/f/n/r/t/v decode to a placeholder that can never spell bd), NUL (value 0) truncates the word per bash C-string semantics, and unknown escapes drop the backslash per bash. Values outside printable ASCII decode to a placeholder so they can neither falsely match nor collide into `bd`. Regression coverage added to the existing lint-interface test test_rejects_direct_beads_cli_invocations in tests/fm-lint.test.sh: $'\x62\x64', $'\142\144' (octal), and b$'\x64' (split word), each written into a fixture and asserted rejected through the real fm-lint.sh executable. Verified regression property: with the fix stashed, the new hex case passes lint (reproduces the reported bypass); with the fix, it is rejected. Verification: all 29 tests in tests/fm-lint.test.sh pass, and the full CI-parity lint (CI=true bin/fm-lint.sh: ShellCheck 0.11.0 full analysis over the canonical set, backend-purity check, and actionlint workflow validation) exits 0. One iteration was needed because an awk comment containing a literal $'\x62\x64' sample broke shell-level quoting (bash -n / SC1001/SC2026); the comment was reworded without quotes. 2. PR must be raised via no-mistakes — NOT code-caused. The check failed with 'Pipeline attestation head_sha does not match the current PR head': the PR body attestation binds to 82b41c7 while the PR head is 1995cdb because a later pipeline push moved the head. This is exactly the stale-attestation condition the user intent describes; it clears when the outer pipeline re-runs 'git push no-mistakes' and re-binds the attestation to the new head (which now includes this Greptile fix). No code change can or should address it. Files changed: bin/fm-lint.sh (ANSI-C escape decoding in the backend-purity normalizer), tests/fm-lint.test.sh (three encoded-bd rejection cases) * fix(bin): preserve captain calls during teardown (kunchenguid#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(backlog): honor configured task adapters * no-mistakes(document): Update lifecycle backend documentation * fix(backlog): close adapter routing gaps * no-mistakes(review): Honor configured markdown paths and path-qualified Beads lint * no-mistakes(lint): Fix empty local variable assignment * fix(backlog): close quoted path gaps * test(backlog): keep structural cases markdown-only * no-mistakes(review): Harden markdown lifecycle routing and close recovery * no-mistakes(review): fix tasks.toml hang, root authorization, lint quoting * no-mistakes(review): validate tasks config before exemption; fix lint quote gap * fix(backlog): address the markdown backlog as <data>/backlog.md Resolving the markdown backlog through a configured `[markdown] path` was scope this task never asked for. It is absent from main, which addresses `<data>/backlog.md` everywhere, and it came from an earlier review round rather than the task brief. Making it effective on the transition path alone put that path at odds with every other consumer of the same backlog - fm-captain-hold.sh, fm-session-start.sh, fm-fleet-snapshot.sh, fm-inbox.sh, fm-backlog-handoff.sh - which all still address `<data>/backlog.md`. In fm-captain-hold.sh the split was live: its reads had already moved to the shared gate while its writes had not, so the two could address different files. Address `<data>/backlog.md` from the shared gate, delete the unused resolver, and drop the two tests that pinned the withdrawn behaviour. What this task actually changes is unaffected: a configured non-markdown adapter is still addressed by its own root, without `--file`. * no-mistakes(review): restore home boundary guard and tighten purity lint * no-mistakes(review): authorize home boundary for every backlog adapter * no-mistakes(test): complete tasks-axi stubs in fm-gotmp teardown fixtures * no-mistakes(document): align backlog transition docs with adapter-neutral addressing * no-mistakes(review): label adapter data-dir authorization, drop dead row_probe local * no-mistakes(review): pin markdown backend at relocated-data addressing roots * no-mistakes(document): point lint-definition mention at fm-lint.sh header * no-mistakes(document): point mutate comment at adapter addressing owner * no-mistakes(review): Fix leftover-symlink refusal on non-markdown homes; hoist config check and lint/dedup cleanups * no-mistakes(document): Align fm-lint purity scope header with bin/backends --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…guid#4119) * fix(bin): stop false watcher-down alarms on long Claude turns A healthy Stop auto-arm rewake or open claim already explains a mid-turn beacon that has aged past grace, because turn-end will re-arm. Keep the supervision-off banner for a missing, failed, or exhausted generation. * no-mistakes(review): Bind Claude rewakes to active recovery generation * no-mistakes(document): Document Claude long-turn supervision exception * no-mistakes(lint): Fix empty ShellCheck assignment * no-mistakes(ci): Fixed all reported CI failures: quoted the hyphenated recovery-delivery value to satisfy ShellCheck SC2100, and updated the session-lock auto-arm fixture to emit the recovery marker and watcher beacon now required for a valid rewake. Verified fm-session-lock-ancestry, fm-test-run, stale-banner, Claude auto-arm, targeted lint, ShellCheck, and workflow lint checks pass
* fix(herdr): classify a gone session's endpoint as recoverable A task whose Herdr endpoint could not be read was classified `unreadable`, which blocks recovery by design. The commonest reason that read fails is that the recorded session's server is not running at all - a host reboot, a server exit, a session never restored - and that is authoritative absence for every pane in that session, not an ambiguous answer about one of them. Tasks in that state had no sanctioned way back. The recovery-grade read now settles an uninterpretable pane read with the session server's own `.server.running` state: positively stopped reads `missing`, while a running server, or a server state that cannot itself be read, still reads `unreadable`. Resting the verdict on that field rather than on the `server_not_running` error code is what keeps it working across Herdr 0.8.x and 0.9.0, since the field is present on both and the code is not. Only that one boundary is widened. The husk classifier under it stays strict, so duplicate prevention, rollback, and teardown - the paths that can destroy something - keep refusing on exactly the reads they refused on before. Separately, a relaunch refused outright when the endpoint's shell had drifted out of the recorded worktree. An agent's own exit routinely leaves its shell somewhere else, so that refusal stranded tasks whose work was sitting untouched on disk. The shell is now told once to return, and only a shell that will not go refuses; the replacement still never starts outside the copy holding the work. Herdr 0.8.x is not installed on this host, so protocol-20 coverage is structural plus the adapter fixture exercising both response shapes, and is recorded as such rather than as a live result. Fixes kunchenguid#4091. * no-mistakes(review): Restrict drift recovery to Herdr endpoints * no-mistakes(review): Correct Herdr recovery verification coverage * no-mistakes(document): Document Herdr endpoint recovery boundaries
kunchenguid#4033) * fix(bin): keep an escalated undelivered handoff wake retryable A remote backlog handoff holds its outbox until the backlog receipt and the receiver wake are both confirmed, and retries the wake under the same pending-reply correlation on every resume. When that wake's remote transport was lost, the correlation stayed undelivered in delivery_unknown and the watcher's next pending-reply tick escalated it. Both the reuse predicate and the known-undelivered reset refused an escalated record, so the resume refused to resend the wake forever and every later handoff to that mate jammed behind the outbox. Treat an escalated record with no confirmed delivery as the undelivered correlation it is: fm_pending_reply_corr_reusable accepts it for its own task and fm_pending_reply_reset_known_undelivered returns it to awaiting_report for the idempotent remote resend, while a delivered record is still never reset and a missed-report escalation keeps its meaning. The published delivery-unknown decision stays open until the record resolves, so a repeat loss neither re-notifies nor strands it. Reproduce the deadlock end to end in the remote handoff test (lost wake transport, watcher escalation, resume) and pin the predicate contract in the pending-reply suite; the fm-send fixture that pinned the refusal now uses a genuinely stale delivered escalation. * no-mistakes(review): Decouple durable outboxes from best-effort wake retries * no-mistakes(review): Align handoff documentation with durable receipt release policy * no-mistakes(review): Handle unrecordable wake state as dropped * no-mistakes(review): Prevent stale wake markers blocking handoffs * no-mistakes(review): Prevent stale delivered markers suppressing new wakes * no-mistakes(document): Clarify retry escalation decision lifecycle * no-mistakes(document): Document pending receiver wake retries
* docs: bound the mandatory captain address to the chat channel AGENTS.md's opening address rule said "address the user as captain at least once in every response" and never said what a response is. The artefact exclusion two lines below governed only the optional nautical seasoning, not the mandatory address. An agent that reads this file without being the first mate - a pipeline corrector agent running inside a copy of this repo - therefore read the obligation as applying everywhere and the exclusion as applying only to flavour, and opened its delivery message with "Captain,". That reading was correct. Patch the existing owner rather than adding a rule elsewhere: - bound the obligation to chat messages sent to the captain; - state the artefact exclusion once, explicitly binding every agent that reads this file whether or not it is the first mate, and naming commit messages, PR and issue descriptions, briefs, code and comments; - fold the seasoning under the same bound instead of carrying a second, narrower copy of the exclusion. The obligation itself is unchanged: the captain is still addressed in every chat message. AGENTS.md goes from 603 to 602 lines: the redundant "never send a response with zero direct address" clause and the duplicated seasoning exclusion pay for the new bound. The two cross-references that paraphrased the unbounded wording (bin/fm-parent-channel-lib.sh's header and docs/secondmate-parent-channel.md's problem statement) now match the owner; neither restates the rule. * fix(review): Limit address exclusions to artifacts while preserving public replies * fix(document): Consolidate captain address guidance
…nguid#4131) * fix(herdr): close persisted-focused tabs when no live client is attached The teardown active-tab guard treated Herdr's last-focused pointer as a live viewer, so detached sessions could not close panes on that tab. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): allow detached seeded-tab prune after live-client gate Projection create still restored the persisted focused tab after a successful prune, so a detached last-focused seeded tab still quarantined the spawn. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(herdr): probe live client after seeded prune only when that tab was focused The extra title-clear read after every prune shifted canned CLI fixtures and failed projection create. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Tighten Herdr active-tab close guard * no-mistakes(review): Guard Herdr mutations with fresh target focus * no-mistakes(document): Document Herdr live-viewer teardown guard --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…guid#4169) * fix(bin): escalate decision-owned wakes once as the decision The away-mode daemon treated a needs-decision: queued payload as an unknown wake, so suppression markers never committed and the same open decision re-escalated on every poll. Classify that payload through the existing signal path so it escalates once, labelled as the decision, and an unchanged repeat is suppressed on the same terms as any other signal. Fixes kunchenguid#4096 * no-mistakes(review): Escalate captain-held decision-owned rows once as the decision * no-mistakes(review): Self-handle captain-held decision-owned rows instead of escalating them * no-mistakes(document): Name away daemon as needs-decision payload reader
* test(herdr): pin leftover-shell vs live-idle via agent get Herdr 0.9.0 already distinguishes a Pi that exits to a surviving pane shell from a sibling live idle occupant. Pin that pair through agent get and the recovery classifier so a lagged pane-get status cannot silently reclaim the leftover shell as alive. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Document Herdr leftover-shell liveness regression * no-mistakes(ci): Fixed Lint failure SC2034 by replacing the unused wait-loop variable with `_`. Verified with the pinned project lint command, Bash syntax check, and git diff check --------- Co-authored-by: Cursor <cursoragent@cursor.com>
Merge the exact contiguous upstream prefix with both parents intact. Preserve fork TRACK contracts and Firstmate-resolved away ownership, declared-wait cadence, and supervisor-only chat address. Reconcile overlapping runtime owners and retain the fork's validation, publishing, and cleanup safeguards.
Reconcile fork main 90ca53b as directed by Firstmate while retaining the original upstream merge through 4768e98 and its full parentage. Restore the presentation move audit from its evidence owner, keep live fixture calls behind the lab helper, and include native discovery in the loaded-marker fixture. Rebalance the same portable isolated test set from completed CI timings without increasing workflow or per-script bounds.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge the contiguous upstream prefix through
4768e98d469b7216569acf6e7a5cc696ecd15c2e, preserving fork contracts and both parents.All 16 substantive checks pass on pushed head
a5ee7e365d44e3e4abfbe27376e826dbc584f7c4; only the expected direct-PR delivery-policy check is red.Fork parent:
ea6b934ad220c15c861a9cafe81badc10e192248.Measured merge base:
b84e0e362face25f3dd8945297a3df1320d7668c.The exact round contains 18 first-parent changes and 86 overlapping paths; overlap marks collision risk, including clean textual merges.
No later upstream commit is included.
c034b268746fc0aa0fe226c940c50ea89ba29db255d406911c7a713f78318e2c28153d1b00334ef2269f8fe6b6b51b9c3e817d3fb1ad702f34bd4189527aa7c15300a4f54768e98dContract-file reasoning (runtime/schema owners are additionally covered by the applicability and divergence rows):
AGENTS.md.agents/skills/afk/SKILL.md.agents/skills/bearings/SKILL.md.agents/skills/bootstrap-diagnostics/SKILL.md.agents/skills/captain-hold-lifecycle/SKILL.md.agents/skills/firstmate-coding-guidelines/SKILL.md.agents/skills/fmx-respond/SKILL.md.agents/skills/harness-adapters/references/common/model-and-effort.md.agents/skills/harness-adapters/references/harness/pi.md.agents/skills/secondmate-provisioning/SKILL.mdRuntime and workflow contract-file reasoning:
tests/herdr-presentation-fixture.shtests/fm-backend-herdr-presentation-e2e.test.shtests/fm-pi-loaded-marker.test.shCONTRIBUTING.md.github/workflows/ci.yml.pi/extensions/fm-branch-supervision.ts.pi/extensions/fm-primary-pi-watch.ts.pi/extensions/lib/fm-native-contract.tsbin/backends/herdr.shbin/fm-afk-contract.shbin/fm-afk-launch.shbin/fm-afk-return.shbin/fm-backend.shbin/fm-backlog-handoff.shbin/fm-backlog-transition-lib.shbin/fm-bearings-snapshot.shbin/fm-bootstrap.shbin/fm-brief.shbin/fm-busy-event.shbin/fm-captain-hold.shbin/fm-classify-lib.shbin/fm-claude-stop-autoarm.shbin/fm-control.shbin/fm-crew-state.shbin/fm-fleet-snapshot.shbin/fm-guard.shbin/fm-harness.shbin/fm-landed-lib.shbin/fm-launch-lib.shbin/fm-lint.shbin/fm-mail-check.shbin/fm-mail.pybin/fm-mail.shbin/fm-merge-local.shbin/fm-nm-run-lib.shbin/fm-parent-channel-lib.shbin/fm-pending-reply-lib.shbin/fm-pr-merge.shbin/fm-procevent-lib.shbin/fm-procevent.shbin/fm-remote-doctor.shbin/fm-remote-entrypoint.shbin/fm-remote-herdr-guard.shbin/fm-remote-herdr-owner-lib.shbin/fm-remote-secondmate-control.shbin/fm-secondmate-restart.shbin/fm-session-start.shbin/fm-spawn.shbin/fm-supervise-daemon.shbin/fm-tasks-axi-lib.shbin/fm-teardown.shbin/fm-test-isolation-proof.shbin/fm-test-run.shbin/fm-wake-lib.shbin/fm-watch.shDocumentation reasoning (audiences checked against
docs/documentation-audiences.json; task delivery evidence stays in this PR):docs/agent-control.mddocs/architecture.mddocs/captain-hold-lifecycle.mddocs/configuration.mddocs/fm-test-portable-shards.mddocs/herdr-backend.mddocs/pi-supervision-branch.mddocs/remote-secondmates.mddocs/scripts.mddocs/secondmate-parent-channel.mddocs/supervision-protocols/claude.mddocs/supervision-protocols/pi.mddocs/turnend-guard.mddocs/verification/process-event-sources.mddocs/verification/runtime-backends.mddocs/verification/supervision.mddocs/fork-divergence.mdea6b934a4768e98d90ca53b3CI=true bin/fm-lint.shFork failures: brief's single-entry PATH fixture, GBrain pin fixture exposing the installed binary through jq's directory, and private-state rejection under umask 002.
Upstream failures: missing Ruby in 2 YAML-parser tests (both pass with installed Ruby), the same umask fixture (passes under 022), public-followup teardown's directory-removal race after its behavioral assertions (unchanged rerun passes), fm-on's symlinked Nix PATH expectation, and remote-doctor's assumption that no system harness exists despite
/usr/bin/pi.The last 2 have passing fork fixtures that account for the actual host configuration.
Parent and final suites use the public
bin/fm-test-run.sh --all --per-script-timeout-secs 480entrypoint, with JSON timing output.The fork baseline combines 30 completed initial scripts with the disjoint remaining 209 after a supervised stop/resume; their union is exactly the parent's 239-script manifest.
Canonical lint uses
CI=true bin/fm-lint.sh, pinned ShellCheck 0.11.0 with extended analysis and actionlint 1.7.12.Final validation uses explicit umask 022 and isolated tool directories without changing either parent's tracked source.
Fork-parent brief fixture initially failed under a single-entry isolated PATH; its unchanged public suite passes with the same directory duplicated, because its PATH-copy loop requires a terminated record.
Herdr is deliberately absent from parent and final local test PATHs: unsafe direct lifecycle fixtures cannot reach the fleet; capability skips are not live evidence.
Pi/Gemini comparison: Pi native-main progress and explicit ultra apply in this prefix; supervision-branch provider/session isolation and away standby remain independent guarantees.
Gemini CLI and agy remain separate executable, authentication, trust, and supervision adapters; Gemini is already landed, and fork bootstrap recognition of both survives.
Prior comparison sources: #270 and #275.
Those earlier measurements are historical; final tests below own this round's verdict.
The merged Gemini suite passes executable/crew-scope/settings-hook separation, and the shared launcher suite passes Pi-native ultra admission while confirming Gemini omits unsupported effort.
Pi/OMP extension suites pass daemon standby, pending-wake retention, and resumption; gated live vendor/Herdr tests do not become fresh local proof.
Parked branches remain excluded under
docs/fork-divergence.md:fm/fm-afk-injection-wedge,fm/fm-crew-state-blind-during-fix-round,fm/fm-parked-decision-stale-noise,fm/fm-subagent-model-routing-guard, andfm/fm-vault-drift-check.None was merged, resurrected, rebased onto, or cherry-picked.
Exact round overlap paths (86)
Ledger changes in this round:
sync-tail-chat-role-boundary.sync-tail-pi-afk-owner.Merged validation corrections:
tests/fm-brief.test.sh: regenerate brief fingerprints through the public brief command for upstream's new scout/secondmate deadline guidance; all four ship fingerprints remain identical and the complete brief suite passes on rerun.tests/fm-dashboard-events.test.sh: bring the synthetic Claude Stop fixture into the new recovery-generation contract, assert actual session binding, and normalize only process-specific identities for the instrumentation comparison.tests/fm-issue-writeback.test.shandtests/fm-merge-local.test.sh: give the fixtures a resolvable data directory and explicit Markdown configuration so the held-task guard proves absence instead of failing before the milestone path; both complete suites pass on rerun.tests/fm-remote-doctor.test.sh: restore executable setup of the new Directory Services stub in the merged remote-doctor fixture; the complete login-shell/Aqua suite passes on rerun.tests/fm-spawn-dispatch-profile.test.sh: align the new native-ultra fixture brief with each requested delivery mode; the complete spawn-dispatch suite passes, preserving the fork's mode-mismatch refusal.tests/fm-watch-triage-waits.test.sh: restore the three upstream deadline-fixture helpers in the fork's shared test-helper owner; the complete wait suite passes, including immediate, near-future, and distant deadlines.The fresh final run covers all 249 tracked scripts exactly once, has zero failures, and successfully emits its own JSON summary.
The seven earlier fixture corrections are included in that complete run.
An earlier 246-script attempt lost summary emission after a comment edit changed its running Bash file; its records and counterfactual remain private diagnostic evidence and are superseded by the fresh complete generated summary.
Capability gates in the broad run remain disclosed; full live Herdr presentation/recovery, explicit Pi typechecking, and the Ruby-enabled runner suite provide the supplementary coverage described below.
tests/fm-agy-adapter.test.sh: Crew-only/Herdr-only admission, flag rendering, Gemini separationtests/fm-lint.test.sh: Bounded retry and canonical pinned linter invocationtests/fm-quota-sidecar.test.sh: Eligibility and spend accounting stay independent of upstream quotatests/fm-send-strict.test.sh: Steer reopens the completed scout gatetests/fm-watcher-lock.test.sh: Generation-bound handover and replacement lock ownershiptests/fm-watcher-lock.test.sh,tests/fm-backend-herdr.test.sh: Stop handling and event-wait descriptor ownershiptests/fm-watch-triage.test.sh,tests/fm-watch-triage-waits.test.sh: Keyed suppression with unkeyed recovery and decision-owned wake dedupetests/fm-wake-queue.test.sh: Semantic busy/idle/unknown thresholds and progress resettests/fm-run-progress.test.sh,tests/fm-watch-triage-waits.test.sh,tests/fm-daemon.test.sh: Moving validation defers wedge escalation through fork hold countertests/fm-watch-triage-waits.test.sh,tests/fm-daemon.test.sh: Live wait routing, replacement reset, confirmed-away held recheck and until deadlinestests/fm-backend-herdr.test.sh: Native working baseline avoids inapplicable rendered-footer readtests/fm-remote-job-orphan-reap.test.sh,tests/fm-remote-job.test.sh: Descendants reaped before lane shutdown and ownership quarantinetests/fm-remote-job.test.sh: Open stdin times out before queue admission deadlinetests/fm-test-run.test.sh: 480-second default, opt-out and cancellation attributiontests/fm-test-run.test.sh: Stable coverage comparisons across localetests/fm-wake-queue.test.sh,tests/fm-tangle-guard.test.sh,tests/fm-turnend-guard.test.sh: Durable queued rows keep supervision requiredtests/fm-crew-state.test.sh,tests/fm-teardown.test.sh: Relation table, recorded branch, degraded/abandoned evidence and strict abort boundarytests/fm-trigger-validation.test.sh,tests/fm-brief.test.sh,tests/fm-design-skills.test.sh,tests/fm-control-relaunch.test.sh: Design, continue-branch and ready-to-validate shared renderingtests/fm-backlog-atomicity.test.sh: Unlanded/preserved task cannot retain a false cleanup intenttests/fm-spawn-dispatch-profile.test.sh: Published task remains recoverable after capture failuretests/fm-home-summary-refresh.test.sh,tests/fm-fleet-snapshot-view.test.sh,tests/fm-dashboard-backlog.test.sh: Per-task timeout and durable abandoned statustests/fm-pi-watch-extension.test.sh,tests/fm-omp-harness.test.sh,tests/fm-afk-launch.test.sh: Retired arm, retained pending wake and exactly one resumed extension cycletests/fm-test-fixture-cleanup.test.sh,tests/fm-backend-herdr-presentation-e2e.test.sh: Fixture state and cleanup remain named-session owned; live gate disclosedtests/fm-no-mistakes-required-gate.test.sh: Fork workflow events preserve stable expected direct-PR policy verdicttests/fm-no-mistakes-required-gate.test.sh: Compliance evaluation refreshes live bodytests/fm-pr-merge.test.sh: Landed readback wins failed command; upstream queue acceptance retainedtests/fm-upstream-status.test.sh: Detector remains inert without configuration and performs no publishtests/fm-nm-test-contract.test.sh: Private validation evidence retained outside tracked repository configtests/fm-agents-hard-rules.test.sh: Fork target and explicit upstream write boundary remain enforcedtests/fm-gbrain-lib.test.sh,tests/fm-gbrain-capture.test.sh,tests/fm-gbrain-health.test.sh,tests/fm-recall.test.sh: Home-scoped capture/retrieval and no cross-home writetests/fm-dashboard-events.test.sh,tests/fm-usage.test.sh,tests/fm-outcome-manifest.test.sh: Canonical instrumentation and completion artifacts remain wiredtests/fm-spawn-dispatch-profile.test.sh: Actual generated harness prompt preserves Firstmate reporting and worker address exceptionFor repository-local evidence, the parsed YAML additionally confirms
test.evidence.store_in_repo=false,disable_project_settings=true, no replacementcommands.test, and canonicalcommands.lint; no pipeline was invoked.For upstream-read-only posture, manual review additionally confirms
CONTRIBUTING.mddirects origin and PR publishing to this repository rather than the third-party parent.Herdr fixture cleanup has portable ownership/refusal coverage plus separate local live presentation and recovery runs through the fixed guarded helper and task-generated named labs; both complete suites pass on Herdr 0.8.2 with default-session tripwires intact.
Current fork main
90ca53b37844e1a757d90ba0ccae7813d0772136, landed through #276, is reconciled by a normal merge under Firstmate instruction 005, which also authorized the CI repairs below.The upstream merge
1eea120eae5e1cc70a13f1685480042e89c070e3retains exactly its original parents,ea6b934ad220c15c861a9cafe81badc10e192248and4768e98d469b7216569acf6e7a5cc696ecd15c2e.There is still exactly one upstream merge, ending at the pinned endpoint.
The additional merge parent carries the one already-landed fork-main change; the merge commit also contains the documented CI repairs.
Current main adds exactly one fork commit beyond the original parent and still meets the later upstream target at
b84e0e362face25f3dd8945297a3df1320d7668c; the task branch meets it at the pinned4768e98d469b7216569acf6e7a5cc696ecd15c2e.No squash, rebase, force-push, default-branch rewrite, or parked-branch integration occurred.
Observed CI evidence: https://github.com/HelloWorldSungin/firstmate/actions/runs/34837015817
The active-seeded case is the trigger; separate source and evidence roots expose the wrong restore path; a stale retired-workspace audit entry is the assertion's visible symptom.
The standalone filesystem reproduction returns restore=1 and assertion=1 before correction, then assertion=0 when only the restore source changes.
This is a deterministic fixture defect, not evidence of incorrect production ordering or a timing race.
Both copies now fail immediately on error, and the complete live fixture passes on Herdr 0.8.2, including three repeated concurrent waves and successful default-session tripwire verification.
The fixture supports the generated task lab label and translates its version probe through guarded status, avoiding a direct Herdr invocation outside the lab helper.
The pushed-head Herdr CI lane passes all 16 scripts, with zero failures and one disclosed capability gate; both presentation and recovery complete successfully.
Merging current fork main reproduces the same module-load failure locally before the fix.
Adding the native dependency to the fixture makes the complete loaded-marker suite pass, including the real installed Pi 0.85.1 list-models probe.
Current main's prompt-join and marker owners remain unchanged; the native guarded-tool registration wrapper composes with them.
The corrected loaded-marker suite also passes without a capability gate in the pushed-head serial-4 CI artifact.
The job reached its 600-second cap after five completed scripts, just after starting grok-harness.
Completed measurements from both lanes plus the six retained earlier measurements yield an LPT partition of 544542 ms and 544509 ms, twelve scripts each.
The proven-isolated set, 480-second per-script bound, ten-minute parallel job cap, and eight serial shards remain intact.
On pushed head
a5ee7e36, both complete 12-script lanes pass in CI run 34846962245; total job walls are 9m08s and 8m47s, within the unchanged ten-minute cap.Local limitations: the filesystem reproduction is a diagnostic counterfactual, not a replacement for the real Herdr suite.
The broad normalized local run omits the real Herdr binary so unrelated suites cannot drive the shared fleet outside this task's generated lab contract.
Exact pushed head:
a5ee7e365d44e3e4abfbe27376e826dbc584f7c4.The fork-main merge has parents
97b42e0456c7f2f0164456194cc7fa7bca853af1and90ca53b37844e1a757d90ba0ccae7813d0772136; remote branch identity matches the pushed head.Full canonical lint passes; the final comment-only timing arithmetic refresh also passes explicit canonical runner lint.
Published-head CI run 34846962245 completes successfully on
a5ee7e365d44e3e4abfbe27376e826dbc584f7c4.All 16 substantive checks pass, including lint, both parallel lanes, all eight serial lanes, Herdr, stock macOS Bash compatibility, invariants, coverage, and timing aggregation.
The only red check is the expected
PR must be raised via no-mistakesdelivery-policy check.The portable aggregate covers 233 scripts and Herdr covers 16; their disjoint union matches all 249 tracked scripts exactly, with zero failures and 36 disclosed capability gates.
Live mergeability observed at 2026-09-14T13:23:02Z:
mergeable=true,mergeable_state=unstable, exact heada5ee7e365d44e3e4abfbe27376e826dbc584f7c4, base90ca53b37844e1a757d90ba0ccae7813d0772136.The policy-only red check explains the unstable status; the branch has no merge conflict and remains unmerged for configured authority.
Direct PR is required because no-mistakes rebases and would linearize this merge-only branch.
Only the no-mistakes delivery-policy check may remain red, with every substantive check passing.
Do not squash or rebase this round.
Landing must use
bin/fm-pr-merge.sh fm-upstream-sync-2026-09-11-tail https://github.com/HelloWorldSungin/firstmate/pull/277 -- --merge.The first local recovery attempt placed temporary secondmate homes inside this repository and correctly failed the home-placement guard.
Its lab teardown completed; the unchanged complete suite passes under the normal temporary-root layout, including default-session tripwire verification.
Supplementary Pi typecheck: both advanced-main parent and combined tree pass
tests/fm-pi-primary-types.test.shagainst installed Pi 0.85.1 with task-local TypeScript 7.0.2.The normalized broad-suite PATH intentionally omits that compiler; its capability gate is preserved in the raw summary and supplemented by this complete explicit check.
The complete
tests/fm-test-run.test.shsuite also passes with installed Ruby 3.4.9 exposed through a task-local PATH, including the YAML timeout and shard assertions that the normalized broad-run PATH capability-skips.