Conversation
…or-owed gate (kunchenguid#4974) * fix(watch): recheck a gate awaiting a human instead of wedge-escalating it A lane whose validation run is parked at a gate waiting on a human decision is correctly quiet, but nothing in its status line says so: the evidence is the pipeline's own gate state rather than anything the worker wrote. The wedge timer read that silence as a suspected wedge and climbed the escalation ladder for as long as the wait lasted, and each escalation cost a supervising turn. The landed declared-wait consult does not reach it, because a live ordinary crewmate never reports a declared pause, and raising FM_STALE_ESCALATE_SECS would delay genuine wedge detection for every lane by the same amount. The threshold now reads a second, independent record when the status line accounts for nothing: whether the crew's current state is a gate whose answer is owed by a human. That is minted only from the gate's own findings table, by a row whose `action` column is exactly `ask-user`, located by position out of the table header the way nm_gate_step_row already reads its row - never searched for over the run payload, where a finding's free-text description or a branch name satisfies a search just as well. A gate awaiting the CREWMATE's own answer keeps the unchanged escalation schedule, reason and demand-deep-inspection wording, because a crewmate that goes quiet before answering its own gate is exactly the wedge the ladder exists to catch. Each kind of wait now carries the human it is on, the action that clears it, and whether that human is the captain as data alongside the verdict, rather than as wording chosen per branch where the recheck is written, so the deferral cannot word one kind of wait as another and a new kind cannot ship without deciding all of them. A parked gate has no written record of when its wait began, so its recheck publishes no wait age at all rather than one read from the quiet window this deferral resets on every pass, which would report the same small number for a gate of any age. Like every other captain-facing recheck here it is absorbed in silence while the away-posture record exists, arming no throttle, so the recheck is owed in full the moment the record is archived. The consult runs only in the at-threshold branch that was about to escalate, beside the worktree walk already there, and only for lanes whose status line explained nothing. Closes kunchenguid#3055 * no-mistakes(review): require an unanswered decision before deferring a parked gate * no-mistakes(review): reset the away-silenced timer, fail-safe findings parse, US-joined wait records * test(watch): pass the pane hash wedge_timer_check now takes Upstream gave wedge_timer_check a sixth <pane-hash> argument for its dead-record probe. The malformed-wait-record rounds drive the real function directly, so they pass one, and stub fm_backend_agent_state to a live agent so the probe that runs after a refused deferral keeps the unchanged ladder rather than reading a backend the child shell has none of. * no-mistakes(review): Bind parked-gate wait to its run, owe it firstmate * no-mistakes(document): correct wait-kind count, crew-state reader scope, gate-key coupling * feat(watch): make the parked-gate wait deferral opt-in The wedge timer deferring a lane parked at a validation gate is new supervision behaviour rather than a restored one, and it decides which lanes give up the escalation ladder, so it now ships as a default-off per-home option instead of changing every home on upgrade. config/wedge-defer-parked-gate arms it. The flag is read before the decision fold, so an unconfigured home spends no fold or current-state read, writes no record, and keeps the unchanged escalation schedule, reasons and demand-deep-inspection wording; a test counts the reader calls in both directions to pin that. It is not inherited by secondmate homes: each home supervises its own crew and owns that trade separately, the same reason config/turnend-churn-absorb is home-local. The away-posture absorb returns to leaving the idle timer alone, which it had restarted only because the costly consult could reach it. A parked-gate wait is owed to the supervisor rather than the captain, so it never enters that branch, and the recheck owed on return is again owed in full the moment the record is archived. * test(watch): pin that the away-silenced hold leaves the idle timer alone The absorb no longer restarts the timer, so the recheck owed on return is owed in full rather than a cadence into the return. Nothing asserted that, so a restart could be reintroduced silently. * no-mistakes(review): document away-silence rationale, pin captured gate component * no-mistakes(test): anchor gate row scan to the braced findings header * no-mistakes(document): pin same-block gate row invariant in crew-state comment
…uid#5007) * fix(control): let the owning seat reclaim a task whose endpoint is gone A destroyed pane or workspace made `missing` a terminal state. Relaunch accepted only `dead` and said to stop the agent first; exit refused `missing` and said to reconcile the task first; there is no reconcile verb. Each command named the other as its prerequisite, so a task whose terminal went away could not be reclaimed by anything, and a no-mistakes approval it was parked on had no seat left to answer it. `missing` is agent-free a fortiori: there is no endpoint, so there is no agent in it. Widen the existing guards rather than add a verb. - fm-spawn --relaunch accepts a positively proven `missing` and creates one fresh endpoint in the recorded worktree; the record it already republishes rebinds the task to it. A `dead` endpoint is still adopted in place. - fm-control exit reports `endpoint-gone` instead of dying, so the relaunch transaction's stop step no longer dead-ends, and re-resolves the endpoint from the record before verifying the replacement. The duplicate-agent refusal is untouched: both verdicts come from the same recovery-grade classifier, which claims `missing` only from positive absence, so `alive`, `ambiguous`, and `unreadable` all still refuse. The backends' own create paths refuse a live same-labeled endpoint as a second independent guard. The worktree, its branch, commits, uncommitted changes, armed poll and registration, record rows, and status log are all untouched - a reclaim is a recovery, never a teardown. A secondmate is excluded: its gone-endpoint recovery already has one owner in the session-start liveness sweep, so relaunch refuses and names it rather than becoming a second path to the same outcome. Tests reproduce both halves of the deadlock, the reclaim succeeding, unlanded work surviving it, and the refusals that still hold. * no-mistakes(review): prove endpoint absence per backend before reclaim rebinds * no-mistakes(review): give exit and relaunch one absence proof; pin herdr rebind session * no-mistakes(review): narrow endpoint reclaim to herdr; tmux refuses honestly * no-mistakes(review): stop refusals and docs asserting unestablished causes * no-mistakes(review): stop herdr fixture helper losing tmp-root registration * no-mistakes(review): document workspace drift and absence-probe server residue * no-mistakes(review): correct rebind limitation to its one reachable case * no-mistakes(review): stop claiming reclaim leaves instructions untouched * no-mistakes(document): scope fm-control-lib purity claim, note reclaim coverage * no-mistakes(rebase): read the staged launch file in the herdr fixture Rebasing onto main picked up kunchenguid#4994, which stages a long worker launch command into a script and delivers the short `. '<path>'` line instead of the literal command. The tmux fake and tests/fixtures.sh were updated for that; the herdr fake this branch adds was written before it and still keyed "an agent now exists on this pane" off the literal `encode launch-brief` text, so after the rebase it never marked the rebound pane live and the reclaim's alive-wait read `dead`. Dereference the staged file first, exactly as the tmux fake above does. Test-fixture only; no production path changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(document): note reclaim placement in herdr and scripts inventories --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…3764) * test(status): reproduce missing event emission time * wip(status): preserve optional event emission time * test(status): document indirect clock stub invocation * no-mistakes(review): Preserve historical status bytes during reply recovery * no-mistakes(test): Fix timestamped status assertions and remote fixture dependencies * no-mistakes(review): Preserve captain regex overrides for timestamped status events * no-mistakes(document): Clarify status event timing and publication contracts * no-mistakes(lint): Quote literal done to satisfy ShellCheck * no-mistakes(ci): Captain, updated .github/workflows/ci.yml to expect 19 snapshot tests instead of 18, matching the PR’s added regression. Reproduced the failure before the fix. Stock Bash 3.2.57 verification passed: parse sweep, 19 snapshot tests, 53 Bearings tests, and the public-followup regression. Workflow lint and diff checks passed * no-mistakes(test): Preserve terminal notifications with malformed timestamp tags * no-mistakes(test): Stamp Rovo spawn failures with emission time * no-mistakes(document): Verify status event documentation * no-mistakes(lint): Fix ShellCheck quoting in status emission-time tests * no-mistakes(ci): Captain, fixed four lifecycle assertions to accept emission timestamps while preserving publication and retry checks. Reproduced the CI failure before the fix. The lifecycle suite now passes with six Beads capability skips; syntax, targeted ShellCheck, and diff checks passed * no-mistakes(ci): Captain, fixed malformed timestamp colons hiding actionable events using shared normalization. Original bytes and unknown ages are preserved. Regression reproduced before the fix; classifier and remote-reply suites, targeted lint, syntax, and diff checks passed * no-mistakes(review): Stamp remote escalations at call sites, drop new flag * no-mistakes(review): Accept stamped escalation and close lines in test assertions * no-mistakes(review): Restore reserved-key answered-note guard for stamped closes * test(status): accept optional emission time in PR-provenance assertions The kunchenguid#4148 provenance test landed on main with exact unstamped greps. Parent-channel lines from this branch carry [at=<epoch>], so strip only that tag before the same exact match. No production change. * no-mistakes(review): Accept stamped ready signal in PR fallback scrape * no-mistakes(review): Drop relay flag, stamp parent events at call sites * no-mistakes(review): Stamp worker terminal-signal instructions, revert fm-on fixture * no-mistakes(review): Accept optional stamp in live cmux drift guard * no-mistakes(review): Restore original test invocation order in two suites * no-mistakes(review): Strip only well-formed numeric status time tags * no-mistakes(document): Drop stale unstamped PR-ready line spelling from channel doc * no-mistakes(review): Stamp agy spawn-failure status lines with event time * fix(bin): normalize status event times in-shell and freeze the budget test clock Two paths made a status event's emission time cost more than it should. The captain-relevance fallback piped every line through awk to drop a well-formed `[at=<epoch>]` tag before matching, so a supervisor sweep paid a fork per line just to prepare a regex match. Shell parameter expansion does the same strip with no fork, and the retry-dedup scan now reuses that one helper instead of carrying a second copy of the rule in awk. The copies had already drifted: the shell side stripped tags from lines with no colon, which the awk rule left whole, so a colonless line could be mistaken for one already recorded. One definition, checked against the awk rule it replaces over the edge cases and a 4000-line fuzz. tests/fm-contributions.test.sh froze its fixture clock only in exhaust mode. In hang mode the poll set DEADLINE to the real now plus a one-second budget, and when the second ticked before the first forge call the loop broke without ever calling gh: forge/calls was never written and the assertion failed reading a missing file. Freezing the clock in both modes removes the dependence on wall time; the bounded call is still cut by the real timeout, so the observation the test asserts still starts. Emission time stays optional on new status records, and legacy or malformed lines keep an unknown age. * no-mistakes(review): Stamp ask-user escalation line and fix Kimi status assertion * no-mistakes(document): Drop stale unstamped done-line spelling from watcher docs * test: fold emission-time snapshot coverage into the fixture case Drop the incidental ci.yml 18-to-19 count hunk so the PR no longer touches workflows. Keep every emission-time assertion by folding it into test_fixture_snapshot_json. * no-mistakes(review): replace brief date substitution with epoch placeholder; drop emitted_at_epoch * no-mistakes(review): align untimed normalizer with epoch parser; tolerate placeholder stamp in PR scrape * no-mistakes(review): strip undelimited at-tags; correct brief stamp header * no-mistakes(review): normalize stamps at both captain-regex sites; restore mtime freshness * no-mistakes(review): strip colon-bearing stamps for relevance; fix headers and test oracles * no-mistakes(review): narrow escalation match to stamp tolerance; pin note verb * no-mistakes(review): read note and key past colon-bearing stamps * test(status): keep inactive reconcile assertions stamp-tolerant These two oracles were made stamp-tolerant while resolving one of the branch's merges from main. The rebase drops merge commits, so that adaptation was lost and both assertions went back to matching an exact substring that a stamped line no longer contains: the tag lands before the colon, so "failed [key=k]: ..." is now "failed [key=k] [at=N]: ...". Strip a well-formed tag before matching, as the branch's other oracles do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * no-mistakes(review): unstamp fold colon tests; reserve stamp width in cap * no-mistakes(document): correct stale unstamped status-line spellings in docs * no-mistakes(document): quote brief-test literals for lint; correct stamp-helper contract comments * no-mistakes(ci): rename subshell-local epoch in delivery-race stub The serialization test overrides fm_pending_reply_mark_delivered inside a (..) subshell. Its `epoch` local collided with the same name in status_line_at_epoch/status_stamp_line, which this branch added and this suite now calls at top level, so ShellCheck 0.11.0 reported SC2030 and failed Lint 2. The stub already prefixes its other locals with `pending_` for the same reason; `epoch` was the leftover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: ship clean Lavish host fixes * no-mistakes(review): Fix Lavish classifications and fail-closed host loading * no-mistakes(review): Restore Lavish host state across retries and launches * no-mistakes(review): Preserve destination Lavish host when configuration is absent * no-mistakes(document): Document Lavish status and host guarantees
…#5076) * feat(afk): make the captain's away words the whole mandate Retire the clause fields, verb list, never-set scan, refused records, and the per-task merge-grant list from the away-posture record. The record is now version 2: the captain's words verbatim plus expected return, spend cap, and reach line; a version 1 record still validates, reads, and archives so a live away window is never broken by the upgrade. The supervision branch reads the words at the tail of every wake and acts on them by its own judgment through the guarded scripts under standing authority, never by analogy, holding for the return on doubt, and opens each such outcome summary with "per your away instructions:" so the return brief can render the words beside the session's account. While the record exists any green merge runs under away authority (ledger tag "away"); red merges, --allow-red, asynchronous and queued merges, and local-only landing stay refused. The branch may file a backlog item the words explicitly call for before dispatching it under the spend cap. Tests drive fm-afk-contract.sh, fm-afk-launch.sh, fm-afk-return.sh, and fm-pr-merge.sh as commands: version 2 written, version 1 read, retired flags and subcommands refused by name, green merges landing under the record, red and waived-red refused, the record lock still closing the authority-read window, and the Pi away tail carrying the words. * no-mistakes(review): carry the away read-back to the session verbatim * no-mistakes(review): match the exact away-action marker in the return brief * no-mistakes(review): refuse a words block truncated by a damaged line * no-mistakes(document): Refresh away-role contract documentation
…unchenguid#5049) * fix(bin): render the remote charter's steering-inbox path host-local A freshly provisioned remote secondmate read a parent-home absolute steering-inbox path in its charter - a location that exists on no route - and spent its first turn discovering the gap and filing a blocked decision for what was a render defect. The seed's remote-copy rewrite now maps the inbox to the route's host-local parent-route inbox, exactly as it already maps the reply-log path, so every mention - bare path, listing, and handled/ acknowledgement - lands host-local. Both rewrites also become plain assignments, because a quoted substitution nested inside a double-quoted printf argument leaks literal quotes into the replacement text on stock macOS bash. The lifecycle suite pins the corrected render both directions against the real seed, provisioning, and delivery route, sharing one fixture value between the render truth and the delivery truth. Closes kunchenguid#5012 * no-mistakes(document): document remote charter's host-local steering inbox
) * feat(procevent): route worker-owned Lavish rounds * no-mistakes(review): drop duplicate artifact field from task-owned registration * no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic * no-mistakes(review): keep worker board owned until terminal round acknowledged * no-mistakes(review): refuse every retirement of an open worker-owned round * no-mistakes(review): use real lavish reply flag, isolate reply generations * no-mistakes(review): drop .posted marker for best-effort reply posting * no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures * no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms * no-mistakes(review): re-arm only to acknowledge an open round * no-mistakes(review): conclude only a still-open terminal round * no-mistakes(review): record the acknowledgement before retiring the board * no-mistakes(review): retain the registration across a conclude, qualify terminal docs * no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107) * fix(bin): reserve contribution observation budget * no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103) * feat(bin): add idempotent inbox orders, receipts, replies, and readiness Let a caller supply a request id when publishing a captain inbox note so a retry returns the original note instead of creating a second one, including across the crash window between save and wake announcement. Separate saved from announced so a failed wake is repairable without enqueueing again. Add bounded receipts JSON with omission disclosure, a durable primary reply against a note id, and a read-only readiness projection that can say unknown instead of inferring liveness from a lock file. * no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict * fix(bin): resolve ready from lock-holder ancestry; drop lock status --json Remove the extra JSON surface from fm-lock.sh so its human status still always exits zero. Have the readiness projection classify the inspected home from the lock-holder pid via fm-harness.sh ancestry, with an explicit FM_SUPERVISION_MODEL still winning and an unknown model when there is no holder. Prove the yes path when that ancestry names a known harness. * no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor * no-mistakes(document): Note read-only lock inspection in scripts inventory * no-mistakes(lint): Pass missing id argument to malformed-reply test printf --------- Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118) * fix(composer): stop a harness footer row from reading as a composer holding text A harness draws its own furniture below the composer - a user statusLine, a permission-mode hint - and the cursorless "bottom-most shape wins" rule looks exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text everywhere else, so a statusLine opening with `→` was selected as a bare composer, swallowed the hint row beneath it as wrapped input, and answered `pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first doorbell and every retry were skipped and the worker never saw the steer. Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on Herdr 0.8.0 had genuinely empty composers and every one of them was refused. A separator pair that closed over a bare agent-glyph row is a proven composer container, so the contiguous non-blank rows below its closing rule are that composer's footer and are no longer composer candidates. The demotion is bounded by all three of its own preconditions: a blank row ends the zone, a pair that closed over no glyph row demotes nothing, and a shape with no separator pair at all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same composer, including a stray SGR mouse report left by a click in the pane, still reads `pending`. Pinned by two portable regressions and by a new cursorless arm on the live composer-matrix guard, which re-reads each harness's already-proven-idle pane the way every non-tmux backend reads it and fails naming the harness and version when that read is `pending`. * no-mistakes(review): make composer footer-zone demotion shape-independent * no-mistakes(review): make footer-zone demotion refuse-only and drop rescan * no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100 --------- Co-authored-by: Koen Muller <koen@catapult.nl>
…5115) Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114) * fix(bin): stop misreading a no-run branch as an unreadable runs table Defect: when `no-mistakes axi status`'s overview is truncated (a task's own branch has zero rows among the shown ones), fm_nm_select_run's Python fallback derived the repo identity for its direct SQLite query from a `repo: <path>` line it expected in the overview text. The real CLI never emits that line, truncated or not (see the genuine capture at tests/captures/no-mistakes-v1.70.1/overview.toon, which has only `count:`/`runs[...]:`), so the lookup always failed and reported "unreadable runs table" for a task that simply has no run on its branch. On a fleet with many concurrent runs, every idle-branch task hits the truncated-overview path routinely, so this fired every few minutes and drowned genuine unreadable/blocked verdicts in noise. Fix: derive the repo identity from the task worktree path instead, which is exactly the value `no-mistakes` records as a repo's `working_path` (confirmed against the existing capped-overview test fixtures, which already register repos by worktree path). A worktree path that is not absolute cannot be matched and still reads as unreadable rather than being guessed at. Also raise the reader's SQLite busy timeout from 1s to 30s so ordinary lock contention on a busy fleet cannot masquerade as an unreadable database. Safety: every other verdict byte-for-byte unchanged - the repo lookup still requires exactly one matching row (a genuinely corrupt or mismatched repos table still reports unreadable, per the existing `repo` failure-mode test), the branch query and row validation are untouched, and a zero-row result for the branch still flows through the same recursive re-parse that already turns an empty `runs[0]{...}` table into `absent`. Added a regression test (test_capped_overview_without_repo_line_and_no_runs_reports_absent) that reproduces the real overview shape - capped, zero rows for the task's branch, no `repo: ` line - and asserts the crew state falls through to the pane/busy verdict instead of reporting unknown or "unreadable". Full fm-crew-state.test.sh suite passes unchanged otherwise. * fix: recovered same-branch inventory awk misreads empty result as unreadable fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:` overview and re-runs it through the same awk selection pass. When that rebuilt inventory has zero rows for the branch, the row-matching loop never executes, so its counters (`seen`) stay at awk's uninitialized empty string while `expected` and `shown` are plain strings parsed from the header text. Comparing an uninitialized value against a non-numeric string uses string comparison, so "" != "0" is true, and the END block takes the "unreadable runs table" branch instead of falling through to the correct "absent" verdict for a branch with genuinely zero runs. Coerce the affected END comparisons with `+0` so they are always numeric, matching seen/expected/shown/total regardless of whether awk classified them as strings or numeric strings. A truncated or genuinely malformed inventory still differs numerically and still reports unreadable. * no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup * no-mistakes(review): match recorded repo path first, tolerate duplicate spellings * no-mistakes(review): revert repo lookup to exact working_path match * no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141) * fix(bin): require a non-draft pull request before a PR-based done report A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge. The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done. bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before. The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged. Closes kunchenguid#4757 * fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more than one account: every provider row carries an accountKey and one provider id may appear on several rows. fm_quota_json_valid accepted only schema 5 with unique provider ids, so fm-dispatch-resolve.sh, fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live snapshot and quota-informed dispatch was dead against the current tool. - bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with accountKey required on every row and uniqueness on provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ is the one join every consumer uses: schema 5 binds by provider alone, schema 6 binds to the row keyed by the candidate's Pi lane, else the provider's default row, else no row (unmeasured, never blocked, never by position or summed across accounts). - bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey column, and joins through the shared function. - bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through the shared function; an expanded provider with no row for the candidate's account is reported as such. - tests: schema 6 fixtures shaped like the real snapshot, each paired with a schema 5 case on the same path; every new case fails on the previous scripts and passes now. - docs: the two sentences naming the row join describe the schema 6 key. * no-mistakes(review): Fix native Codex quota and expanded provider watches * no-mistakes(review): Align native Codex account matching across dispatch paths * no-mistakes(document): Align quota documentation with account-aware snapshots * no-mistakes(document): Align quota dispatch documentation with account matching * fix(bin): keep CI lint and the quota watch test portable - bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts that source this library, so full-mode ShellCheck reported SC2034 on the assignment; mark it alongside the existing SC2016 disable. - tests/fm-procevent-quota.test.sh: the schema 6 provider-watch assertions used rg, which CI runners do not install, so the case failed with 'rg: command not found' rather than on behavior; use grep like the rest of the file. * no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression * no-mistakes(review): Assert SessionStart digest completeness within its hook_response event * no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
…nchenguid#5174) * fix: preserve Pi watcher ownership across session replacement * no-mistakes(document): Scope Pi predecessor retention away from omp * no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236) * fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown Catch-up correctly refuses while a leftover task record has no status file. Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers * no-mistakes(review): Validate windowless leftover identity via shared endpoint validator * no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity * no-mistakes(document): Clarify windowless teardown retry documentation --------- Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250) * fix: surface parked launch prompts as not started * no-mistakes(document): docs: record launch-prompt busy backstop classification * no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write Collapse the propose-then-confirm away entry into one 'enter' step that writes state/.afk-contract immediately and prints the announcement and read-back after the record exists, never asking for a go. The retired propose, confirm, and --proposal inputs are refused by name, and a stale proposal left by an older version is removed rather than promoted. Refresh and replace semantics, verbatim words, the single writer, the never-set, and per-harness launch behavior are unchanged. * no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294) * fix(bin): map passed-with-override to done instead of unknown no-mistakes' axi status emits outcome: passed-with-override for a run that finished with an explicitly approved Test or CI exception. Both bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's pre-teardown terminal-run check only matched the literal passed and checks-passed tokens, so this outcome fell through to unknown/parked and a finished worker awaiting merge kept getting re-alerted as stale, while an abort race during teardown could also leave a finished run misreported as still parked. Map passed-with-override to the same done/terminal handling as a clean passed in both places. * fix(document): Replace stale outcome mapping with authoritative pointer * fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return During the 2026-09-22 away window every exemption worker whose pull request had merged was left sitting for nine hours. The supervision branch received the stale wake, the merge-landed check, and the hourly inactive-outcome row for each of them, ran the recovery playbook, found nothing to recover, and reported "no further action". The branch prompt granted ordinary teardown of a confirmed-landed task without ever naming the moment or the command, and the playbook has no landed exit, so the stale path ended at "nothing to recover". The return brief then listed only blockers, decisions, and the latest five routine outcomes, so the landed workers stayed invisible after the captain came back. - bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale, inactive-outcome, or heartbeat row on a done task with a merged PR, as the moment to claim the lease and run bin/fm-teardown.sh with no flags; a refusal is reported, never forced or worked around. Add teardown to the handling tool list. - stuck-crewmate-recovery: a landed worker is not a recovery case; point at the ordinary teardown owner for each actor. - bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable records only (a live task record whose recorded PR carries the merge-notification marker), between could-not-fix and handled, without holding the gate; the afk skill's return step closes each listed task through ordinary teardown once the check clears. - tests: pin the prompt rule in fm-branch-supervision and the brief section in fm-afk-return through the real marker writer. * no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring A green PR could sit unreported because neither the worker nor the supervisor could observe checks-green while the ci step kept monitoring for the merge. Supervisor read: fm_nm_select_run's capped-overview inventory reader looked the repository up by the task worktree path, but no-mistakes registers a repository once by its main clone path and resolves every linked worktree to it, so on every task copy of a busy repo the lookup matched no row and each read reported "complete same-branch run inventory unreadable". Key the lookup on the overview's own top-level `repo:` line, which every axi release emits as the resolved working_path. Even with a readable run, the ci-log classifier treated "base branch advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs a checks state only when it changes and a base advance does not clear readiness, so a green PR read as still validating for as long as main kept advancing. Stop treating that line as a marker, matching no-mistakes' own ci-log parser, and name the run's PR URL in the held-for-merge reading so the existing inactive-outcome path can act on it without a worker report. Worker contract: `axi status` never reports checks-passed while the ci step monitors for merge, so the definition of done no longer makes a status poll the wait for the next gate or outcome; the drive call's own return is the green signal, reattached with `no-mistakes axi run` after a bounded return. * no-mistakes(review): read the full ci log when checking checks-green * no-mistakes(review): correct stale ci log tail wording in docs * no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session * no-mistakes(document): Document session-derived Lavish polling * no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900) * fix(bin): ignore vanished state scratch files on secondmate relaunch Relaunch refused when find(1) exited non-zero while listing a secondmate home's state directory. A live watcher can delete scratch files between readdir and processing, which is not evidence that child *.meta records are unreadable. Prove the directory is listable from its mode and keep the existing readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765. * no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907) * fix(bin): treat home-owned status closes as already read Self-announced bookkeeping appends now record their exact byte ranges. Later drains and signal scans skip those ranges, so two distinct --resolve-key answers after an OPEN DECISIONS fold do not each wake the supervisor. Worker-authored lines outside that ledger still signal. * no-mistakes(review): Keep owned closes in unread status; lock ledger writes * no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake * no-mistakes(review): Require real owned growth before ledger marks status seen * no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS * no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test * no-mistakes(review): Align ledger docs and scope ledger to wake path only * no-mistakes(review): Restore stranded historical-annotation test comment to its function * no-mistakes(review): Retire the home-appends lock alongside its ledger * no-mistakes(document): Note ledger's lock-helper dependency in classify library * no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions * no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger * no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger * no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350) * chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the floor-boundary test fixtures to the new versions. tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final expecting pr-merged, so add the regression test: a bound work that ends failed reports its honest outcome text through fm-public-followup-emit.sh, consume marks the commitment ready, and deliver posts that text exactly once. Also make two hang-guard tests in fm-backlog-atomicity portable to hosts without coreutils timeout, and stop an installed herdr from leaking into the secondmate-liveness husk classifier test. * no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test * no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures * no-mistakes(document): Document failed public-followup delivery behavior * no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878) * fix(bin): refuse ship done: when the named head lives only in the worker copy A ship done: is not current-state done until that exact commit is reachable outside the disposable copy. The check tests the named head, not whether some branch moved. * fix(bin): gate CI-ready ship done: on named-head reachability, not handoff Keep no-mistakes' first done: as the pipeline handoff, apply the same shared check when registering a PR and when a secondmate publishes ledger-first, treat a recorded merged PR as landed after prune, and name the PR head instead of scanning free-text SHAs. * no-mistakes(review): Bind named-head gate to recorded PR and forge heads * no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries * no-mistakes(review): Align worker done wording, test mapping, pending-retry test * no-mistakes(test): Raise watcher test time limit to stop load flake * no-mistakes(document): Restore ledger-path fact and name named-head gate coverage * ci: re-attest named-head ship-done gate for a fresh serial-3 verdict * no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery * no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204) * fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows. * no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#6002) * fix(tests): disable Claude Code's auto-updater during live harness runs fm_live_gate let a live run proceed without ever setting DISABLE_AUTOUPDATER, so a live Claude test could let the real updater repoint ~/.local/bin/claude into a temporary directory and stop every Claude process on the machine from starting. Export DISABLE_AUTOUPDATER=1 on every path where the gate lets a live run proceed, and assert the export in tests/fm-live-gate.test.sh, including that it reaches a child process the same way a real harness pane would inherit it. * no-mistakes(ci): Greptile flagged that the PR's DISABLE_AUTOUPDATER inheritance test only checked a `bash -c` direct child, not the fm-spawn.sh launch path. The user chose to fix it with a regression on that path. In tests/fm-live-gate.test.sh I replaced the generic child test with test_disable_autoupdater_reaches_the_claude_pane_on_the_fm_spawn_launch_path: it drives the real fm-spawn claude launch through the spawn fixtures, captures the exact staged launch command, and runs it as a synthetic pane whose only `claude` is a stub recording the inherited DISABLE_AUTOUPDATER, asserting it saw 1. Switched the file to source fixtures.sh (pulls in lib.sh, guarded) for the spawn helpers. Verified it is a real guard: the stub records `1` when the ambient var is set and `unset` when absent, so it fails if fm-spawn ever scrubbed the variable (e.g. env -i or -u). This confirms fm-spawn's launch construction never references the name and passes it through via ordinary ambient inheritance with no allowlist. Full suite passes (12 tests ok), shellcheck clean. Note for the outer executor: I embedded the daemon caveat as a code comment in the test, but the finding also asks the PR body to state that a backend daemon already running before the gate exported the variable does not inherit it and fully covering that would need launcher support - that forge-side PR-body sentence is outside this CI phase's scope * no-mistakes(ci): Fixed Greptile finding ci-1. Root cause: fm-spawn.sh handed its launch command to an already-running backend daemon that never inherited the test process's exported DISABLE_AUTOUPDATER, so ambient inheritance dropped it and Claude's auto-updater could still run. Fix (bin/fm-spawn.sh): when DISABLE_AUTOUPDATER is set in the spawn's own environment, embed `export DISABLE_AUTOUPDATER=<value>;` into the LAUNCH command text (same idiom as the adjacent COMPACT_ADVISER_DISABLE export), so it survives a daemon-built pane, the env -i allowlist path, and relaunch alike; gated on presence so ordinary spawns are unchanged. Added regression test test_disable_autoupdater_survives_a_daemon_pane_that_never_inherited_it in tests/fm-live-gate.test.sh: stages a real claude launch with DISABLE_AUTOUPDATER set in the spawn env, then runs that exact command in a synthetic pane with `env -u DISABLE_AUTOUPDATER` and asserts the claude stub still recorded autoupdater=1. Verified the test fails (autoupdater=unset) without the fix and passes with it; the round-1 ambient test stays green either way. Full suite passes (13 ok); test file and isolated snippet shellcheck-clean (full fm-spawn.sh shellcheck kept getting terminated by the memory-constrained host, not by findings). Forge-side note for the outer executor: the PR-body caveat that fully covering the daemon case would need launcher support no longer applies to the Claude launch path and should be corrected
…cevent record (kunchenguid#6010) * fix(bin): ring the inbox doorbell only for a newly published procevent result publish_result rewrote a worker's captured Lavish round idempotently on every reconcile, unconditionally moved an already-acknowledged inbox record back out of handled/, and rang the doorbell every time - so an already-processed round rang the owning worker on every cycle. Snapshot the existing active and handled records before the idempotent write and ring, or move anything, only when the write actually created a fresh record; re-delivery of a still-open round is left to the inbox's own re-ring ladder. * no-mistakes(document): docs: reflect worker-board doorbell rings only on fresh inbox record * no-mistakes(ci): Fixed Greptile finding ci-1 in tests/fm-procevent.test.sh (test-only change). The redelivery regression previously moved the delivered note into handled/ before any repeated reconciles, so it only proved an acknowledged note stays quiet and would still pass if an unchanged active note rang every cycle. Per the user's instruction, I inserted (before the mv into handled/) five repeated `pe reconcile` runs with the note still in the active inbox and asserted the ring log holds exactly one line and 001.msg remains active; the existing acknowledged-note assertion after the move is kept unchanged. No product code changed. bash -n confirms syntax is valid; the block mirrors the already-passing post-move reconcile/ring-count assertion directly below it
…unchenguid#6032) * fix(bin): make the Claude Stop auto-arm refuse arguments before arming A model running bin/fm-claude-stop-autoarm.sh --help mid-turn armed a real supervision-host park owned by its short-lived tool process, leaving supervision down once that process exited. The Stop hook passes no arguments, so -h/--help now prints usage and any other argument is refused before anything is sourced, read, or armed. * no-mistakes(document): Clarify Claude Stop hook documentation for manual invocations * no-mistakes(ci): Updated the argument-run regression test to compare checksums of state files as well as entry names. The Stop auto-arm test suite passes, and git diff --check is clean * docs: restore the bin/ toolbelt intro's manual-use clause The document step dropped "interactive entrypoints work by hand too" from docs/scripts.md, which still holds for most bin/ scripts. * no-mistakes(document): Clarify Claude auto-arm manual-use guidance
…uid#6039) * feat(calm): show supervision sailboat and anchor notes on Claude Code The Calm mod follows a bounded display tail copy of the outcome store, which bin/fm-branch-outcome.sh append now refreshes, and the supervision host's latch, and appends one dim transcript line per visible routine outcome, captain outcome, and latch change, replaying unread and unprocessed outcomes at session start. It shows them whenever the mod is active, regardless of config/calm, and never marks anything read. * fix(calm): show each supervision note once per session on Claude Code Claude Code 2.1.283 stores ui.log lines in the session and restores them on --continue, so the mod records how far each session has followed the outcome store and a resume replays only newer outcomes. It also checks file existence before reads so absent files do not log debug errors. The live guard gains the supervision-notes scenario and the dated 2.1.283 record documents the observed behavior. * docs: name the Claude supervision note row as the engine draws it * no-mistakes(review): Seed outcome tail on present and anchor first tail on markers * no-mistakes(review): Seed outcome tail at session start; replay against start markers * no-mistakes(review): Bound outcome tail by bytes; reread recently changed files * no-mistakes(review): Skip store validation when outcome tail already exists * no-mistakes(document): Clarify bounded Claude supervision note replay * no-mistakes(ci): Fixed seed-tail to validate only a bounded suffix of complete store rows and write it through the existing byte- and row-limited tail writer. Added a regression test with malformed history outside that window and updated the script header. Outcome tests and shellcheck passed; the full session-start suite timed out after 240 seconds
…enguid#6033) * fix(bin): read a quiet-mode record as a present captain, never hold-for-return Daemon-backed quiet mode writes the away-posture record marked mode: quiet, but the entry announcement, read-back, and session-start digest rendered it as "hold-for-return only", and the spend cap and PR merge gate treated it as away. A present captain's requested actions could then be held for a return that was not coming. bin/fm-afk-contract.sh now owns which posture a record is (the mode subcommand, fm_afk_contract_mode, fm_afk_contract_away_present). A quiet record announces, reads back, and appears in the digest as a present captain holding nothing; merges under it stay attended and it binds no spend cap. An away record is unchanged, an /afk entry over quiet mode rewrites the record as away, and a quiet entry never turns a standing away record quiet. * no-mistakes(document): Clarify quiet-mode authority and remove stale away guidance * no-mistakes(ci): The CI failure came from a race in the supervision-host test: its restart fixture could observe a watcher left by the preceding cycle. The test now retires that watcher and waits for the fixture arm to report its own started cycle. The focused test passed three times; the full suite was attempted but stopped at a separate intermittent test failure * no-mistakes(ci): Fixed daemon refresh mode selection so an unset-mode refresh follows the posture record: /afk over a running quiet daemon now changes state/.afk to away, while a plain quiet refresh stays quiet. Added script-level regression coverage for start and start-native and corrected a quiet-refresh fixture. The launch test suite, syntax checks, and diff check passed * no-mistakes(ci): Herdr was blocked before tests ran by a GitHub HTTP 500 downloading pinned Treehouse; no code change was warranted for that check. Fixed the Lint 1 ShellCheck warning in tests/fm-afk-launch.test.sh by annotating the intentional background PID capture. The focused test suite, ShellCheck, syntax check, and diff check passed
…merge (kunchenguid#6053) * fix(bin): accept a task's next PR once fm-pr-merge confirms the bound one merged require_recorded_pr_identity now checks fm_pr_poll_merge_already_notified for the recorded pr= before refusing a different URL, so a task's later PR is accepted once its earlier PR's merge is confirmed, while it keeps refusing while the bound PR is still unmerged. * no-mistakes(document): docs(fm-pr-merge): note next-PR accepted after bound PR merges
…6064) * fix(bin): read a live quiet record as a present captain at the host and watcher A quiet record left without its daemon (a quiet start that never ran or was interrupted) was read as away by the supervision host, so it parked a present captain's main and held captain outcomes for a return that never comes, and the watcher and daemon silenced captain-held rechecks on record presence. The host's posture checks, the watcher's and daemon's captain-held silencing, and the host's outcome path (branch report, drain BRANCH OUTCOMES, relocated branch authority, the owners' away wake note, and the Codex checkpoint bound) now ask the record owner's away-or-quiet reading, so only an away record is away. A live away record keeps today's behavior. * no-mistakes(document): Correct quiet-record documentation and supervision guidance * no-mistakes(document): Clarify quiet-record posture and captain-held rechecks * no-mistakes(document): Clarify quiet-record posture in documentation
…kunchenguid#6043) * fix(bin): name an in-window engine latch in the return brief and drop the false handling GAP line The away return brief said nothing had failed after the supervision host latched on engine errors during the window, and printed a GAP: watcher downtime line whenever a wake was merely being handled or queued at return. The failures section now reads the host ledger and latch record and names the latch time, the window's engine-error count, and whether the session is still paused or recovered. An open recovery episode is reported as information, and as a gap only when a queued episode outlived the return grace or the marker cannot be read. * no-mistakes(review): Fix latch trip time, drop marker-age grace, bound error count * no-mistakes(review): Report paused latch without ledger trip row; bound errors * no-mistakes(review): Never report a failed probe's latch row as trip time * no-mistakes(review): Only a retained trip row marks a pre-window latch * no-mistakes(document): Clarify return-brief latch and watcher-gap documentation * no-mistakes(ci): Fixed Lint 1 by marking the shared cooldown constant as used by sourcing scripts. The repository lint command and diff check pass; the return test run was stopped by a 180-second timeout after its completed cases passed * no-mistakes(ci): Fixed the return brief so the trip time and error count come from the same initial latch row, and ledger rows before the current session’s lock boundary cannot affect its latch report. Added real-script regressions for both findings. The return test suite, repository lint, and diff check pass * no-mistakes(ci): Fixed the return brief’s restart cutoff so it retains in-window failures, prints one line per initial-trip row, and omits zero-error count wording. Added real-script restart regressions. The return test suite, ShellCheck, and diff check pass * no-mistakes(ci): Fixed the return brief so a recorded trip followed by recovery stays recovered, while a later pause with a lost trip append gets a separate “trip time unavailable” line. Added a real-script regression that failed before the fix. The return test suite, ShellCheck, syntax checks, and diff check pass * no-mistakes(ci): Fixed the false second latch during recovery. A real-script regression failed before the fix and passes now; the lost-second-trip test still passes. The return test suite, ShellCheck, syntax checks, and diff check pass
* fix(calm): name the Claude Code Calm plugin fm so supervision notes read "fm: " Claude Code labels every mod transcript line with the plugin name, so the notes rendered as "firstmate-calm: ⚓ ...". Rename the plugin to fm, update the live guard to assert the fm: label, and document the one-time replay for sessions resumed across the rename. * no-mistakes(document): Clarify Calm plugin rename in documentation
…#6037) * feat(bin): add fm-live-lab.sh, a one-command live supervision lab builder * fix(bin): exact lab windows, per-lab task ids, self-safe teardown * fix(bin): target lab windows by id, stop lab descendants, add readiness tests * fix(bin): keep Claude's auto-updater off in live labs; list fm-live-lab.sh * fix(bin): start the lab tmux server without user config * no-mistakes(review): Scope lab teardown to its store, root, and task ids * no-mistakes(review): Record selected user stores at up for check and down * no-mistakes(document): Clarify live lab documentation and remove stale narratives * no-mistakes(ci): Fixed the CI failure by checking for an existing lab root before looking up the harness executable. The affected behavioral test and shell syntax check pass; the refusal also works with Claude absent from PATH * no-mistakes(ci): Fixed all four Greptile findings: teardown signals only recorded lab processes and their descendants; the worker gate is in its granted task directory and its path is exposed; readiness uses current crew state; and mate and worker IDs use 12 nonce hex digits. The CLI behavior tests pass, as do shell syntax, ShellCheck, and diff checks. The Claude no-host path is unchanged * no-mistakes(ci): Fixed the CI test’s dependence on an installed Claude binary by supplying a test-local stub. The full fm-live-lab test, shell syntax check, and diff check pass * no-mistakes(ci): Fixed all three selected findings in bin/fm-live-lab.sh: down waits for recorded processes and escalates before cleanup, PID roots are checked against recorded start times, and Claude primary trust is rechecked after mate/worker readiness. Added behavioral tests in tests/fm-live-lab.test.sh. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed the pre-primary settle wait, worker gate instructions, unused retry variable, and teardown PID revalidation in bin/fm-live-lab.sh. Added behavioral tests in tests/fm-live-lab.test.sh. Both requested commands pass: tests/fm-live-lab.test.sh and bin/fm-lint.sh * no-mistakes(ci): Fixed teardown to track pre-kill lab processes by PID and start time, including children orphaned when a root exits. Up now rejects an empty pane PID before calling ps. Added regression tests and a Linux-safe worker fixture. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass * no-mistakes(ci): Fixed teardown tracking for children spawned during shutdown and made the worker fixture verify its exact window with a Linux-available shell. Both requested checks pass. The lab test takes about 66 seconds locally, so the under-one-minute target remains unmet * no-mistakes(ci): Fixed ci-2 and ci-4 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. Teardown now tracks identity-checked members of captured lab process groups, including children orphaned during shutdown, without signaling the caller’s group or unrelated processes. Lint passed, and the lab test passed four times * no-mistakes(ci): Fixed teardown so an observed-empty process group is permanently dropped, preventing a reused group ID from signalling unrelated work. Added a ps-shim regression test. The lab test, lint, and diff checks pass * no-mistakes(ci): Fixed ci-1 in bin/fm-live-lab.sh and tests/fm-live-lab.test.sh. The TERM-born-child fixture now waits until its handler is installed before calling down. Down sends SIGKILL to identity-valid survivors on every pass from pass 20 onward and includes survivor process details if it must refuse cleanup. bin/fm-lint.sh and tests/fm-live-lab.test.sh pass locally; Linux CI remains to be verified * no-mistakes(ci): Fixed down’s teardown wait to require two empty identity-checked scans separated by 0.5 seconds, and removed the unused test loop variable without changing the TERM-born-child test. The lab test, lint, and diff check pass locally
…unchenguid#6103) * fix(bin): keep slow watcher cycles and preempted reply polls from breaking supervision - fm_pending_reply_tick selects the records it has work for in one awk pass, so settled records cost no lock or fork and the walk no longer grows with the never-pruned store. - An attached arm keeps following a live, identity-matched holder whose beacon went stale until the lock changes or the shared stall bound (fm_watcher_stall_bound), then fails with a typed stalled-holder line so the retry replaces the holder. - The remote-reply adapter reports the job worker's preemption (exit 76) as a closed window, so the listener keeps its claim and polls again instead of being relaunched every watcher cycle. * no-mistakes(document): Clarify watcher grace and attached-arm documentation
…geable is UNKNOWN (kunchenguid#6110) * fix(bin): retry a bounded number of times when GitHub mergeable is UNKNOWN Fixes kunchenguid#6020 bin/fm-pr-merge.sh refused a GitHub merge whenever the pull request's mergeable field was not literally MERGEABLE. GitHub reports UNKNOWN for a short while after a push or a base-branch change while it recomputes mergeability, so a green, conflict-free pull request was refused as if it could not be merged. github_verify_mergeable now returns a distinct status when mergeable is the only failing condition and reads UNKNOWN. The caller retries up to 5 times, 3 seconds apart (overridable in tests), re-reading and re-checking every live condition on each attempt. Once the bound is spent it reports mergeability as still being computed rather than unmergeable, with the same nonzero exit as before. Every other refusal (closed, draft, conflicting, red or missing checks, away authority, queue protection) is unchanged and never retried. * no-mistakes(ci): I fixed both review findings the way you asked. The full suite (`bash tests/fm-pr-merge.test.sh`) ran to completion. Its last lines showed all `ok`, and any failure would have stopped the run early. I watched the output through `tail`, so I didn't see the new test's own `ok` line directly. **ci-2 (`bin/fm-pr-merge.sh`), retry delay not validated.** What must hold: the retry wait is always a short, valid `sleep` argument, so a bad `FM_PR_GITHUB_MERGEABLE_RETRY_DELAY` can never trip `set -e` or hold the task lock for a long time. The retry loop is the only place that reads this variable. The script now reads the value once before the loop and accepts only whole numbers from 0 to 10. Anything else (empty, `abc`, `-1`, `1.5`, `11`, a huge number, leading spaces) falls back to 3. I ran those values through the check by hand and each came out as expected. The loop now sleeps on that checked value. **ci-1 (`tests/fm-pr-merge.test.sh`), no test for a check changing between UNKNOWN reads.** What must hold: every retry re-checks all live conditions, not just mergeable. The fake `gh pr view` in the test can now take an optional second word on each line of the mergeable sequence, which sets the first check's result. The new test `test_github_mergeable_unknown_retry_rechecks_checks` feeds `UNKNOWN`, then `UNKNOWN FAILURE`. It asserts: - exit code 1 after exactly 2 reads, - the refusal names `check 'ci' is not green`, - the message does not say mergeability is still being computed, - `pr merge` was never called. If a later change made the retry look only at mergeable, the loop would read UNKNOWN 5 times, end with the "still being computed" message, and this test would fail. I didn't run it against a deliberately broken script to confirm that. `bash -n` passes. `shellcheck` reports only the existing info-level notes about files it can't follow. Only `bin/fm-pr-merge.sh` and `tests/fm-pr-merge.test.sh` changed
kunchenguid#6112) * fix(bin): converge every open owner onto a known terminal contribution settle_final only cleared a stale error on retry, so an owner whose saved row still said open kept projecting a merged or closed pull request as open after another owner's row had already recorded the terminal observation. Copy the known terminal observation to every owner whose saved row is not itself terminal, keeping that owner's own pending and notified state, and clear its error. * no-mistakes(review): Carry terminal checked_at when converging existing owner rows * no-mistakes(ci): I fixed Greptile finding ci-2 as you asked, with a change to tests/fm-contributions.test.sh only. The rule it enforces: when a retry converges an owner onto a URL that is already merged or closed, that owner gets the terminal owner's whole observation, not just its state. The same weak check appeared twice in test_interrupted_multi_owner_poll_settles_every_owner, so I fixed both: - **Open owner (line 784):** the check now also requires `.observation == $terminal[0].records[0].observation`. The existing checks for error, checked_at, pending and notified are unchanged. - **Errored owner (just below):** it only checked state and error before. It now reads the terminal owner's file and makes the same full-observation comparison. Adding the comparison alone would not have caught anything. The test fixtures gave both owners identical observations apart from `state`, so copying only the state would still have passed. In both cases I also set the terminal owner's observation head to HEAD_B, so the two observations now really differ. Verification: - The focused test passes against the current bin/fm-contributions.sh. - I temporarily changed `settle_final` so it copied only the state. The test then failed, reporting the owner still on the old head (HEAD_A). I restored the file afterwards, and `git status` shows only the test file modified. - The full tests/fm-contributions.test.sh suite exits 0. No product code changed. The other CI finding (ci-1, "Behavior portable serial 9") was left alone because you chose to ignore it
…nguid#6124) * feat: run the supervision host by default on a Claude primary An absent config/supervision-host on a Claude primary now reads as on with the default engine, and a file holding `off` opts any home out. Cursor, OpenCode, omp, Grok, and Codex stay file-gated, with `off` read as disabled there too. Every reader asks fm_supervision_host_enabled instead of testing the file, and non-bash readers query it through the lib's `enabled` entry. A primary's `off` is not inherited by secondmates: each home keeps its own supervision posture. * test: pin the watcher-path posture in fixtures that assume no supervision host Fixtures that drive the watcher arm or assert a non-host drain now write an explicit off file, and fixtures that copy the Stop auto-arm or the supervision instructions carry the engine lib they now source. The two drain suites also stop reading the code root's config. * fix: name the opt-out when an off home passes an attended wake to main A host parked when the home writes off now logs that the home does not run the supervision host, rather than claiming it has no engine. * no-mistakes(document): Clarify Claude supervision defaults and historical evidence * no-mistakes(ci): Fixed process leaks in the two added host tests. Each case now stops its recorded watcher and host/arm processes; fake hook sessions exit through session.stop. The full host suite passed before the final cleanup refinement, and both affected cases, bash syntax, ShellCheck, and diff checks passed afterward. CI runtime still needs confirmation
…start scope check (kunchenguid#6125) * fix(bin): create the state dir on a fresh primary before the session-start scope check fm_primary_scope_matches required an already-existing state directory, so bin/fm-sessionstart-run.sh stood down on a fresh clone before anything could create it. Split out fm_primary_root_matches so the run wrapper can confirm primary-home identity first, create the gitignored state dir when it is missing, and only then run the unchanged scope check. * no-mistakes(document): Document session-start state dir creation on fresh clones * no-mistakes(ci): I fixed the Greptile P1 the way you asked. When a fresh primary can't create `state/`, the run wrapper no longer stands down silently. **Invariant:** when an otherwise eligible fresh primary cannot create `state/`, startup must never fail silently. This path has only one site: the mkdir in `bin/fm-sessionstart-run.sh`. Other hooks and the nudge wrapper never create `state/`, so they have no equivalent failure. **What changed:** - **Run wrapper** (`bin/fm-sessionstart-run.sh`): it captures mkdir's error and prints one line to stderr before standing down as before (exit 0, or 3 for the Pi prerequisite). The line looks like `fm-sessionstart-run: startup could not create the state directory <path>: <reason>`. - **Test** (`tests/fm-sessionstart-nudge.test.sh`): the new case `test_run_reports_a_state_dir_it_cannot_create` uses a fresh primary with no `state/` and a read-only (0500) root. It checks four things: exit 0, no digest on stdout, no state dir created, and exactly one stderr line ending in "Permission denied". It fails without the fix and passes with it. - **Docs** (`docs/sessionstart-nudge.md`): I added one sentence describing the stderr line and one describing what the new test proves. **Verification:** I ran `tests/fm-sessionstart-nudge.test.sh`, and every test passes. `bin/fm-lint.sh` on the changed scripts (pinned ShellCheck 0.11.0) and `tests/fm-documentation-audiences.test.sh` also pass. As you asked, the wrapper still stands down with the ineligible-checkout status afterwards. It does not report this as a failed eligible startup, which is what the bot suggested
…ery (kunchenguid#6126) * fix(bin): measure pending-reply grace from turn completion, not delivery Fixes kunchenguid#6057 The pending-reply guard demanded a repost ("REPOST REQUIRED: previous marked request had no correlated parent report") while the second mate's correlated reply was already on its way. fm_pending_reply_send_recovery measured its grace window from delivery instead of from the request turn's completion, so any turn longer than the grace fired the demand the moment the turn ended, before the reply could have landed. The missed-report escalation had the same gap: it fired the instant the recovery turn's completion was observed, with no grace at all. Both now measure grace from the relevant turn's completion (request turn for the recovery repost, recovery turn for the escalation), and both take one fresh, uncached read of the parent status file immediately before firing, accepting a correlated line regardless of its verb. Transport-failure escalations stay immediate, and the one-repost limit is unchanged. * no-mistakes(review): Document grace window as measured from turn completion * no-mistakes(ci): Both Greptile findings were real and caused by this PR, so I fixed them. The full `tests/fm-pending-reply.test.sh` suite passes. **ci-1 (a reply could be overwritten by a repost).** The rule that must hold: a recovery send is recorded only if the record is still unresolved, checked under the same per-correlation lock that resolution uses. The escalation path already did this (`_fm_pending_reply_maybe_escalate_locked` reads fresh and publishes under one lock). The recovery path did not: `fm_pending_reply_send_recovery` did its fresh read through `fm_pending_reply_try_resolve`, which let go of the lock before the send was recorded. A reply landing in that gap could be overwritten, and the repost would go out anyway. Now `send_recovery` takes the lock once and, while holding it, re-checks that the phase is still `awaiting_report`, runs the fresh uncached read, and records the send (sender pid and identity, attempt time, phase `recovery_sending`). It releases the lock before actually sending, so the lock is not held during the send. It uses the same lock helpers the other lock wrappers use. Grace timing, the one-repost limit and the escalation path are unchanged. **ci-2 (the test would pass even without the fix).** In `test_recovery_fresh_status_read_resolves_before_firing`, the reply is still appended to the status file, but the stored file signature is then set to the file's new signature. That stands in for a same-size rewrite that the signature cache cannot see. The test first checks that a normal cached read misses the reply, then that the fresh read before sending catches it. I also added the same check for the fresh read before escalation, which the review said was uncovered. The test now sets its own send hook, so it no longer depends on one left over from an earlier test (that leftover had made failures exit silently). **Checks:** - I removed the fresh-read bypass at each site in turn and reran the suite. With it gone from recovery, the test fails with "recovery must not fire once a correlated reply has landed". With it gone from escalation, it fails with "the fresh pre-escalation read should have resolved the record, got escalated". With both in place, all tests pass. - Shellcheck with `-x` timed out locally. Without `-x` and ignoring SC1091, the only warnings are SC2034 on the existing `maybe_escalate` lock wrapper, which is not part of this change. The new code adds no warnings. Changes are in `bin/fm-pending-reply-lib.sh` and `tests/fm-pending-reply.test.sh`. Nothing is committed yet; a plain commit message such as "fix(bin): record the pending-reply recovery send under the fresh-read lock" fits the instruction * no-mistakes(ci): ci-1 was real and caused by this PR. The same bug was also in the escalation path, so both are fixed. The full tests/fm-pending-reply.test.sh suite passes. The rule that must hold: a recovery repost or an escalation goes out only if the record's phase, read after the fresh-read resolve, is still what it was before. The resolver writes phase=resolved first and only then writes the other resolution fields. If one of those later writes fails, it returns an error even though the record is already resolved. Places this rule applies, both fixed: - Recovery (fm_pending_reply_send_recovery): the fresh-read resolve now runs first, and the phase is re-read right after it, whatever it returned. The send is recorded and made only if the phase is still exactly awaiting_report. This replaces the earlier phase check rather than adding a second one. - Escalation (_fm_pending_reply_maybe_escalate_locked): same bug. After a failed resolve it went on to publish the blocked line and set phase=escalated. One added line after the resolve call returns 1 without publishing if the phase has changed. Test: added test_partial_resolve_write_blocks_firing. It forces a failure on the resolved_epoch write after a correlated reply has landed. It checks that the recovery send hook is never called, that no escalation line is published, and that the phase stays resolved. The forced failure runs in a subshell so it can't affect later tests. Checks: - With the recovery fix reverted, the new test fails with "recovery must not fire after a partial resolve". - With the escalation fix reverted, it fails with "partial resolve should block escalation, got escalated". - With both fixes in, every test passes. - Shellcheck was run with SC1091 excluded and without -x, not through the repo's lint script. The only new message is one SC2329 info on the test's override function; other test overrides in the same file already get that same info, unsuppressed. Changed files: bin/fm-pending-reply-lib.sh and tests/fm-pending-reply.test.sh. Nothing is committed. Suggested plain commit message: "fix(bin): recheck pending-reply phase after the fresh read before sending
…to stderr (kunchenguid#6001) * fix: provider-table lookup never writes a broken-pipe error to stderr Fixes kunchenguid#5956 fm_quota_single_provider_for_harness returned from its while read loop as soon as it found a match, closing the pipe while fm_quota_single_provider_table's printf could still be writing. Where SIGPIPE is ignored, as on GitHub Actions runners, bash then prints "printf: write error: Broken pipe" on the resolver's stderr, which intermittently broke the one-diagnostic-line assertions in tests/fm-dispatch-resolve.test.sh. Read the whole table before answering, the way fm_control_harness_supported already does, so the writer always finishes. Return values and output are unchanged. Reproduced by running tests/fm-dispatch-resolve.test.sh with SIGPIPE ignored on a single pinned core under CPU contention: 30 of 30 runs failed before the fix, 0 of 30 after. Note: reproducing requires setting the trap inside the tested shell because nice(1) resets an inherited SIGPIPE ignore to SIG_DFL. tests/fm-quota-choose.test.sh passes and bin/fm-lint.sh is clean. * no-mistakes(ci): Fixed both Greptile findings the user chose to address. ci-1 (bin/fm-quota-axi-lib.sh:154). Invariant: looking up a harness must always end with status 0 and print the provider, even when the caller runs under `set -e`. The loop body `[ -z "$found" ] && [ "$harness" = "$1" ] && found=$provider` now ends in `|| :`. Every iteration succeeds and the whole table is still read. Only `fm_quota_single_provider_for_harness` loops over the table this way, so this is the one place the fix was needed. One caveat: on bash 5.3 the old code did not actually exit under `set -e`, because the `while` loop is not the function's last command, so the new `set -e` test would have passed before this fix too. The change makes the loop's success explicit, as the user asked. ci-2 (regression coverage). I added three cases to the existing `tests/fm-quota-choose.test.sh`, all calling the public lookup function after sourcing the library: 1. With SIGPIPE ignored (`trap "" PIPE`), it looks up every harness 200 times and checks that nothing reaches stderr. 2. A deterministic version of the race: the table function is wrapped so it writes the first row, pauses 0.2 s, then writes the rest. With SIGPIPE ignored, it checks that looking up `claude` prints `claude` and writes nothing to stderr. The stress loop alone reproduced the bug in only about 1 of 5 local runs, which is why this case exists. 3. A direct call under `set -e` prints `claude`. Verification: - `bash tests/fm-quota-choose.test.sh`: all pass. - Same test against the pre-PR library (fa48367, via `FM_ROOT_OVERRIDE`): fails with `printf: write error: Broken pipe`. The deterministic case failed in one run and the stress loop caught it in another. - `shellcheck` on both files: clean. - `tests/fm-dispatch-resolve.test.sh`: passes
…isioning (kunchenguid#6162) * fix: survive Pi 0.99 rendering and Git 2.55 local-clone races Pi 0.99 puts arguments on the stock tool header and leaves hidden custom messages in the export conversation column. Match that header, and keep Calm's boundary on the visible column. Clone a remote home with --no-local so a prune during Git's loose-object copy cannot fail the seed. * no-mistakes(review): Stop SIGPIPE write errors; cover older Pi export and project clones * no-mistakes(document): Clarify Calm export visibility and tool rendering * no-mistakes(ci): Fixed the dispatch diagnostic to list every provider-less use/default profile in one line and added a multi-profile behavior test. Shortened supervision fixtures using the existing engine-grace and park-clock knobs; removed stray scratch files. Dispatch tests, syntax checks, and three targeted supervision cases passed. CI’s prior supervision duration was 751s; the single permitted local full-suite run timed out at 1200s, so an after-duration is not established. The cancelled serial check had no failure verdict. The outer executor should record the measured before/after duration in the PR body when available * no-mistakes(review): Gate Pi 0.99 call headers by version; drop hidden-row assertion * no-mistakes(review): Test stock call headers under Pi 0.87 and 0.99 stubs * no-mistakes(test): Fix older-Pi queued-row test and verify park-boundary behavior * no-mistakes(document): Clarify Pi Calm export and queued-turn documentation * no-mistakes(ci): Fixed the stock macOS Bash 3.2 parse failure in tests/fm-calm-pi-extension.test.sh; its parse check passes. The watcher CI failure is in unchanged code: the isolated five-minute/66-minute case passes locally, but the CI log omits the drain error needed to establish its cause. No speculative watcher fix was made. The full local watcher suite timed out after 500 seconds
…6169) * Prevent premature Lavish board handoffs * Prove Lavish arm lacks reply acknowledgement * Confirm Lavish replies before arming worker boards * no-mistakes(review): Post Lavish reply only after locked arm eligibility checks * no-mistakes(review): Fail Lavish reply closed on unknown version * no-mistakes(document): Correct Lavish reply documentation and remove stale guidance * no-mistakes(document): Clarify Lavish reply routing and remove duplicate version guidance
…#6154) * feat: inherit the supervision-host opt-out from the primary Move the supervision host's off opt-out out of config/supervision-host into its own presence flag, config/supervision-host-off, and add that flag to the primary-authoritative inherited config set. A primary that opts out now opts every secondmate home out at spawn and convergence, and clearing it converges them back. config/supervision-host stays the home-local engine choice. Shape: config/supervision-host mixed two things, a fleet posture (off) and a per-home engine and model. Only the posture should follow the primary, so it becomes a separate presence flag that rides the existing inherited-config mechanism (FM_INHERITABLE_CONFIG in bin/fm-config-inherit-lib.sh) with no new machinery, while the engine line stays local. The parse stays in its one owner, fm_supervision_host_enabled. There is no migration or compatibility handling for a home that still holds off in config/supervision-host. Primary off, mate on: inherited material is primary-authoritative by design, so a mate cannot keep the host while the primary is opted out, and a mate's own opt-out is removed at the next convergence while the primary has none. Running the host on a mate is the primary's choice for the fleet; no override mechanism is added. Live validation (disposable bin/fm-live-lab.sh lab, Claude primary with a real seeded secondmate, --supervision-host off): - up: every readiness check ok, including "host: none running, as expected" and a live mate session; the spawned mate home held the inherited config/supervision-host-off and the gate read primary OFF, mate OFF. - primary removed its opt-out, then bin/fm-config-push.sh reported "supervision-host-off: pushed - mirrored primary absence" and a config reread sent; the gate read primary ON, mate ON, and the live mate handled the reread. - primary opted out again and pushed: "supervision-host-off: pushed", mate gate OFF. - down stopped every lab process and left no lab process running. Out of scope, follow-up: default-on for the other harnesses, away-daemon retirement, rollout. * no-mistakes(document): Document inherited supervision-host opt-out ownership * no-mistakes(ci): Fixed ci-4: with `--supervision-host off --mate`, lab readiness now requires the inherited flag in the mate home and a disabled mate supervision-host gate. The focused behavior test, shellcheck, and diff checks pass. Left ci-1–ci-3 untouched as directed * no-mistakes(test): Fix mate readiness HOST_OFF initialization in lab up * no-mistakes(ci): Fixed Lint 2 by making the new test’s fixtures source resolvable to ShellCheck; its off/on readiness test and ShellCheck now pass locally. Behavior portable serial 5 failed in the unchanged remote-reply test at generation 7. That test passes locally, and no PR-caused defect was identified, so no remote-reply code was changed
…nguid#6179) * fix(tests): cut the fixed sleeps in supervision-host cycles The serial CI lane keeps brushing its 30-minute cap because fm-supervision-host.test.sh spends ~903s of the job, and per the run-36635306527 case profile the top nine cases are all multi-cycle ones (3-10 park/close/turn cycles each): every close waits out the host's sleep $POLL in await_close plus a watcher sleep $FM_POLL scan cycle, and every engine turn waits out the fixed sleep 1 descendant snapshot. That is ~3s of pure sleep per cycle before any real work. The host poll now accepts positive decimal seconds through a new seconds_or validator (FM_SUPERVISION_HOST_POLL), and the engine turn's snapshot loop takes FM_SUPERVISION_ENGINE_SNAPSHOT_SECONDS, also a positive decimal defaulting to one second - the smallest seam at each wait's single owner. The suite drives them at 0.2 alongside the existing FM_POLL=0.5 and FM_ARM_ATTACH_POLL=0.2 knobs, so the real poll loops still run. The park-boundary case moves onto the injected test clock instead of a real 3s wait, per-case cleanup polls the host pid rather than sleeping a full second, and the proof-by-absence windows (flood re-escalation, successor re-announce, watcher persistence, recovery staying off main) shrink from 2-3s to 1s, which still spans two watcher polls at the test cadence. Every assertion, process lifecycle, and reaping path is unchanged; production defaults stay at one second. Isolated case timings on a contended host, base vs branch: attended-latch 54.3->34.6s, undelivered-dialog 67.7->59.1s, away-latch 46.5->30.5s, held-cadence 47.9->21.6s, unreadable-mirror 39.2->38.5s, park-limit 18.2->12.3s, registration-fallback 14.1->10.0s, first-cycle-status 12.6->8.4s, latch-scope 16.7->16.3s. Full suite: 65/65 pass. fm-lint and shellcheck clean. * no-mistakes(review): Wait for scan lock release before duplicate check * no-mistakes(document): Correct supervision snapshot cadence documentation * fix(tests): keep production poll cadence, probe exits at 0.1s The fractional poll cadences multiplied the cost of each loop body: full process-table scans in the engine turn and process refreshes in await_close ran five times more often, which swamped the thin CI runner and nearly doubled every multi-cycle case (serial 5 was cancelled at its 30-minute limit on run 36635306527's successor). Restore the production cadence and notice arm/engine exits with a cheap kill -0 probe at a tenth of a second between the one-second bodies instead: strictly less dead time than baseline with no added CPU. Also hold each injected-clock park bound well past its case's wall-clock checks so a host that ignored the test clock fails instead of silently passing at a real-time boundary, and restore the shortened proof windows (watcher liveness, recovery-off-main absence, first-cycle stream) to their baseline depth. * no-mistakes(document): Clarify supervision engine snapshot documentation
…henguid#6192) * fix: rebalance portable CI from current duration measurements * no-mistakes(test): Test serial packing boundary and verify endpoint timeout cleanup * no-mistakes(document): Clarify timeout guidance and remove duplicated packing estimates
…nguid#6216) * fix(bin): run no repository hook when core.hooksPath is empty The per-task hook wrapper refused every commit in a repository whose own config sets core.hooksPath to the empty string, because git rev-parse --git-path hooks fails on it. Plain git reads that setting as no hooks, so the wrapper now runs none; every other lookup failure still refuses and shows git's error. Fixes kunchenguid#6171 * no-mistakes(review): Refuse commits when core.hooksPath is a valueless key * no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs * no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files
…kunchenguid#6213) * fix(bin): let a stale record on a reassigned slot retire records-only When a pool slot's owner claim names another task, the stale record's teardown touches nothing under the slot, so the exclusive-slot record scan no longer refuses it. Full teardowns of a slot this task still claims, or one with no claim, keep the refusal. Fixes kunchenguid#6184 * no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots
…uid#6240) * fix(bin): keep the steering doorbell short under deep homes The doorbell printed the task inbox's absolute path twice, so under a deep home it grew to about 290 characters and a Herdr submit reported it never reached the pane on every re-ring. It now names the inbox once by its short <task>.inbox name and points at the full path the worker's brief already gives, so its length no longer depends on the home's depth. Fixes kunchenguid#6120 * no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell * no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh
Brings in four upstream fixes: - kunchenguid#6192 ci: rebalance portable test groups and enforce a packing budget - kunchenguid#6213 fix(bin): let a stale record on a reassigned slot retire records-only - kunchenguid#6216 fix(bin): run no repository hook when core.hooksPath is empty - kunchenguid#6240 fix(bin): keep the steering doorbell short under deep homes Naive merge would pick a stale history merge base because fork main carries the #11 and #14 syncs as squashes; an explicit content merge base of b3d4133 (last upstream commit whose content fork main shared) is used instead. Two genuine conflicts (bin/fm-teardown.sh, docs/architecture.md) were 3-way merged: upstream kunchenguid#6213's records-only code path survives, fork #13's claim-first semantics and endpoint-cleared behavior are kept, and the overlapping header/doc prose keeps the fork's newer wording. Content delta vs fork main is exactly the 22 files of kunchenguid#6192+kunchenguid#6213+kunchenguid#6216+kunchenguid#6240 (docs/architecture.md is byte-identical to fork main after resolution).
…Path refusal cases
keenvc
force-pushed
the
fm/fm-upstream-sync-2026-09-30-2
branch
from
October 1, 2026 01:09
c69b16d to
b9ebe4c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Second fork sync PR for keenvc/firstmate: bring in four upstream fixes merged since PR #14's sync point (kunchenguid#6192 rebalance portable test groups and packing budget, kunchenguid#6213 stale record on a reassigned slot retires records-only, kunchenguid#6216 run no repository hook when core.hooksPath is empty, kunchenguid#6240 keep the steering doorbell short under deep homes). The user chose: merge PR #14 first, then raise this second small sync. Fork main carries the #11 and #14 syncs as squashes, so the naive history merge base is stale; the merge uses an explicit content merge base b3d4133. Two genuine conflicts (bin/fm-teardown.sh, docs/architecture.md) were 3-way resolved: upstream kunchenguid#6213's records-only code hunk survives, the fork's #13 claim-first semantics and endpoint-cleared behavior are kept, and the overlapping header/doc prose keeps the fork's newer wording. A prior run validated the merge content (teardown suite 27/27, kunchenguid#6192 packing budget verified, kunchenguid#6216 intent verified) and committed a test-fix (probe git before the broken-hooksPath refusal cases) before its test agent hit the Claude session limit; this run resumes from that head.
What Changed
kunchenguid#6192 (ci: rebalance portable test groups and enforce a packing budget),
kunchenguid#6213 (stale record on a reassigned slot retires records-only),
kunchenguid#6216 (run no repository hook when core.hooksPath is empty), and
kunchenguid#6240 (keep the steering doorbell short under deep homes).
The merge uses an explicit content merge base
b3d413323because fork maincarries the #11 and #14 syncs as squashes, which makes git's history merge
base stale. Two genuine conflicts (
bin/fm-teardown.sh,docs/architecture.md) were 3-way resolved: upstream kunchenguid#6213's records-onlycode hunk survives, the fork's #13 claim-first semantics and
endpoint-cleared behavior are kept, and the overlapping header/doc prose
keeps the fork's newer wording. One pipeline test fix (
fm-git-strip-ai-trailers.shprobes git before its broken-hooksPath refusal cases) and one documentation
correction were added by the validation pipeline. Content delta vs fork main
is exactly these 21 files:
Risk Assessment
✅ Low: The change is a faithful, verified-complete sync of four upstream fixes whose content delta is exactly the 21 expected files, both conflict resolutions leave code and prose consistent, every consumer of the changed launch text and doorbell format was updated, and the new packing refusal and empty-hooksPath branch both behave correctly under a concrete traced case.
Testing
I derived the scenarios from the four upstream fixes named in the intent and drove each against the real scripts rather than relying on suite exit codes. kunchenguid#6240 is the headline result and is proven three independent ways: the real doorbell line drops from 623 to 166 characters and becomes constant across home depth; in a live 80-column tmux pane that is 9 wrapped rows down to 3, which is precisely the wrapping failure the fix cites; and the repo's own live e2e put the new line in front of real agents launched with no brief at all, where claude, opencode and pi each resolved the inbox from $FM_TASK_INBOX, acted, and acknowledged with the mv. I also proved the load-bearing coupling that spawn really exports FM_TASK_INBOX for ship and secondmate launches even under the cleared allowlist. kunchenguid#6216 reproduces as a true regression pair: a real commit with an empty core.hooksPath is refused outright on pre-fix code and succeeds on this branch with the AI trailer stripped, the human trailer kept, and no repository hook run; adversarially, a broken non-empty hooksPath still refuses with HEAD unmoved. kunchenguid#6192 holds at the runner's own CLI, where the nine CI shards form an exact gap-free, duplicate-free partition of the 211-script serial lane with the largest shard inside budget. kunchenguid#6213's behaviour holds through the real teardown CLI, though reverting the adopted hunk leaves the case passing, confirming the redundancy the review round already recorded and declined. Two live harnesses failed: codex and grok never honored the doorbell, but an A/B against the fully reverted pre-sync code shows both fail identically there, and grok's own pane reports its weekly limit is exhausted, so neither is merge fallout and I recorded them as untested rather than guessing. The two fm-test-run failures are likewise the previously recorded host issue, confirmed by checking that the implicated test file is outside this diff and that this host reparents orphans to a systemd subreaper instead of PID 1. This change has no graphical surface, so the reviewer-visible artifacts are terminal captures of the actual pane rendering rather than screenshots. All lab directories were removed and the worktree carries no changes from my testing.
git committhrough an installed wrapper: pre-fix exit 1 with HEAD unmoved ('cannot resolve this repository's hooks directory'), this branch exit 0 with the ant…Evidence: kunchenguid#6240 doorbell length vs home depth, pre-sync vs this branch (623 -> 166 chars, now constant)
Source: #6240 doorbell length vs home depth, pre-sync vs this branch (623 -> 166 chars, now constant)
Evidence: kunchenguid#6240 the doorbell rendered in a LIVE 80-col tmux pane: 9 wrapped rows before vs 3 after, plus dead-pane inertness and $FM_TASK_INBOX resolution
Source: #6240 the doorbell rendered in a LIVE 80-col tmux pane: 9 wrapped rows before vs 3 after, plus dead-pane inertness and $FM_TASK_INBOX resolution
Evidence: kunchenguid#6240 LIVE real-agent doorbell e2e: claude, opencode and pi acted and acked with no brief; codex and grok did not
Source: #6240 LIVE real-agent doorbell e2e: claude, opencode and pi acted and acked with no brief; codex and grok did not
Evidence: kunchenguid#6240 A/B conclusion: codex and grok fail identically on pre-sync code; grok's pane reports 'Weekly limit left: 0%'
Source: #6240 A/B conclusion: codex and grok fail identically on pre-sync code; grok's pane reports 'Weekly limit left: 0%'
Evidence: kunchenguid#6240 A/B raw run against fully reverted pre-sync lib, spawn and test (codex + grok)
Source: #6240 A/B raw run against fully reverted pre-sync lib, spawn and test (codex + grok)
Evidence: kunchenguid#6216 real git commit with empty core.hooksPath: pre-fix REFUSES the commit, post-fix commits with the AI trailer stripped and no repository hook run
Source: #6216 real git commit with empty core.hooksPath: pre-fix REFUSES the commit, post-fix commits with the AI trailer stripped and no repository hook run
Evidence: kunchenguid#6216 adversarial: unresolvable and valueless core.hooksPath still refuse, HEAD does not move, no hook silently skipped
Source: #6216 adversarial: unresolvable and valueless core.hooksPath still refuse, HEAD does not move, no hook silently skipped
Evidence: kunchenguid#6216 full git-strip suite 16/16 with the committed test-fix's honest env-probed skips
Source: #6216 full git-strip suite 16/16 with the committed test-fix's honest env-probed skips
Evidence: kunchenguid#6192 the 9 CI serial shards are an exact disjoint cover of all 211 serial scripts (0 gaps, 0 duplicates)
Source: #6192 the 9 CI serial shards are an exact disjoint cover of all 211 serial scripts (0 gaps, 0 duplicates)
Evidence: kunchenguid#6192 coverage guard and packing budget: serial_max_ms 1074843 within serial_budget_ms 1200000, exact-budget boundary case passes
Source: #6192 coverage guard and packing budget: serial_max_ms 1074843 within serial_budget_ms 1200000, exact-budget boundary case passes
Evidence: kunchenguid#6192 lane composition from bin/fm-test-run.sh --list-lanes and --check-coverage
Source: #6192 lane composition from bin/fm-test-run.sh --list-lanes and --check-coverage
Evidence: kunchenguid#6213 teardown endpoint-safety 28/28 including the stale-record-on-a-claimed-slot case
Source: #6213 teardown endpoint-safety 28/28 including the stale-record-on-a-claimed-slot case
Evidence: kunchenguid#6213 control run with the hunk reverted: the case still passes, confirming it is a no-delta on this fork
Source: #6213 control run with the hunk reverted: the case still passes, confirming it is a no-delta on this fork
Evidence: kunchenguid#6240 spawn exports the absolute steering inbox as FM_TASK_INBOX for ship and secondmate, proven by executing the emitted launch
Source: #6240 spawn exports the absolute steering inbox as FM_TASK_INBOX for ship and secondmate, proven by executing the emitted launch
Evidence: fm-teardown full suite 107/107, 0 failures
Source: fm-teardown full suite 107/107, 0 failures
Evidence: Control: the 2 fm-test-run failures are the recorded host issue (test file absent from this diff; host reparents orphans to a systemd subreaper, not PID 1)
Source: Control: the 2 fm-test-run failures are the recorded host issue (test file absent from this diff; host reparents orphans to a systemd subreaper, not PID 1)
Pipeline
Updates from git push no-mistakes
... (10 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)
bash tests/fm-git-strip-ai-trailers.test.sh- 16/16 pass, including the committed test-fix's env-probed skips for the two broken-hooksPath refusal cases on this host's git 2.43.0Manual live drive: realgit commitwithcore.hooksPath=""against an installed wrapper, pre-fix vs post-fix, with a canary repository commit-msg hook to detect silent delegationManual live drive: adversarial unresolvable (~nonexistent-user-xyz/hooks) and valuelesscore.hooksPath, asserting the commit is refused and HEAD does not moveManual live drive: real doorbell line generated from bin/fm-task-inbox-lib.sh at a shallow vs 236-char-deep state dir, pre-sync lib vs this branch, comparing line lengthManual live drive: the real doorbell typed into a live 80x24 tmux pane on a private socket, counting wrapped terminal rows pre-sync vs post-sync; then Enter in the bare (dead-agent) pane to confirm inertness; thenls "$FM_TASK_INBOX"/*.msgin that same pane to confirm the worker can resolve its inboxFM_SEND_INBOX_LIVE_E2E=1 FM_SEND_INBOX_LIVE_TIMEOUT=200 bash tests/fm-send-inbox-doorbell-live-e2e.test.sh- real agents, no brief, isolated tmux server: claude 2.1.286 PASS, opencode v2.0.19 PASS, pi 0.99.1 PASS, codex 0.118.0 FAIL, grok 1.0.44 FAILA/B control: the same live e2e with bin/fm-task-inbox-lib.sh, bin/fm-spawn.sh and the test reverted to pre-sync 4aecf233,FM_SEND_INBOX_LIVE_HARNESSES='codex grok'- both fail identically, proving the two failures are not merge falloutbash tests/fm-spawn-compact-adviser-disable.test.sh- 7/7, including the new case that executes the emitted launch with an env probe to prove ship and secondmate agents start with FM_TASK_INBOX set, under the cleared launch-env-allowlist posturebash tests/fm-task-inbox.test.sh- 21/21, including the hostile-inbox-name no-op and terminal-control rejection casesbash tests/fm-send-inbox.test.sh- 14/14bash tests/fm-teardown-endpoint-safety.test.sh- 28/28, includinga stale record on a claimed slot retires, then the claimant tears down(#6213)bash tests/fm-teardown.test.sh- 107/107, 0 failuresRegression control: the same teardown endpoint-safety suite run against bin/fm-teardown.sh reverted to pre-#6213, showing the adopted hunk is a no-delta on this fork's claim-first orderingbin/fm-test-run.sh --check-coverage- total=251, serial_max_ms=1074843 within serial_budget_ms=1200000bin/fm-test-run.sh --list --lane portable-serial-<k>of9for k=1..9, set-compared against--list --lane portable-serialto prove an exact disjoint cover of all 211 scriptsbash tests/fm-test-run.test.sh- 49 pass, 2 pre-existing host failures; all six #6192 packing cases pass includingserial packing accepts the exact budget and refuses one millisecond above itHost-environment control for those 2 failures: confirmed tests/fm-session-lock-ancestry.test.sh is absent from the sync diff and that this host reparents orphans to a systemd --user subreaper (ppid 1285), not PID 1docs/fm-test-portable-shards.md:82- Judgment call, recorded for awareness rather than action: to resolve the contradiction created by correcting the coverage claim, the hint-refresh snippet's-R kunchenguid/firstmatewas generalized to-R <owner>/firstmate. That is a deliberate one-token divergence from upstream in a PR whose stated intent is to keep upstream drift cheap. The alternative - leaving the snippet pinned to upstream - would tell a fork maintainer to download artifacts from runs that never execute the fork-only serial members the same section now says need measuring. If the fork prefers zero drift here over internal consistency, reverting this one line is safe; the corrected coverage sentence above stands on its own either way.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.