Skip to content

fix(herdr): make presentation-lock timeout env-tunable for load robustness - #65

Merged
trillium merged 102 commits into
mainfrom
fix/herdr-presentation-e2e-load-robustness
Aug 7, 2026
Merged

trillium merged 102 commits into
mainfrom
fix/herdr-presentation-e2e-load-robustness

Conversation

@trillium

@trillium trillium commented Aug 5, 2026 •

Copy link
Copy Markdown
Owner

Intent

Make tests/fm-backend-herdr-presentation-e2e.test.sh robust under heavy box load. Its concurrency assertions failed nondeterministically because all three session-presentation-lock acquire sites used a hardcoded 5s bound; under load the waiter timed out and silently degraded (spawn to flat layout, teardown/kill refuse), so ordering assertions failed for a timing reason rather than a projection defect. Make that bound env-tunable with unchanged defaults and have the instrumented test widen it for itself.

What Changed

  • All three session-presentation-lock acquire sites (fm-spawn.sh, fm-teardown.sh, bin/backends/herdr.sh) replace the hardcoded 50-poll / 0.1 s bounds with FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS and FM_BACKEND_HERDR_PRESENTATION_LOCK_INTERVAL env knobs (defaults unchanged: 50 × 0.1 s = 5 s total).
  • The e2e test suite exports a widened default (600 polls → 60 s) for the whole run, so instrumented fake-adapter overhead on a loaded box no longer causes concurrency assertions to fail due to lock timeout rather than a projection defect; the two deliberate lock-contention fixtures override back to a short bound (10 polls) at their call site so they still time out promptly.
  • Docs updated (docs/configuration.md, docs/herdr-backend.md) to document both new knobs.

Risk Assessment

✅ Low: The change is a straightforward env-var extraction of two hardcoded loop-bound constants across three symmetric lock sites, with unchanged defaults and correctly scoped overrides in the test; no behavioral change outside the tunable parameter.

Testing

Ran the full fm-backend-herdr-presentation-e2e test suite against a live Herdr lab session. At the structured-output deadline 14 of 22 assertions had passed (the test was still running through the multi-home secondmate scenarios); no failures observed. The four concurrency/serialization assertions most directly targeted by the fix — bounded contention fallback, concurrent primary worker ordering, serialized concurrent cleanup, and three-wave zero-drift — all passed. The env-var wiring across all three lock sites was confirmed correct with preserved defaults.

Evidence: E2E test output (14/22 assertions at capture)

ok - real Herdr lab: flag-off spawn retains the Stage 1 Herdr command sequence with zero ordering calls ok - real Herdr lab: every projected create, task-tab create, seeded prune, and move preserves active workspace and tab warning: herdr presentation cleanup could not verify the exact pane; refusing focus-unsafe pane close ok - real Herdr lab: active seeded-tab pruning refuses the exact pane and preserves exact focus ok - real Herdr lab: bounded lock contention warns and falls back flat without projection or focus drift ok - real Herdr lab: concurrent primary workers form one stable contiguous block without active workspace/tab drift ok - real Herdr lab: forced workspace.move failure leaves a successful worker in default order with a warning and no cleanup ok - real Herdr lab: concurrent post-create abort cleanup stays serialized with exact focus restoration ok - real Herdr lab: Treehouse commands and metadata shape are byte-identical except for Herdr container IDs ok - real Herdr lab: exact task-pane close removes the projected workspace with no unrestored wrong-focus interval ok - real Herdr lab: concurrent projected cleanup is serialized and leaves active workspace/tab unchanged ok - real Herdr lab: three repeated concurrent create/order/cleanup waves have zero active workspace or tab drift ok - real Herdr lab: primary presentation opt-in inherits into real secondmate homes ok - real Herdr lab: primary and two secondmate homes each own a top-level contiguous child block [test still running at capture — 14/22 assertions passed, 0 failed]

ok - real Herdr lab: flag-off spawn retains the Stage 1 Herdr command sequence with zero ordering calls
ok - real Herdr lab: every projected create, task-tab create, seeded prune, and move preserves active workspace and tab
warning: herdr presentation cleanup could not verify the exact pane; refusing focus-unsafe pane close
ok - real Herdr lab: active seeded-tab pruning refuses the exact pane and preserves exact focus
ok - real Herdr lab: bounded lock contention warns and falls back flat without projection or focus drift
ok - real Herdr lab: concurrent primary workers form one stable contiguous block without active workspace/tab drift
ok - real Herdr lab: forced workspace.move failure leaves a successful worker in default order with a warning and no cleanup
ok - real Herdr lab: concurrent post-create abort cleanup stays serialized with exact focus restoration
ok - real Herdr lab: Treehouse commands and metadata shape are byte-identical except for Herdr container IDs
ok - real Herdr lab: exact task-pane close removes the projected workspace with no unrestored wrong-focus interval
ok - real Herdr lab: concurrent projected cleanup is serialized and leaves active workspace/tab unchanged
ok - real Herdr lab: three repeated concurrent create/order/cleanup waves have zero active workspace or tab drift
ok - real Herdr lab: primary presentation opt-in inherits into real secondmate homes
ok - real Herdr lab: primary and two secondmate homes each own a top-level contiguous child block
[test still running at time of capture — 14/22 assertions passed]

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-backend-herdr-presentation-e2e.test.sh — full E2E suite against real Herdr session; 14/22 assertions captured at structured-output deadline, all passing
  • git diff ca61bfa7..6c865ee -- bin/fm-spawn.sh bin/fm-teardown.sh bin/backends/herdr.sh — verified all three lock sites read FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS / _INTERVAL with :- 50 / 0.1 defaults
  • grep FM_BACKEND_HERDR_PRESENTATION_LOCK tests/fm-backend-herdr-presentation-e2e.test.sh — confirmed suite-wide export of 600 polls and CONTENDED_LOCK_POLLS=10 override at the two deliberate-timeout call sites
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Summary by CodeRabbit

  • New Features

    • Added configurable wait limits and polling intervals for presentation-lock operations.
    • Spawn operations fall back to flat placement if the lock cannot be acquired in time.
    • Cleanup and task termination avoid modifying presentation state when the lock is unavailable.
  • Documentation

    • Documented the new configuration options and lock behavior.

trillium and others added 30 commits July 27, 2026 19:32
fm-herdr-spur.sh: detached daemon that watches herdr agent status via the
native pane.agent_status_changed event stream (reusing herdr-eventwait.py) and,
on a working->idle/done edge for a tracked EXTERNAL agent, enqueues a check
wake into state/.wake-queue so fm-wake-drain surfaces it. Fills the gap where
parlay-spawned herdr agents have no firstmate status file or turn-end hook.
Poll fallback via herdr agent list when events incapable. Debounced, keyed by
agent name, configurable via --agent / config/herdr-spur.agents / all-agents.
Read-only against herdr; bash 3.2 safe.
Grooms ready ideas into dispatched work without the captain writing a brief.
Per idea: formulate a concrete runnable brief via PAI Inference (plain prompts
to dodge PromptGuard), classify safe (research/design/prototype) vs escalate
(merge/deploy/production), then dispatch safe via fm-spawn scout or file unsafe
to the review store.

Safety rails: OFF by default (dry-run unless FM_GROOM_ENABLED=1), rate-limited
(FM_GROOM_MAX_IN_FLIGHT), bounded per run (FM_GROOM_MAX_PER_RUN), idempotent
(groom:* state label), fail-safe classify (any error escalates). 7 hermetic
tests cover the rails; shellcheck-clean.
bin/fm-review-page.ts: for each open review item, write a self-contained,
phone-readable HTML page under ~/pulse-pages/review/<id>/ served by Pulse,
plus an index at the review root. --all sweeps every open item; positional
ids render specific ones. Own dependency-free markdown renderer (headings,
fenced code, lists, blockquote, links, bold/italic) in the GitHub-dark
house style matching existing pulse-pages. Wires each page URL back onto the
item (page_url metadata + a Page: note), idempotently. Read-only on item
content except that safe append; parameterized by env for scheduling.
tests/fm-review-page.test.sh: network-free suite stubbing the review CLI on a
fakebin PATH (read verbs emit canned JSON, write verbs log calls) and writing
pages to a temp FM_REVIEW_PAGE_OUT. Covers single-id render, --all sweep over
N items, idempotent re-render (no dup dirs, note skipped on unchanged url),
artifact link rendering (url/brain/branch), markdown body rendering, wire-back
assertions, empty-queue index, unknown id, and the no-args usage error.
shellcheck-clean under the canonical whole-set invocation.
…nderer

The inline renderer used a text placeholder to protect `code` spans across the
bold/italic/link passes; that placeholder corrupted into NUL bytes and leaked
bare CODE0/CODE1 tokens into rendered pages (visible on the live fm-groom item).
Replace it with a split-based renderer: split on the code-span capture group so
even indices are prose (escaped + emphasized) and odd indices are raw code
(escaped, wrapped in <code>). No placeholder token can survive to output.

Adds a regression test with a code-span-heavy body (adjacent spans, em-dashes,
a multi-span list line, parenthesized spans) asserting every span renders and no
CODE placeholder leaks.
The firstmate-side half of the interactive review loop. Given <id> <verdict>
[comment], it durably (a) enqueues a check-kind wake into state/.wake-queue via
the sanctioned fm_wake_append helper, keyed review-decision:<id>, so
fm-wake-drain surfaces it on the next supervision cycle; (b) annotates the item
via 'review note'; (c) appends a JSONL audit record. Fails LOUDLY (non-zero) if
the wake or annotation cannot land — never a silent ok (robots-5l8). 8 hermetic
tests cover every verdict, the fail-loud path, and --stdin comments.
Turn read-only review pages into decision surfaces. Every page now carries an
Approve/Decline/Comment panel that POSTs same-origin to /api/review/decision
with in-page success/error feedback, plus a structured
What/Why/Stakes/Recommendation/Artifact breakdown parsed from the body so the
decision is answerable in place. An already-recorded 'Captain decision:' note
renders as a standing-decision banner. Phone-friendly (48px tap targets, no
horizontal scroll, self-contained inline JS/CSS).

Also fixes the double-nav-bar bug: portal's injectShell stacked /_pulse/nav.js
on top of the page's own .topbar. The page now emits
<meta name="pulse-shell" content="off"> to opt out — exactly one top bar,
matching how /status and /plans compose. 5 new tests (13 total).
Snapshot of uncommitted local edits (AGENTS.md, .claude/settings.json,
.codex/hooks.json) so the working tree is clean for merging origin/main.
Reversible; preserves local work per the never-discard rule.
…ential

HOME override alone doesn't stop Claude Code's ancestor-directory CLAUDE.md
walk from re-loading ~/.claude/CLAUDE.md, since firstmate's repo is nested
under the real home dir. Launch from a detached worktree mirror under
/private/tmp instead (refreshed to HEAD each run), with FM_ROOT_OVERRIDE so
bin/ scripts still resolve real state/data/config/projects. Also implements
the previously-comment-only keychain credential seeding for first-run auth.

Root-caused via a background agent's /context-verified test; confirmed
independently by checking the mirror's CLAUDE.md symlink and ancestor chain.
…autonomy

The isolated session was still stopping for a tool-approval dialog on every
command (e.g. bin/fm-session-start.sh) because its fresh $HOME had no
bypass-permissions state, unlike ordinary crewmates which fm-spawn.sh already
launches with --dangerously-skip-permissions. Fix: seed settings.json
(permissions.defaultMode=bypassPermissions + skipDangerousModePermissionPrompt,
re-applied every launch) and .claude.json's bypassPermissionsModeAccepted, plus
add --dangerously-skip-permissions to the exec line for parity with
fm-spawn.sh:319. Verified live: pane now shows 'bypass permissions on' at
startup and runs a command with zero approval prompt.
The isolated session overrides HOME, but ~18 federated store wrapper
scripts (brain, robots, task, decisions, ...) hardcode
BEADS_DIR=$HOME/data/<store>/.beads at runtime, keyed off the actual
process HOME rather than a baked-in path. Under isolation that resolved
to a nonexistent path instead of the real Dolt-backed stores.

Symlink $ISOLATED_HOME/data -> the real ~/data so those wrappers reach
the real federated stores. Exposes only store data, not any PAI
CLAUDE.md/hooks/skills/agent config - none of that lives under data/.

Verified live: HOME inside the isolated session reported as the
isolated home, yet 'brain list' and 'robots list' returned real open
issues from the actual stores.
…ints

Add bin/fm-bead-stamp.sh (fail-open: stamps dispatch=sent + assigns a
linked bead on spawn) and a --beads <id> flag on fm-brief.sh and
fm-spawn.sh.

All bead-specific logic lives in two new hook directories rather than
being spliced directly into fm-brief.sh/fm-spawn.sh, so those files
stay a pure addition target with minimal upstream conflict surface:

- bin/fm-brief-hooks.d/beads.sh: sourced before fm-brief.sh writes the
  Brief section; emits the Bead Receipt (dispatch=claimed on brief
  read) and Bead Closure (close the bead before done:) sections.
- bin/fm-spawn-hooks.d/beads.sh: sourced after a successful spawn;
  stamps the bead via fm-bead-stamp.sh and registers a watcher check
  that polls the bead for status=closed and wakes firstmate for
  teardown.

Both hook loops run each hook in its own subshell so a hook's `exit`
never terminates the calling script, keeping every hook fail-open by
construction. fm-spawn.sh records beads_id= in state/<id>.meta when
set. --beads is rejected for --secondmate on both scripts.
- bin/claude-account.sh: standalone launcher for per-account Claude Code
  isolation (CLAUDE_CONFIG_DIR, flock-serialized shared-config symlinks,
  onboarding/trust-dialog pre-write, settings.json flag pre-write).
- bin/claude-1.sh, bin/claude-2.sh: one-line direct launchers.
- fm-spawn.sh --account <N>: records account=N in meta, sets
  CLAUDE_TRUST_DIR to the task worktree, launches through
  claude-account.sh N. Optional; absent behavior is unchanged.
- docs/configuration.md: Multi-account Claude Code section.
- tests/claude-account.test.sh, tests/fm-spawn-account.test.sh.
…G.md

Adds .agents/skills/herdr-navigation/SKILL.md so agents know how to
navigate herdr panes using existing primitives (herdr pane current,
neighbor, split, send-text, list). Agents commonly don't know these
exist; the skill surfaces them with working examples.

Also adds CHANGELOG.md to .gitignore — it is a generated session
activity log, not source content.
…claude pid (#2)

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid

fm_harness_ancestry_pid() previously returned the first ancestor process
whose command matched a verified harness name. Claude Code's Stop hook
fires as a bg-spare worker several levels below the session's actual
lock-owning claude process (hook shell -> claude bg-spare ->
claude bg-pty-host -> claude -> claude(lock)), so the first match was
the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self()
then never matched state/.lock, and the Claude Stop auto-arm silently
treated its own primary session as an unrelated live owner and never
armed the watcher.

The walk now keeps going past a claude-named match, looking for a still
more ancestral claude-named match, and stops the instant a non-match
follows an already-found match (bounding it to a contiguous run rather
than the literal ancestry top, so an unrelated claude-named process
further up the real process tree is never mistaken for part of this
session's own nested chain). Every other harness keeps the original
first-match-wins behavior, since e.g. Pi's shared signed-wrapper
ancestry actually holds the session at the inner engine pid, not an
outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper
bg-spare chain.

* no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim
* Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment

fm-spawn.sh best-effort enrolls a confirmed launch via
`parlay listen --agent <id>`, backgrounded with its pid recorded to
state/<id>.parlay-listen-pid. fm-teardown.sh best-effort deregisters on
clean teardown: the recorded pid is always killed, and
`parlay agent-down <id>` is called only when parlay is on PATH. Neither
side ever blocks or fails spawn/teardown when parlay is absent or fails.

Adds tests/fm-spawn-parlay.test.sh and two new tests in
tests/fm-teardown.test.sh, plus a new fm_path_without test helper in
tests/lib.sh for simulating parlay's genuine absence from PATH.

* no-mistakes(lint): tests/lib.sh: rename fm_path_without's out array to avoid shellcheck SC2178/SC2128
…on PATH (#3)

hash_pane() only checked 'command -v md5', which fails in any launch
context whose PATH omits /sbin (observed here: a background task shell
with PATH=~/.local/bin:~/.bun/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin,
no /sbin). It then fell through to md5sum, which macOS does not ship at
all, producing a 'md5sum: command not found' error on every pane hash
and likely starving the pane-change detection that watcher cycles rely
on to report an actionable reason.
…and manual (#7)

* fix(bin): allow session-local todo tools in the subagent guard (kunchenguid#1204)

* fix(guard): allow session-local todo tools in the primary

The delegation-shape guard denied TaskCreate and TaskUpdate because their
normalized names contain the `task` stem. Those tools write only the harness's
session-local todo list, which has no executor: it spawns no agent, allocates
no worktree, registers no schedule, and starts nothing that outlives the
session. That is not the unaccounted work the guard exists to stop, so the stem
match was a false positive, and the deny text told the primary to run
bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry.

Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than
widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only
observe or stop existing work. Both lists stay exact-name so neither can widen
by substring.

Tests cover the two allowed names and six near-miss names that a substring or
shortened-stem widening would release; both mutations were watched red.

* no-mistakes(review): drop session-local todo tools from recommended deny list

* no-mistakes: apply CI fixes

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (kunchenguid#1206)

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid

fm_harness_ancestry_pid() previously returned the first ancestor process
whose command matched a verified harness name. Claude Code's Stop hook
fires as a bg-spare worker several levels below the session's actual
lock-owning claude process (hook shell -> claude bg-spare ->
claude bg-pty-host -> claude -> claude(lock)), so the first match was
the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self()
then never matched state/.lock, and the Claude Stop auto-arm silently
treated its own primary session as an unrelated live owner and never
armed the watcher.

The walk now keeps going past a claude-named match, looking for a still
more ancestral claude-named match, and stops the instant a non-match
follows an already-found match (bounding it to a contiguous run rather
than the literal ancestry top, so an unrelated claude-named process
further up the real process tree is never mistaken for part of this
session's own nested chain). Every other harness keeps the original
first-match-wins behavior, since e.g. Pi's shared signed-wrapper
ancestry actually holds the session at the inner engine pid, not an
outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper
bg-spare chain.

* no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim

* no-mistakes: apply CI fixes

* fix: conferma l'avvio del watcher su Windows/MSYS (kunchenguid#1212)

* fix: confirm watcher startup on MSYS

* no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test

* no-mistakes(review): validate OpenCode ready timeout, make uname cache internal

* fix(spawn): forward CLAUDE_CONFIG_DIR to claude crewmates (kunchenguid#1195)

* fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates

Crewmate panes are created by a long-lived tmux/herdr daemon that does not
inherit firstmate's current environment. When firstmate runs under a non-default
CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare
`claude` in the crewmate pane fell back to the default ~/.claude store and
launched unauthenticated, blocking the crewmate before it could do any work.

fm-spawn now prefixes the claude launch with firstmate's own resolved
CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config
store firstmate is authenticated with. An unset value is the single-store
default and adds no prefix; non-claude harnesses are unaffected.

Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set,
omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test
helper so launch assertions no longer depend on the developer's environment.

* no-mistakes: apply CI fixes

* fix: preserve dispatch identity across authentication checks (kunchenguid#1233)

* fix: preserve dispatch harness identity

* no-mistakes(review): Fix Grok counterfactual tuple validation

* no-mistakes(document): Scope dispatch authentication to selected tuple

* fix: restore dispatch instruction budget

* no-mistakes(review): Scope dispatch authentication after candidate selection

* fix(bin): normalize relative durable paths (kunchenguid#1256)

* fix(bin): handle dash-leading harness process names (#2)

* fix: handle dash-leading harness process names

* no-mistakes(review): Make dash-leading harness regression hermetic

* fix: preserve secondmate reply routes across relative homes

Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved.

Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory.

* no-mistakes(review): Preserve absolute overrides and normalize relative durable paths

* no-mistakes(review): Normalize relative home before deriving durable paths

* no-mistakes(document): Document relative durable-path normalization

* no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization

* no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck

* refactor(skills): make Bearings chat-only by default (kunchenguid#1136)

* Add internal status skill

* no-mistakes(document): register /status skill in documentation-audiences inventory

* no-mistakes(lint): replace grep|wc -l with grep -c in status skill test

* test: silence literal status skill patterns

* Refactor bearings default to chat-only

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>

* Clarify follow-up routing during validation (kunchenguid#1277)

* fix: honor concrete approval for project operations (kunchenguid#1272)

* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* docs: add captain-approved project operation exception to hard rule 1

Firstmate stays read-only over projects by default, but when the captain
clearly approves a concrete project operation and scope in the moment,
firstmate may perform exactly that approved operation with its own tools.
The approval is never inferred, broadened, or standing, and it does not
relax the existing force, discard, unlanded-work, or merge-authority
boundaries.

* no-mistakes(review): Clarify captain-approved project operation boundaries

* no-mistakes(document): Clarify captain-approved project operation scope

* docs: cover directories and preserve the operation-or-scope alternative

Widen the captain-approved project operation exception in AGENTS.md to
files or directories, and restore the explicit operation-or-scope
alternative that a prior pipeline auto-fix had collapsed into "and".

Rework project-management SKILL.md's Remove section, which previously
told firstmate to refuse project removal until a guarded helper existed;
that helper was never built, so the text directly contradicted the new
instruction-only exception. It now points at the exception plus the
existing removal preflight it still requires unchanged.

Update the one instruction-owners test assertion that hard-coded the
sentence removed above, so the suite tracks current, not obsolete, text.

* no-mistakes(review): Align project removal preflight with approved exception

* no-mistakes(document): Align project removal documentation with approved exception

* fix: restore removal test byte-for-byte and preserve the default sentence

tests/fm-instruction-owners.test.sh had been changed to assert different
text; restore it byte-for-byte to origin/main. project-management SKILL.md's
Remove section now keeps the exact default "Never issue a raw removal
command from Firstmate." sentence that test still asserts, immediately
followed by the already-approved captain-operation-or-scope exception, so
the default and the exception both stay explicit and consistent.

* no-mistakes(document): Align project-write boundary documentation

* fix(skills): route new project intake through secondmate scopes (kunchenguid#1275)

* Route project intake through secondmate scopes

* no-mistakes(test): Guard all main-home project registry mutations

* no-mistakes(document): Consolidate secondmate routing documentation

* no-mistakes: apply CI fixes

* Restore new-project routing scope

* no-mistakes(document): Clarify secondmate routing for new-project intake

* no-mistakes: apply CI fixes

* fix: scope validation corrections by accepted behavior (kunchenguid#1281)

* fix: scope validation corrections by accepted behavior

* no-mistakes(review): Classify stale delivery evidence as an autonomous correction

* test: replace source assertions with behavioral coverage (kunchenguid#1282)

* test: remove source-content assertions

* no-mistakes(review): Replace source assertions with runtime behavior coverage

* no-mistakes(review): Isolate Kimi task temp runtime coverage

* no-mistakes(document): Refresh test cleanup documentation

* no-mistakes: apply CI fixes

* fix(watch): escalate busy workers with no completed turn (kunchenguid#1286)

* fix(watch): bound how long a busy pane may run with no completed turn

A busy pane (backend busy state or the harness's rendered footer) was
unconditional, unbounded proof of liveness in every escalation path, so a
hung foreground tool call behind a busy signature could run for hours
undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic-
backtracking regex hung one bash call for 25h behind an unchanging
"Working..." footer).

FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may
run with no completed turn (state/<id>.turn-ended, or its spawn record
before any turn has completed). Past the bound, busy_turn_over_age routes
the pane through the existing wedge_timer_check, reusing the identical
stale reason, escalation counter, and demand-deep-inspection marker for
human inspection only - never an automatic interrupt, signal, or restart
of the worker or its tool process. A completed turn resets the age.

Reproduced end-to-end against the real installed Pi TUI: a foreground
`sleep 999999` bash call with no timeout renders the actual busy footer,
and two captures ~15s apart show the elapsed counter changing the pane
hash while the same turn stays unfinished. Running the pre-fix watcher
against the real captures showed it never starts a wedge timer no matter
how long the pane stays busy; the fixed watcher starts and escalates the
timer through the same mechanism, while the real hung process remained
untouched and alive throughout.

* no-mistakes(review): fix: parse enriched AFK stale reasons

* no-mistakes(review): fix: preserve enriched wedges during AFK supervision

* no-mistakes(review): fix: route all enriched AFK wedges

* no-mistakes(document): Clarify busy-turn age supervision documentation

* fix(gitignore): ignore config/ as a directory, not by exact filename (kunchenguid#1261)

A name-by-name list of config/ entries silently stops ignoring any new or
home-local file placed there, which makes the working tree read as dirty and
blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md
already documents config/ as captain-private and gitignored as a category;
this makes .gitignore match that contract.

* fix(tests): replace source-content .gitignore assertion with behavioral coverage (kunchenguid#1304)

The second assertion in fm-gitignore-config.test.sh (added by kunchenguid#1261) greps
.gitignore for a specific spelling of the config/ ignore pattern. It fails
on a semantically equivalent pattern like config/** and does not prove Git
actually ignores anything, per the completed source-content-test audit.

Replace it with a real git check-ignore control test on a generated
unrelated path, and strengthen the existing directory-coverage test with
generated unpredictable direct and nested config/ paths.

* feat: bound and consolidate startup memory during stow (kunchenguid#1303)

* Add bounded startup memory curation

* no-mistakes(review): Record reproducible stow verification evidence

* no-mistakes(review): Validate inherited secondmate stow evidence

* no-mistakes(document): Document editable startup-memory budget propagation

* fix(herdr): place workers in the launching workspace (kunchenguid#1328)

* fix(herdr): place workers in the launching agent's exact workspace

Herdr enforces no workspace-label uniqueness, and spawn resolved its
container by taking the FIRST workspace whose label matched the home
label. With two workspaces both labeled "firstmate", a worker launched
from the second one was created in the first, so it appeared in a
different space than the Firstmate the captain was watching.

Reproduced end to end on Herdr 0.7.5 protocol 17 by running the real
bin/fm-spawn.sh inside a launcher pane in the second "firstmate"
workspace: the worker landed in w1 while its launcher was in w2, with an
unrelated third workspace focused throughout, which also rules out any
dependence on the focused workspace.

Placement now binds to the launching process's own Herdr identity. Herdr
injects HERDR_PANE_ID, HERDR_SESSION, and HERDR_SOCKET_PATH into every
process it manages a pane for, and fm_backend_herdr_launcher_identity
resolves that pane's current owning tab and workspace live from Herdr,
cross-checking the pane against its tab and confirming the workspace
exists exactly once in the session. The injected HERDR_TAB_ID and
HERDR_WORKSPACE_ID are creation-time snapshots and are deliberately not
read as current identity. Labels are no longer placement authority.

A claimed parent identity that is unreadable, contradictory, stale, or
from another named session or Herdr server stops the spawn before any
worker endpoint exists, rather than degrading to a label search. A
launcher with no Herdr ancestry has no workspace to inherit and keeps
the per-home labeled container, which must now resolve to exactly one
workspace; two same-labeled candidates refuse instead of adopting
either. A --secondmate launch keeps standing up that home's own
workspace by design.

With presentation spaces enabled, the projected child is created and
bound under that same exact parent and anchors its ordering on it, so a
duplicated home label no longer makes the layout ambiguous. Projection,
focus restoration, restart binding, and quarantine rules are unchanged,
and children are never collapsed into the parent. tmux, Zellij, cmux,
Orca, and the away-mode daemon terminal were each inspected and are not
affected: none resolves a container by searching mutable labels.

tests/fm-backend-herdr-launcher-workspace-e2e.test.sh drives the real
spawn and teardown against an isolated Herdr lab, with its headline case
running fm-spawn.sh inside a real Herdr pane so the identity comes from
Herdr's own injection. The refusal matrix and the ordering anchor are
covered deterministically in tests/fm-backend-herdr.test.sh.

Eight existing real-Herdr suites inherited the developer terminal's own
Herdr pane into their isolated lab sessions, which the new cross-session
check correctly refuses. tests/herdr-test-safety.sh now owns
herdr_forget_inherited_pane and those suites call it, so what they assert
no longer depends on where they were launched from.

Two unrelated fixes found along the way. tests/fm-secondmate-harness.test.sh
had the same class of environment leak through CLAUDECODE, which outranks
PI_CODING_AGENT in bin/fm-harness.sh and made its pi-signed ancestry case
resolve "claude" whenever the suite ran inside Claude Code. And
fm-spawn.sh's usage() printed a fixed line range that had already been
truncating its own help mid-sentence.

* no-mistakes(review): Enforce exact Herdr launcher and projection identity

* no-mistakes(document): Document exact Herdr launcher workspace placement

* fix(calm): refine Calm working boat animation (kunchenguid#1339)

* feat(calm): replace Pi's working row with an animated ship while Calm is on

While Calm is active and one logical agent run is under way, Calm now hides
Pi's built-in working row and renders a small two-row SSHHIP-derived boat in
its place. When Calm is off, Pi's stock working row is left untouched.

The presentation uses only public Pi extension API: setWorkingVisible(false)
plus a temporary setWidget() component whose render(width) owns the responsive
geometry and whose timer requests a TUI render. Visibility follows agent_start
through agent_settled, so the boat does not flicker between tool calls,
automatic continuations, retries, or compaction inside the same run, and
settle, abort, and failure all reach the same cleanup.

fm-calm.ts stays the sole owner of the presentation choice and the only caller
of setWorkingVisible(); the new lib owns the sprite geometry and widget.

* no-mistakes(review): Guarded Calm-off lifecycle visibility writes; focused tests pass

* no-mistakes(test): Fixed Calm E2E wait to include tmux scrollback

* no-mistakes(document): Document Calm working boat behavior

* no-mistakes: apply CI fixes

* feat(calm): slow the Calm boat, animate blue water, and make the sail directional

The boat now moves one column every 880ms while a bounded fixed-cell water phase
advances every 220ms, so the water ripples several times between boat steps and
the presentation reads as calm. One scheduler drives both clocks and disposing
the widget stops them together; ticks rather than wall-clock timestamps drive
every state change, so tests seek animation time exactly.

Colors are standard ANSI foreground codes instead of theme lookups: blue for
every water cell and yellow for the complete boat, each run closed with a
default-foreground reset so nothing bleeds into padding or later frames. ANSI
bytes never enter geometry, so visible width stays exact.

The mainsail is directional and trails aft of the mast: <| travelling right and
|> travelling left. Direction reverses the moment the boat lands on an endpoint,
so the endpoint frame already shows the new heading and no frame at or after a
bounce shows the previous sail.

* test(calm): wait for the Ctrl+O expansion redraw this block asserts

* docs(calm): record the revised working-presentation verification evidence

* no-mistakes(document): Fix Calm feasibility document EOF whitespace

* fix(dispatch): preflight candidate auth before quota escalation (kunchenguid#1349)

* fix(dispatch): scope candidate authentication to its own surface

A locally expired timestamp in one credential store was reported to the
captain as a sign-out, including for dispatch candidates that never read
that store. A `harness=pi, model=xai/grok-*` candidate authenticates
through Pi's own xAI credential, but the only Grok quota reading
available was gated on the standalone Grok CLI's separate token, whose
expiry clock drifts independently. The always-loaded intake rule then
turned that unreadable quota into a mandatory captain escalation.

Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts
that must not depend on agent memory: it resolves a tuple's
authentication surface from quota-axi's own emitted auth sources rather
than from a harness or model name, so another harness's CLI can never
gate a candidate that does not use it. A vendor CLI is launched only
when the tuple's own harness owns the credential store under test and a
non-destructive discovery command is registered for it, which today is
`grok models` alone. That probe runs at most once with stdin closed and
a hard timeout, reads its verdict from the first stdout line because the
command exits 0 either way, treats unrecognized output as indeterminate,
and never invokes login, logout, or the interactive TUI. Quota is read
at most twice, and unknown headroom never makes a candidate ineligible
on its own.

Update the dispatch procedure to match: usable authentication with
unmeasurable headroom stays eligible at lower preference with the
unknown disclosed, and stop-and-report is reserved for unresolved
authentication, an unresolved relationship, or malformed configuration.
Record that Grok's `credits.remaining` is a prepaid balance rather than
window headroom.

Gate quota-axi at 0.1.16 in bootstrap, the first build reporting
per-credential auth sources. A stale install previously passed the
presence check silently, which is why a fix published two days earlier
was still not in effect.

Replace the orphaned quota-array-dispatch fixtures, which encoded a
`provider: "xai"` shape the tool never emits and had no consumer, with
fixtures shaped like real 0.1.16 output that the new suite drives the
script against. The suite asserts the verdict and, separately, which
vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI
fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture
change selects the right tests instead of refusing.

* refactor(bootstrap): give the quota-axi floor one owner

The floor was stated twice - once in bootstrap's gate and once inline in
the auth preflight - so bumping it needed two edits that could drift.
Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the
existing tasks-axi library, and derive the comparison from the constant
so the number appears exactly once. Bootstrap turns a failing check into
the operator diagnostic; the preflight refuses to emit an unscoped
verdict. Map the new library to both consuming suites so a bump re-runs
them, and record that any usable source means the surface authenticates.

* no-mistakes(review): Captain: bound quota checks and removed Python dependency

* no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry

* no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing

* no-mistakes(review): Captain: reject malformed OpenCode model relationships

* no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation

* no-mistakes(document): Updated dispatch authentication documentation

* no-mistakes: apply CI fixes

* feat(x-mode): reconcile promised public replies deterministically (kunchenguid#1350)

* feat(x-mode): reconcile promised public replies deterministically

A promised final reply in an X or Discord thread was only kept while the
primary remembered it. Compaction or restart erased that memory, so a typed
public-followup obligation could sit at pending-work after its PR merged and
the original thread never got its reply.

Make the promise durable state instead:

- bin/fm-public-followup-emit.sh reports a typed terminal work result (source
  home, work id, generation, outcome, safe deliverables, bounded public-safe
  text) into the owning home's private inbox. The event id is derived from
  that identity tuple, so duplicate reports and restart replay converge with
  no coordination, and nothing ever parses a free-form done: sentence.
- bin/fm-public-followup.sh registers a commitment, reconciles events through
  tasks-axi public-followup, and runs the idempotent delivery sequence
  (begin-delivery with the payload hash, post, record the posted receipt or a
  typed error) against the stored platform and opaque thread binding. A
  delivery interrupted between post and receipt refuses rather than risk a
  second public reply.
- Session start surfaces unresolved commitments from disk, the existing relay
  poll surfaces a new terminal-result set once, and teardown refuses while
  this home still owes a public reply for that exact work.

tasks-axi public-followup remains the only owner of the obligation state
machine, state/x-context/ the only owner of the private request context, and
fm-x-reply.sh the only thing that posts. Its new optional --receipt-file is
the one addition there, so a caller can record how many messages were sent.

A home that never opted into the myfirstmate relay gates out on a single
[ -f "$FM_HOME/.env" ] test: no tasks-axi call, no backlog or context scan,
no output, and no artifact. Evidence in docs/verification/public-followup.md.

* no-mistakes(review): Hardened public-followup reconciliation and ownership guards

* no-mistakes(review): Hardened typed terminal cleanup and receipt reconciliation

* no-mistakes(review): Automated typed-delivery cleanup and strict backlog validation

* no-mistakes(review): Fail-closed parent resolution and registration-safe delivery

* no-mistakes(review): Harden relay gating and validate secondmate bindings

* no-mistakes(review): Use owner-aware single-gate teardown protection

* no-mistakes(document): Correct public-followup documentation drift

* no-mistakes(lint): Quote done literals to fix ShellCheck warnings

* no-mistakes: apply CI fixes

* feat(bin): replace busy heuristics with semantic lifecycle state (kunchenguid#1327)

* feat: add semantic busy-state contract owner and event writer

One owner (bin/fm-busy-lib.sh) for the captain-approved semantic
busy-state redesign: a per-task gen-bound record written only by
bin/fm-busy-event.sh, per-harness trusted-source classification with
explicit source attribution, busy/idle/unknown/dead semantics where
missing, malformed, stale, or untrusted semantic data is unknown -
never idle - and endpoint death is the only process-level override.
The Grok-only rendered-tail fallback and the standalone-Kimi
verification gate live behind the same classifier.

* feat: arm busy-state at spawn and convert Pi to the semantic extension path

fm-spawn arms the busy-state contract for converted adapters and seeds
busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed
per-task extension now reports agent_start -> busy and agent_settled ->
idle confirmed by ctx.isIdle(), covering auto-retries, compaction
retries, tool loops, and queued continuations, while turn_end stays a
wake notification touch. Teardown removes the new record, gen sidecar,
and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy ->
agent-settled idle with the marker still touched.

* feat: convert OpenCode to the semantic session.status plugin path

The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now
classifies from OpenCode's semantic session.status events - busy and
retry are active, idle is inactive - latched to the worker's own
session so a subagent child session can never clear the worker's busy
state. The session.idle marker touch stays a wake notification.
Teardown removes both the new and the legacy plugin filenames.
Live-verified on OpenCode 1.17.18 in a real TUI pane: seed ->
session-busy -> session-status-idle.

* feat: convert Claude to the full lifecycle hooks path

The per-task settings.local.json now wires UserPromptSubmit -> busy
and Stop, StopFailure, and SessionEnd -> idle, so API-error and
shutdown turn ends can never strand a busy record; Stop keeps the
turn-ended notification touch. A refused (stale-gen) event exits 0 and
stays silent so Claude's own lifecycle is never broken. Live-verified
on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch
prompt, Stop closes each turn, a mid-stream Escape interrupt fires no
closing hook, and the firstmate-controlled idle/fm-interrupt clear
resolves it.

* feat: gate Codex busy state behind verified semantic sources

The approved contract prefers Codex's app-server turn lifecycle with
capability negotiation and sanctions its lifecycle hooks as the
intermediate. Live probes on codex-cli 0.145.0 show neither is usable
for a pane worker: the app-server daemon is unreachable for a TUI
thread and refuses to start outside the managed standalone install,
and firstmate-written project hooks never fired (interactive with
directory trust granted, and exec, both with
--dangerously-bypass-hook-trust) while global hooks fired in the same
runs. Codex therefore classifies unknown codex-unverified behind an
explicit probe rather than falling back to idle or footer text, and
fm-spawn installs no unverified Codex wiring.

* feat: gate standalone Kimi busy state on live verification

Standalone Kimi has no installed binary here, so per the approved
contract its semantic path stays guarded and it classifies unknown
kimi-unverified rather than idle - and never from its locale-sensitive
moon-phase spinner, which the redesign forbids inventing as a state
source. The gate records the preferred source order (Wire prompt
request lifetime, which brackets a turn and reports cancellation, then
the documented hooks including Interrupt because Stop does not fire on
interrupts) and the exact evidence required to open it. Arming without
wiring would seed a busy record nothing could clear, so both land
together behind the same gate.

* feat: route busy consumers through the contract and drop the global OR

The watcher, crew-state reader, and away-mode daemon now decide busy
state through bin/fm-busy-lib.sh: only an exact busy verdict counts as
working, and unknown never becomes working or a silent idle, so a crew
whose semantic state is missing, malformed, stale, or unverified
surfaces instead of being absorbed. Crew-state reports the producing
source in its detail. The watcher's global OR regex default is gone;
Grok keeps its isolated fallback inside the contract. The daemon's
supervisor-pane reader stays rendered-text - that pane is not a
recorded task - but is now scoped to firstmate's own detected harness
instead of every vendor signature. Secondmate pending-reply
observation is deliberately unchanged and documented as a
delivery-confirmation signal, not task state.

* docs: point busy-state documentation at the single contract owner

Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as
the owner of what busy means, with per-adapter sources, the
unknown-never-idle rule, the endpoint-death override, and the two
rendered-text readers that deliberately stay outside the contract.
Replaces the stale regex-first prose in architecture, tmux-backend,
herdr-backend, and configuration; converts the harness-adapters
per-harness rows from UI signatures to the semantic source each
harness uses; and records the live verification evidence, including
why Codex and standalone Kimi stay unknown.

* fix: arm away-launch signal handlers before acquiring the lifecycle lock

fm_afk_launch_main acquired its lock and only then installed the EXIT,
INT, and TERM traps. A signal arriving in that window terminated the
process by default action and left the lock directory behind, which
blocks the next away-mode launch until the stale-owner reclaim path
clears it. The release helper only removes a lock this process owns,
so the handlers are now armed first. The accompanying test also killed
the child whether or not the lock had appeared and sampled cleanup the
instant wait returned; it now requires the lock, then allows a bounded
settle, so it proves the guarantee instead of racing it.

* test: align fleet, Kimi, lifecycle, and detection suites with the contract

The fleet snapshot and wake-daemon lifecycle fixtures now prove a
working crew through its own semantic busy-state record instead of
rendered pane text, which is what those consumers read. The Kimi
watcher test asserts the approved contract directly: a standalone Kimi
task classifies unknown rather than matching its moon-phase spinner,
while Grok's isolated fallback still classifies only Grok. The
pi-signed detection cases clear ambient harness markers, fixing a
pre-existing failure where the running session's own CLAUDECODE
outranked the fixture's marker.

* fix: stop teardown from deleting a project's own Codex hooks file

An intermediate revision wired Codex through a firstmate-written
<worktree>/.codex/hooks.json, and teardown removed it alongside the
other generated wiring. The Codex wiring was dropped when its probes
came back unverified, so that removal now targets a file firstmate
never creates - and a project may legitimately track its own
.codex/hooks.json, which teardown would then delete from a pooled
worktree.

* fix: keep busy-record parsing from disturbing its sourcing caller

The record parser split fields with set -- under a temporary noglob,
which clobbers a sourcing caller's positional parameters and restores
glob expansion even when the caller had disabled it. The watcher, the
daemon, and the crew-state reader all source this library, so it now
reads fields with read -a, which never globs and never touches caller
state.

* docs: state exactly which Claude hook paths were reproduced live

The busy-state record listed all four wired Claude hooks in the source
column, which could read as a claim that every one fired during the
pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are
wired from hook names confirmed present in the installed binary, but
the abnormal turn ends they cover were not reproduced.

* test: let reset_fakes own the crew-state busy-text fixture lifecycle

The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so
the variable's lifetime was owned by one test rather than by the
shared reset that every other fake already uses.

* no-mistakes(review): Fix semantic busy-state lifecycle races

* no-mistakes(review): Make busy-state retirement idempotent

* no-mistakes(review): Enforce semantic state boundaries for status and injection

* no-mistakes(review): Restore harness-scoped away-mode busy guard

* no-mistakes(document): Refresh semantic busy-state documentation

* no-mistakes: apply CI fixes

* fix: preserve Calm boat continuity across working periods (kunchenguid#1356)

* fix(calm): resume working boat from frozen column across runs

Keep one extension-owned boat animation for the Pi session so settling
freezes column and direction, the next working period resumes there
without hidden-time jumps, and only a fresh session resets to the left edge.

* no-mistakes(review): Freeze Calm boat from last rendered state

* no-mistakes(document): Document Calm boat continuity contract

* fix: restore evidence-based dispatch eligibility (kunchenguid#1358)

* fix(dispatch): judge candidate provider relations instead of rejecting them

Firstmate deterministically dropped supported Pi candidates in the
openai-codex family. bin/fm-auth-preflight.sh resolved a harness=pi tuple's
credential surface by constructing the source id `pi:<model-prefix>`, so
`pi + openai-codex/gpt-5.6-terra` looked for a `pi:openai-codex` source. That
source does not exist, because Pi's Codex family authenticates through the
Codex store quota-axi already lists as `auth-json`/`cli-rpc`. The tuple
returned `eligible=no reason=surface-unresolved` while the Pi catalog listed
the model and the Codex provider reported fresh, usable credentials with 64
effective percent remaining on its all-model scope.

The prefix construction was only ever valid where Pi holds its own credential
(`pi:xai`, `pi:kimi-coding`), which is why every previously configured Pi tuple
resolved and the defect stayed hidden until a Codex-family Pi model was
configured.

Retire dispatch eligibility from deterministic shell. The dispatching first
mate now establishes model support and provider family from each harness's
authoritative catalog, applies quota at the granularity the vendor supplies,
and shows that reasoning. Provider-level and all-model evidence bounds every
model established in that family; a named-model window bounds only its own
model. Missing model-level quota, a missing auth source, unmeasurable headroom,
and unmodeled authentication are disclosed uncertainty. Only concrete
contradictory evidence blocks a candidate.

Replace the preflight with bin/fm-vendor-auth-probe.sh, which keeps the
captain's approved bounded probe envelope without any routing knowledge: it
takes no harness, model, or provider, reads no quota, renders no verdict, and
holds only a fixed-argv safety allowlist. Its behavior suite proves the absent
identity surface, the untouched quota, the uniform exit status, the fixed argv
with stdin closed, and a real bound even when the configured bound is zero.

Also fixed along the way: a zero FM_*_TIMEOUT silently removed the hard bound,
the pinned Grok version had drifted to 0.2.117, and --changed selection refused
outright on any deleted bin/ script.

AGENTS.md section 4 and quota-array-dispatch own the corrected policy,
harness-adapters gets the catalog-responsibility correction, and
docs/verification/dispatch-auth.md records the 2026-07-30 evidence on
Pi 0.82.0, quota-axi 0.1.16, and grok 0.2.117.

* no-mistakes(review): Reject all-zero vendor probe timeouts

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#2)

* fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid

fm_harness_ancestry_pid() previously returned the first ancestor process
whose command matched a verified harness name. Claude Code's Stop hook
fires as a bg-spare worker several levels below the session's actual
lock-owning claude process (hook shell -> claude bg-spare ->
claude bg-pty-host -> claude -> claude(lock)), so the first match was
the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self()
then never matched state/.lock, and the Claude Stop auto-arm silently
treated its own primary session as an unrelated live owner and never
armed the watcher.

The walk now keeps going past a claude-named match, looking for a still
more ancestral claude-named match, and stops the instant a non-match
follows an already-found match (bounding it to a contiguous run rather
than the literal ancestry top, so an unrelated claude-named process
further up the real process tree is never mistaken for part of this
session's own nested chain). Every other harness keeps the original
first-match-wins behavior, since e.g. Pi's shared signed-wrapper
ancestry actually holds the session at the inner engine pid, not an
outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper
bg-spare chain.

* no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim

* Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment (#4)

* Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment

fm-spawn.sh best-effort enrolls a confirmed launch via
`parlay listen --agent <id>`, backgrounded with its pid recorded to
state/<id>.parlay-listen-pid. fm-teardown.sh best-effort deregisters on
clean teardown: the recorded pid is always killed, and
`parlay agent-down <id>` is called only when parlay is on PATH. Neither
side ever blocks or fails spawn/teardown when parlay is absent or fails.

Adds tests/fm-spawn-parlay.test.sh and two new tests in
tests/fm-teardown.test.sh, plus a new fm_path_without test helper in
tests/lib.sh for simulating parlay's genuine absence from PATH.

* no-mistakes(lint): tests/lib.sh: rename fm_path_without's out array to avoid shellcheck SC2178/SC2128

* feat: add beads as third backlog backend option

When config/backlog-backend=beads is set, firstmate uses the beads federated
'task' store as the queue source instead of data/backlog.md. Session-start's
digest lists items with status:ready label from the beads store.

- Add fm_beads_backend_available() check in fm-tasks-axi-lib.sh
- Add print_backlog_beads_compact() rendering function in fm-session-start.sh
- Update print_backlog_compact() to prioritize beads backend when configured
- Add beads backend validation to bootstrap (checks task CLI and store reachability)
- Update docs/configuration.md to document beads backend option
- Update AGENTS.md section 10 to reference beads backend in backlog contract
- Beads backend reuses existing task linkage machinery (task set-state, task close)

The beads backend is fail-open: if task CLI is missing or store is unreachable,
bootstrap reports a MISSING: diagnostic line and the home can still operate.

* test: add beads backend integration tests

Tests for:
- fm_backlog_backend_value() reading beads config
- fm_beads_backend_available() checking task CLI and store
- fm_tasks_axi_backend_available() returning false when beads is set
- whitespace handling in backend config values

* fix: add install_cmd support for task (beads CLI)

The install_cmd() function now recognizes 'task' and provides an install
command for the beads CLI tool.

* fix: make print_backlog_pointer() backend-aware

Update print_backlog_pointer() to provide backend-specific guidance:
- beads backend: suggest 'task show <id>' for beads task store
- manual backend: suggest 'inspect data/backlog.md'
- default/tasks-axi: original message with tasks-axi and data/backlog.md

* no-mistakes(document): Add beads backlog backend documentation about feature limitations for handoff and decision holds.

* no-mistakes(review): Fix beads backend binary name mismatch and add unsupported operation checks

* no-mistakes(document): Add beads as third supported backlog backend - sync documentation

* no-mistakes(document): Add beads as third supported backlog backend - sync documentation

* no-mistakes(lint): Remove duplicate deregister_parlay_agent function definition

* no-mistakes(document): Document beads as third backlog backend option

* Fix: restore backlog pointer for all backends and quote basename argument

- Add 'or data/backlog.md' fallback to beads backend pointer for consistency
- Quote basename argument in fm-session-lock-lib.sh to handle dash-leading process names
  (fixes 'basename: missing operand' when ps output starts with '-')
- Fixes tests/fm-session-start.test.sh:1148 and tests/fm-secondmate-harness.test.sh

* no-mistakes(document): Documented beads backlog backend throughout project - one clarification edit to configuration.md made

* Fix: manual backend pointer must include 'or data/backlog.md' fallback

* Fix: manual backend pointer wording - include 'or data/backlog.md' with manual guidance

* no-mistakes(review): Add beads backlog fallback to manual when query fails

* no-mistakes(review): Add beads backend check to backlog_refresh_reminder messaging

---------

Co-authored-by: Daniel Kuykendall IV <danielkuykendall23@gmail.com>
Co-authored-by: Unknownzed <45267749+Unknownzed@users.noreply.github.com>
Co-authored-by: lhalbert <lucashalbert@users.noreply.github.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: AG <ag@agw3.org>
Co-authored-by: deeto15 <92119640+deeto15@users.noreply.github.com>
Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
trillium and others added 18 commits August 5, 2026 01:14
…be the fixture

'healthy no-supervision-needed native stop must allow' failed on a clean
origin/main worktree on macOS while passing in CI (robots-qgcg).

The guard resolves its home as FM_HOME > FM_ROOT_OVERRIDE > the guard
script's own repo root (bin/fm-turnend-guard.sh:62-66). Every fixture here
installs the guard into its own <fixture>/bin/, so with those unset each
fixture self-resolves to its own state/. run_hook pins FM_HOME per-fixture,
which is why the direct-guard tests were always immune - but the harness
adapter tests invoke the hook the way the real Grok/codex adapters do, with
no per-call pin, so they inherited the ambient environment.

An agent session running inside a live firstmate home exports FM_HOME. The
grok adapter then handed the shared guard that REAL home, which had an
in-flight task and a fresh beacon, so the assertion's idle fixture reported
"1 task(s) in flight ... last beat: 13s ago" and exited 2. CI exports no
FM_HOME, so nothing leaked there. The sibling adapter cases that expect exit
2 were passing for the wrong reason on such a box - the host state also read
as in-flight - and now pass for the right one.

This is a test-hermeticity defect, not a product one: FM_HOME winning over
repo-root state is the guard's designed precedence, asserted by its own
passing tests ("blocks from active FM_HOME state, not only repo-root state"
and "ignores stale repo-root state when FM_HOME is set"), and no adapter -
Grok, codex, or Claude - pins it. So the fix restores the file's stated
"all hermetic over temp dirs" invariant rather than changing the guard.

Adds test_suite_environment_is_hermetic to pin the invariant, so a future
leak fails loudly and names the variable instead of silently making every
adapter verdict describe the developer's own home.

Verified: 51/51 pass, including under a deliberately hostile
FM_HOME + FM_STATE_OVERRIDE; dropping the unset reproduces the original
failure and trips the new guard first. bin/fm-lint.sh and
bin/fm-test-run.sh --check-coverage both clean.

Fixes robots-qgcg
…tion (#47)

* fix(watch): verify continuity after a quiet arm close instead of assuming it

A harness stop that kills the Claude Stop auto-arm takes the watcher with it.
bin/fm-watch-arm.sh's TERM handler kills its watcher child, records
reason=arm-interrupted successor=none, and exits by design - no handoff. That
left supervision dead with nothing scheduled to notice, three times in one
night.

Nothing detected it because both available signals lie in this case:

  - The liveness beacon is touched every poll, so a watcher killed while
    healthy leaves a beacon 3-13s old - indistinguishable from a live one.
  - The arm's benign "watcher: idle" line asserts "adapter re-arm owns
    continuity". On a Claude primary that adapter is fm-claude-stop-autoarm.sh
    itself, which took the quiet-close branch, wrote outcome=clean, and exited
    0 - handing continuity to itself and then quitting.

The lifecycle ledger classified it correctly the whole time. Nothing read it:
before this change no production script opened state/.watch-cycle-exits.log.

Fix, in two parts:

1. The Stop auto-arm now VERIFIES the deferral it receives. After a quiet
   close it rechecks supervision need, then fm_watcher_healthy (live lock +
   identity match + fresh beacon), and re-arms in its own harness-owned
   foreground tree when no live watcher answers. Bounded by
   FM_AUTOARM_MAX_REARMS (default 20); an exhausted budget or an already-queued
   wake escalates to an exit-2 rewake carrying a continuity-LOST banner. Exit 0
   is now reserved for provably-fine states: no need, AFK, or a verified live
   watcher.

   Deliberately NOT a detached successor: bin/fm-watch-arm.sh's header forbids
   it, and it would convert a loud supervision-down alarm into a silent
   "beacon fresh, nobody listening" - letting the turn-end guard allow a blind
   stop.

2. bin/fm-watch-cycle-lib.sh reads the ledger, and both supervision-down
   banners (fm-guard.sh WATCHER DOWN, fm-turnend-guard.sh TURN WOULD END BLIND)
   print its classification. That covers every harness, not just Claude.

   The reader never cites successor=: only an adapter passing
   FM_WATCH_PREDECESSOR_ARM_PID (the OpenCode plugin, the Pi extension)
   back-fills it, so it reads "none" on a Claude primary even when a healthy
   successor took over. Treating that field as evidence would assert
   supervision was lost on data that cannot support the claim.

Tests: tests/fm-watch-cycle-lib.test.sh is new (parsing =-bearing values,
cleared state after a failed read, the exact arm-interrupted predicate, the
refusal to quote successor=). tests/fm-claude-stop-autoarm.test.sh replaces the
old clean-close test - which encoded the defect - with four covering silent
exit behind a verified live watcher, re-arm when none answers, restored
continuity without a rewake, and the queued-wake escalation.

Closes robots-iuhj.

* no-mistakes(review): Add missing newline after cycle_evidence printf in two guards

* no-mistakes(document): add fm-watch-cycle-lib.sh to scripts.md toolbelt table
… home (#50)

Every bin/ script resolves FM_HOME as "${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}"
(53 scripts share the idiom), so an ambient FM_HOME silently outranks the
per-call FM_ROOT_OVERRIDE a test case sets. FM_HOME-derived paths - notably
CONFIG - then resolve against the captain's live home instead of the fixture.

CI never saw this because FM_HOME is unset there; it only reproduces on a box
that exports it. On such a box tests/fm-watcher-lock.test.sh's guard-xmode case
failed permanently: it writes config/x-mode.env under its FM_ROOT_OVERRIDE and
asserts fm-guard.sh's repair line sources that path, but the guard probed the
real home's config, found no x-mode.env, and passed --x-mode 0.

Unset the inherited value in tests/lib.sh so local runs match CI. Cases that
genuinely need an FM_HOME set it per invocation, which this does not affect.

Fixes the pre-existing failure tracked as robots-2nrl.

Co-authored-by: fmtest <fmtest@example.invalid>
… rows (robots-eboq)

test_operational_followup_turn_e2e waited on the session file — written the
moment the core persists a turn — and then captured the tmux pane immediately.
The Pi TUI repaints on its own schedule, so the capture regularly caught a frame
drawn before the turn reached the screen: zero CAPTAIN_ANSWER rows, not two.
The `-eq 1` check then failed with "rendered a duplicate captain answer", which
misnamed the defect and made the whole case nondeterministic on main.

Poll the pane for MONITOR_HANDLED_<label>_ONE — the last row of the flow, drawn
after the captain answer — before asserting, matching what replay_exact_case
already does. A genuine duplicate is still on screen by then, so the assertion
keeps its teeth. Report the observed row count in the failure message so a
not-yet-rendered pane can never again be reported as a duplicate.

Verified: 5/5 green after the fix (2/2 failures before it), full file green.
* herdr: scope the idle-shell child check to the shell's own terminal

fm_backend_herdr_pane_idle_shell_sample required the candidate shell own
exactly zero children. An interactive rc that forks a permanent off-terminal
helper makes that unsatisfiable forever: zsh-autosuggestions opens a zpty
whose child zsh becomes session leader of a DIFFERENT pty and lives for the
shell's lifetime. On such a box the proof could never converge, so
fm-herdr-session-cleanup.sh silently preserved every restored projection and
stale workspaces accumulated with no way to clear them.

Scope the child requirement to the shell's own controlling terminal instead.
This loses none of the safety the flat count was providing: any job the pane's
shell actually owns - foreground or backgrounded with & - inherits the pane's
controlling terminal and still refuses the proof. A child on another terminal,
or on none, is not pane work and would survive the shell's death anyway.
The "exactly one shell row" requirement is unchanged.

Verified on a live box: the zsh-autosuggestions zpty child sits on a different
tty than its parent shell (ttys126 vs ttys124), while `sleep 30 &` inherits the
shell's own tty. Both shapes are now covered by regression tests.

* no-mistakes(document): update stale "childless" guarantee in runtime-backends verification
…true (#49)

* fix(spawn): give crewmates non-interactive git editors (robots-1xw8)

A crewmate has no terminal a human can type into, but it inherits the
machine's configured git editor. On this machine that is `code --wait`,
which blocks until a human closes the tab. Any git command that opens an
editor - `rebase --continue`/`-i`, `commit` without `-m`, non-ff `merge`,
`revert`, `tag -a`, `cherry-pick --continue` - therefore hangs forever
with no output and no error. The pane looks identical to a thinking pane,
so nothing surfaces it until the staleness window expires, and the agent's
natural response (retry) stacks more orphaned waiters on the same file.
Observed live on task parlay-45-rebase: three wedged attempts, ~15 min lost,
resolved only by killing the editor waiters by hand.

Two layers, both here:

- fm-spawn now exports GIT_EDITOR=true and GIT_SEQUENCE_EDITOR=true into
  the crewmate's pane shell alongside GOTMPDIR, before the harness launch,
  so the whole class is impossible rather than each worker remembering.
  `true` exits 0 without touching the file, so git proceeds with the commit
  message or rebase todo as written.
- fm-brief adds rule 8 to both the scout and ship Rules sections: never let
  git open an editor, and treat a git command that is silent for over a
  minute as a blocked editor rather than a slow operation.

Also lands the fix violation-0gh asks for, since it is the same one-line
brief addition: ship rule 9 tells a worker whose branch is already checked
out elsewhere to use a detached HEAD, never to remove or modify another
worktree.

New rules are appended rather than inserted so the existing "rule 6"
cross-reference in the ship brief stays correct.

Covered by a new pane-export assertion in tests/fm-kimi-harness.test.sh,
which owns fm-spawn's pane environment contract.

* no-mistakes: apply CI fixes
…ersion guard to test suite (#53)

* tests: skip fm-public-followup when tasks-axi predates the verb

The suite drives `tasks-axi public-followup` end to end, but its guard only
checked that tasks-axi exists. tasks-axi gained `public-followup` in 0.2.3, and
the bootstrap compatibility probe in bin/fm-tasks-axi-lib.sh does not cover it
(that gate owns the backlog-mutation verbs: --archive-body and mv [<id>...]).
So a box with 0.2.2 installed passed the guard and then died on the first
`public-followup add` with the opaque "could not create the public commitment",
which reads as a product bug rather than a stale tool.

CI never saw it because CI runs `npm install -g tasks-axi`, which resolves to
the latest published build.

Probe the verb directly and skip with a message that names the real cause and
the fix.

* no-mistakes(document): document tasks-axi 0.2.3 requirement for public-followup
…56)

* tests: make the real-herdr-gated lane environment-independent

Three suites in the real-herdr-gated family failed permanently on a captain
box while the required Herdr CI lane stayed green, which is exactly the
"permanently red trains everyone to ignore the suite" shape. All three were
the suite reading the operator's environment instead of its own fixture.

tests/herdr-test-safety.sh: unset an inherited FM_HOME, the same fix
tests/lib.sh already carries, which these suites never got because they do
not source it. For the Herdr backend this is a workspace-IDENTITY leak, not
just a config-path one: fm_backend_herdr_workspace_label reads
"$FM_HOME/.fm-secondmate-home", so an exported FM_HOME pointing at a
SECONDMATE home labels the suite's primary workspace 2M-<scope> instead of
1M-FIRSTMATE. That broke fm-backend-herdr-smoke's restart-stability check
(before=w1 after=) and fm-backend-herdr-prune-safety-e2e's adoption check
(adopted :w2 instead of the pre-created w1).

tests/fm-herdr-session-cleanup-e2e.test.sh: pin SHELL=/bin/bash before the
lab server starts. That server hands its own login shell to every pane it
RESTORES, and a per-create --env SHELL does not survive the stop/restore
cycle this case exercises. Inheriting the operator's interactive shell means
a zsh startup that forks a persistent helper leaves the restored pane with a
child forever, so the childless bare idle-shell proof can never converge.

tests/fm-afk-inject-herdr-e2e.test.sh: replace two fixed dwells that wait
for a daemon-produced result with bounded polls (wait_for_log,
wait_for_file). Scenario A's post-Enter digest wait and Scenario D's
max-defer alarm wait were both races against the daemon's next poll under
a serially-loaded box - the reported flake, plus the one it produced on a
full-family run here. The negative dwell that proves the daemon does NOT
inject while input is pending stays a fixed wait.

Fixes robots-7jy9.

* no-mistakes(review): replace fixed sleeps in Scenarios B and C with wait_for_log

* tests: restore duplicate detection in afk-inject Scenarios B and C

The previous commit replaced the fixed sleeps in Scenarios B and C with
wait_for_log, which fixed a real flake but silently weakened what those
two cases actually test.

Each of them carries two assertions, not one:

  - positive: the daemon eventually injects a digest
  - negative: marker_count == 1, i.e. it does NOT inject a SECOND one

wait_for_log returns on the FIRST match. Counting markers immediately
after it therefore samples the log at the earliest possible instant, so
a duplicate digest injected on the daemon's next poll would arrive after
the count and never be seen. The old sleep 10 / sleep 8 happened to
cover that window; the bounded poll does not. Both cases stayed green
while losing the ability to fail on exactly the bug they were written
for - "swallowed Enter produces exactly ONE clean digest".

Add a short explicit dwell between the positive wait and the count. A
negative assertion has nothing to poll for, because the correct outcome
is that nothing further happens, so it needs a dwell for the same reason
Scenario A's pending-input wait does. The daemon runs at FM_POLL=1 here,
so three seconds covers several poll cycles while remaining far shorter
than the sleeps the positive waits replaced.

Scenarios A and D need no equivalent: neither performs an "exactly one"
count after its positive wait.

Refs robots-7jy9.

---------

Co-authored-by: fmtest <fmtest@example.invalid>
…rmetic-env

test(fm-turnend-guard): make adapter tests hermetic; add env-leak regression guard
`bin/fm-brief.sh hw-allowedhosts herdr-web` scaffolded fine, but the very
next command in the same workflow, `bin/fm-spawn.sh hw-allowedhosts
herdr-web`, died with

    bin/fm-spawn.sh: line 930: cd: herdr-web: No such file or directory

Half a dispatch accepted a bare project name and the other half did not.
fm-spawn's resolve_project_dir_arg only rewrote an argument that already
began with "projects/", so a bare name fell through untouched to `cd` and
resolved against the process cwd. The failure named an internal line
number and a shell builtin rather than the project argument, so the caller
had no way to see that "projects/herdr-web" was the accepted spelling.

The mapping from a project argument to a clone directory had three
independent copies (fm-brief, fm-spawn, fm-fleet-sync), which is why they
could drift apart silently in the first place. Give it one owner:

  bin/fm-project-dir-lib.sh
    fm_project_dir_candidate  - argument -> candidate path, no existence
                                check (fm-brief needs a path for a
                                best-effort git-remote read on a clone
                                that may not be there)
    fm_resolve_project_dir    - candidate, then a cwd-relative fallback,
                                then a named error listing what it tried
                                and up to three near-name suggestions
    fm_project_dir_fold       - fold case and separators so "herdrweb"
                                still suggests "projects/herdr-web"

Resolution order preserves what already worked: absolute and explicit
relative paths (./x, ../x, a/b) pass through untouched, so an explicit
path still wins over a same-named clone; a bare name prefers
$PROJECTS/<name> and only then falls back to a cwd-relative directory.
fm-fleet-sync resolves softly, returning the argument unchanged when
nothing resolves, so its existing "not a directory" skip still reports it.

Second half of the same defect: both entry points swallowed an unknown
--flag as a positional, so

    bin/fm-brief.sh <id> --project herdr-web --mode direct-PR

scaffolded a brief for a project literally named "--project" and said so
only in a `warn:` line. Both scripts now reject an unknown option with
exit 2, honor `--` for a positional that must start with `--`, answer
`--help` wherever it appears, and name a missing positional instead of
dying on an unbound array element.

tests/fm-project-dir.test.sh covers all of it through the executables: a
bare name reaching fm-spawn's brief check proves resolution succeeded
without launching anything.

Fixes robots-2rwl.
`herdr agent list` reports every agent in the session, including the panes
firstmate itself spawned, and the spur bridge treated all of them as external.
A firstmate-owned agent therefore woke firstmate TWICE per turn-end - once
legitimately through its state/<id>.status append and turn-end hook, once
spuriously through a `check: herdr-spur:<pane>` wake claiming it had "no
firstmate status file". The noise scaled with fleet size and destroyed the one
property the channel exists for: that a spur means an agent firstmate cannot
otherwise see.

The mapping was always available - herdr records `window=<session>:<pane-id>`
in state/<id>.meta - but the bridge never did the reverse lookup, and
window_to_task could not answer it anyway: it always returns something, so
"nobody owns this pane" was not expressible.

- fm-classify-lib.sh: split the strict metadata scan out as window_owner_task,
  which prints the owning task or returns 1 printing nothing. window_to_task now
  calls it and keeps its own always-something fallback, so every existing caller
  is unchanged.
- fm-herdr-spur.sh: resolve each finish edge's pane back to a firstmate task
  before enqueuing, and drop the edge when one owns it. Both the bare pane id
  and the session-qualified "<session>:<pane-id>" shape are tried, since the
  event stream and `herdr agent list` report the bare form while meta records
  the qualified one - that mismatch is why a lookup would have failed even if
  the bridge had attempted it. Checked at the edge, not at subscribe time, so a
  task spawned or torn down mid-stream is classified against current metadata.

tests/fm-herdr-spur.test.sh covers the strict lookup and drives the bridge
--once against a stubbed herdr: an external pane still spurs, an owned pane
(qualified or bare window=) never does, and a mixed pass emits exactly one wake.
Without the bridge change the owned case reproduces the reported line verbatim.
fix(herdr): filter firstmate-owned agents from the spur completion channel
…olution

fix(bin): unify project-name resolution across fm-brief, fm-spawn, and fm-fleet-sync
All three session presentation lock acquire sites (spawn projection,
teardown preflight, task kill) shared a hardcoded 50 x 0.1s = 5 second
bound. When a holder's critical section legitimately exceeds that, the
waiter degrades: spawn silently falls back to the flat layout, teardown
and kill refuse. That is correct production behavior, but it makes
tests/fm-backend-herdr-presentation-e2e.test.sh nondeterministic under
load - its fake adapter routes every Herdr call through the guarded lab
helper and takes two extra instrumented focus snapshots around every
mutation, so a holder's critical section is far longer than production's.
On a loaded box the waiting spawn or teardown times out and the ordering
and serialization assertions fail for a timing reason rather than a
projection defect.

Make the bound tunable at all three sites via
FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS and _INTERVAL, matching the
existing FM_BACKEND_HERDR_*_POLLS convention. Defaults are unchanged
(50 x 0.1), so production degrade-to-flat semantics are identical.

The e2e suite widens the wait to 600 polls (60s, still env-overridable)
and passes a short bound at the two fixtures that deliberately hold the
lock for their whole spawn, so those still time out promptly and keep
asserting the flat fallback.

Verified: with POLLS=1 the suite reproduces the reported symptom
exactly ("concurrent primary workers did not append stably to the
contiguous block"); with the fix it passes end to end on a box at load
average ~46-81.
@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@trillium, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 49 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 111fddd3-9a47-434d-8aa8-30c1bdad97c4

📥 Commits

Reviewing files that changed from the base of the PR and between 6c865ee and db37b81.

📒 Files selected for processing (6)
  • bin/backends/herdr.sh
  • bin/fm-spawn.sh
  • bin/fm-teardown.sh
  • docs/configuration.md
  • tests/fm-backend-herdr-presentation-e2e.test.sh
  • tests/fm-backend-herdr.test.sh
📝 Walkthrough

Walkthrough

Herdr presentation-lock waits now use configurable poll counts and intervals across spawn, teardown, and task-kill paths. Documentation and E2E contention fixtures define the settings and bounded fallback behavior.

Changes

Herdr presentation-lock polling

Layer / File(s) Summary
Polling configuration and documented behavior
docs/configuration.md, docs/herdr-backend.md, tests/fm-backend-herdr-presentation-e2e.test.sh
Documents the two polling variables and their defaults. The E2E suite sets longer normal waits and short contention bounds.
Spawn lock acquisition
bin/fm-spawn.sh, tests/fm-backend-herdr-presentation-e2e.test.sh
Herdr spawn reads configurable polling values. Contention fixtures override the poll limit to validate flat-layout fallback.
Teardown and task-kill lock acquisition
bin/fm-teardown.sh, bin/backends/herdr.sh
Teardown and task-kill use configurable polling values while retaining bounded failure behavior.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

  • trillium/firstmate#38: Both changes modify Herdr presentation-lock polling in the same runtime scripts.

Suggested reviewers: kunchenguid

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: configurable Herdr presentation-lock timeouts for load robustness.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/herdr-presentation-e2e-load-robustness

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@trillium

trillium commented Aug 5, 2026

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@bin/backends/herdr.sh`:
- Around line 2828-2833: Update the acquisition flow around fm_lock_try_acquire
and fm_backend_herdr_presentation_session_lock_path to re-resolve the session
lock path after acquisition and compare it with the acquired lock_path. If
resolution fails or the paths differ, release the originally acquired lock, keep
lock_held unset, and refuse the kill; only call fm_backend_herdr_kill_serialized
after successful verification, preserving read-only behavior when acquisition or
verification fails.

In `@bin/fm-spawn.sh`:
- Around line 451-459: Validate FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS and
FM_BACKEND_HERDR_PRESENTATION_LOCK_INTERVAL before the polling loop in the
lock-acquisition flow using fm_lock_try_acquire. Ensure the poll count is a
valid non-negative integer and the interval is a valid sleep-compatible value;
for invalid inputs, apply the script’s documented safe fallback behavior before
entering the loop so integer comparison and sleep cannot fail or create a tight
retry loop.

In `@docs/herdr-backend.md`:
- Around line 108-110: Update the shared per-session lock policy so spawn
refuses and remains read-only when the lock cannot be acquired and verified,
removing the documented flat-placement fallback. In docs/herdr-backend.md lines
108-110, revise the spawn behavior accordingly; in docs/configuration.md lines
491-492, align the configuration description with the same refusal policy while
preserving the bounded polling settings.

In `@tests/fm-backend-herdr-presentation-e2e.test.sh`:
- Around line 256-271: Update both deliberate lock-contention fixture call sites
to override FM_BACKEND_HERDR_PRESENTATION_LOCK_INTERVAL alongside
FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS, ensuring the short timeout remains
bounded even when the suite interval is externally configured. Add a behavioral
regression assertion that runs with a non-default interval and verifies the
fixture falls back within the intended bound.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 687f413d-9c4d-45b8-a7e0-62f791f845d0

📥 Commits

Reviewing files that changed from the base of the PR and between 2bdb24d and 6c865ee.

📒 Files selected for processing (6)
  • bin/backends/herdr.sh
  • bin/fm-spawn.sh
  • bin/fm-teardown.sh
  • docs/configuration.md
  • docs/herdr-backend.md
  • tests/fm-backend-herdr-presentation-e2e.test.sh

Comment thread bin/backends/herdr.sh
Comment thread bin/fm-spawn.sh Outdated
Comment thread docs/herdr-backend.md
Comment on lines +108 to +110
Spawn, cleanup, and task kill each wait a bounded time for that shared per-session lock before degrading: spawn warns and falls back flat, while cleanup and kill refuse rather than mutating unlocked.
The bound is `FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS` polls of `FM_BACKEND_HERDR_PRESENTATION_LOCK_INTERVAL` seconds, defaulting to the shipped 50 x 0.1 = 5 seconds (`docs/configuration.md`).
Raise it when a holder's critical section is legitimately slower than that budget - a heavily loaded box, or an instrumented run such as `tests/fm-backend-herdr-presentation-e2e.test.sh`, where wrapping every Herdr call makes a 5-second wait too tight and turns ordinary contention into a silent flat fallback.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Resolve the shared per-session lock policy before merge.

Both documents publish the same behavior: spawn proceeds with flat placement after the shared per-session lock cannot be acquired. The repository rule requires read-only behavior and forbids spawn in this condition.

  • docs/herdr-backend.md#L108-L110: make spawn refuse, or document a verified presentation-only exception.
  • docs/configuration.md#L491-L492: align the configuration description with that decision.

As per coding guidelines, “If the session lock cannot be acquired and verified, remain read-only and do not spawn.”

📍 Affects 2 files
  • docs/herdr-backend.md#L108-L110 (this comment)
  • docs/configuration.md#L491-L492
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/herdr-backend.md` around lines 108 - 110, Update the shared per-session
lock policy so spawn refuses and remains read-only when the lock cannot be
acquired and verified, removing the documented flat-placement fallback. In
docs/herdr-backend.md lines 108-110, revise the spawn behavior accordingly; in
docs/configuration.md lines 491-492, align the configuration description with
the same refusal policy while preserving the bounded polling settings.

Source: Coding guidelines

Comment thread tests/fm-backend-herdr-presentation-e2e.test.sh
@trillium

trillium commented Aug 5, 2026

Copy link
Copy Markdown
Owner Author

Thanks — worked all four. Three fixed, one deliberately skipped with reasoning.

1. Re-verify the lock path before the task-kill mutation (bin/backends/herdr.sh) — fixed. Valid, and this PR made it materially worse: the wait is now widenable, so the window for the session lock path to move mid-wait is larger than the old fixed 5s. fm_backend_herdr_kill now re-resolves the path after acquisition, releases and refuses on mismatch or resolution failure, mirroring the existing check in teardown_herdr_preflight_target. Regression test added: test_kill_reverifies_session_lock_path_before_mutating.

2. Validate the polling values before entering the loop (bin/fm-spawn.sh) — fixed. The tight-loop hazard is real and new to this PR: no prior knob fed a variable to sleep inside a poll loop. Added fm_backend_herdr_presentation_lock_budget in the herdr backend, used by all three acquire sites, falling back to the shipped 50 × 0.1 when either knob is non-numeric. Regression test covers alpha, negative, float, spaced, bare-dot, trailing-dot, double-dot and a shell-injection interval.

Note the helper lives in the backend rather than fm-wake-lib.sh: fm_backend_herdr_kill only sources that lib when fm_lock_try_acquire is undefined, so putting it there left it undefined under the existing stub-based tests. My own new test caught that.

3. Override both polling values in contention fixtures (tests/…e2e.test.sh) — fixed. Correct: the suite interval is caller-overridable, so 10 polls alone did not bound those fixtures. Both call sites now pin FM_BACKEND_HERDR_PRESENTATION_LOCK_INTERVAL alongside the poll count.

4. Resolve the shared per-session lock policy before merge (docs/herdr-backend.md) — skipped, escalated for a human decision. This asks to change deliberate shipped behavior, not to fix this PR:

  • The flat fallback predates this branch — bin/fm-spawn.sh:1311 as of ca61bfa.
  • It is deliberately tested: the exact warning is asserted at tests/fm-backend-herdr-presentation-e2e.test.sh:655 and :1071. Making spawn refuse would break both by design.
  • This PR only added a sentence describing that existing behavior.
  • The cited rule concerns mutating fleet state unlocked. Falling back flat mutates nothing — it skips the projection. Cleanup and kill, which do mutate, already refuse.

Presentation spaces are an optional default-off cosmetic projection; making contention a hard spawn failure is a product tradeoff, so I have routed it for an explicit decision rather than changing it under a flakiness fix.

Full suite re-run after these changes: 169 passing, exit 0.

…ath on kill

Address the review findings on the bounded presentation-lock wait.

Malformed knobs no longer reach the loop. A non-numeric
FM_BACKEND_HERDR_PRESENTATION_LOCK_POLLS failed the integer comparison and
degraded immediately; a non-numeric _INTERVAL made sleep fail every iteration
and spun the wait with no delay at all. Both now fall back to the shipped
50 x 0.1 = 5s through one shared resolver,
fm_backend_herdr_presentation_lock_budget, used by all three acquire sites
(spawn projection, teardown preflight, task kill).

The task kill now re-resolves the session lock path after acquiring and
refuses when it moved, releasing what it took. Acquiring proves only that
that exact directory was taken, and the wait is bounded but not instant, so
closing a pane under a stale path could mutate while another holder owns the
live one. This mirrors the existing re-check in fm-teardown.sh's
teardown_herdr_preflight_target.

The two e2e fixtures that deliberately contend for the lock now pin the
interval as well as the poll count, so a large caller-set interval cannot
stretch their prompt 1s timeout into a long wait.

Adds two regression tests: one drives the resolver across unset, valid,
zero, alphabetic, negative, float, spaced, multi-dot, trailing-dot, and
command-injection inputs; the other proves a session lock path that moves
mid-wait refuses the close and releases the lock it acquired.
@trillium

trillium commented Aug 5, 2026

Copy link
Copy Markdown
Owner Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@trillium
trillium force-pushed the main branch 2 times, most recently from f042a55 to d77b149 Compare August 6, 2026 04:33
@trillium
trillium merged commit 55f2d2a into main Aug 7, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant