Skip to content

Merge upstream/main while preserving fork behavior - #9

Merged
eduardstan merged 19 commits into
mainfrom
fm/fm-upstream-align-3
Sep 7, 2026
Merged

eduardstan merged 19 commits into
mainfrom
fm/fm-upstream-align-3

Conversation

@eduardstan

Copy link
Copy Markdown
Owner

Summary

Merged upstream/main at 0b9f518603189469cc3ed858ca39c164f6aa0a50 into fm/fm-upstream-align-3 with a real merge commit.

The resolution keeps the fork-only Prime Agent adapter, per-secondmate durable runtimes, PR-closed-without-merge wakes, decision-key placement fixes, worktree-isolation guards, per-turn quota instrumentation, AI co-author trailer ban, and the remaining fork behavior while adopting upstream's current contracts.

Conflict resolutions

  • .agents/skills/harness-adapters/SKILL.md, AGENTS.md, bin/backends/tmux.sh, bin/fm-bootstrap.sh, bin/fm-harness.sh, and bin/fm-session-lock-lib.sh: unioned the verified prime-agent and omp adapter registration, detection, and liveness paths.
  • bin/fm-control-lib.sh: unioned Prime Agent and omp control, interruption, exit, wiring, and supervision mechanics.
  • bin/fm-session-start.sh and bin/fm-supervision-instructions.sh: retained Prime Agent extension ownership and added omp extension discovery and supervision instructions.
  • bin/fm-spawn.sh: retained Prime Agent launch and extension wiring while integrating omp launch configuration, model/provider validation, and extension wiring.
  • bin/fm-teardown.sh: retires both Prime Agent and omp task extensions.
  • bin/fm-wake-lib.sh: retains Prime Agent extension ownership and adds omp ownership checks.
  • bin/fm-crew-state.sh: retained cancelled-run live-worker handling and adopted upstream orphaned-CI and daemon-down classification.
  • bin/fm-watch.sh: retained merge-wait handling and adopted bounded stale alarms for open captain holds.
  • docs/configuration.md and docs/turnend-guard.md: documented both adapter integrations.
  • tests/fm-calm-pi-extension.test.sh, tests/fm-crew-state.test.sh, and tests/fm-session-start.test.sh: retained fork regressions and added upstream coverage.

Upstream features now present

  • Live harness guards run by default when available.
  • Treehouse locks for remote secondmate homes.
  • Verified omp (Oh My Pi) harness adapter.
  • Herdr maturity labels.
  • Bounded stale alarms for backlog captain holds.
  • Test-run refusal in the primary checkout when a task marker is set.
  • Extension-registered providers in the Pi supervision branch.
  • The other upstream fixes and tests from commits main..upstream/main.

Validation

  • bin/fm-lint.sh passed with ShellCheck 0.11.0 and actionlint 1.7.12.
  • bin/fm-doc-audience-check.sh passed: 99 surfaces and 354 local links.
  • Full suite command on elysium: bin/fm-test-run.sh --all.
  • Elysium run directory: /home/eduard/fm/runs/firstmate/20260907T190914Z-align-tests.
  • Result: total=191 failed=41 skipped_gate=36 duration_ms=2481313.
  • The failed cases are environment-gated on this host rather than merge conflicts: node is absent, ruby is absent, actionlint is absent in the remote PATH, and Herdr/Orca/Zellij live services are unavailable; the log records the exact failures.
  • Live-harness guards skipped 36 cases as recorded by the runner; the remote run also completed the available pure, secondmate, snapshot, and watcher coverage.
  • The runner's bare bin/fm-test-run.sh invocation refused because this revision requires an explicit selection; the full suite was then run with --all.

tiago-peixoto and others added 19 commits September 6, 2026 03:14
…tree poll (kunchenguid#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure
Co-authored-by: Talon Stark <talonstark@gmail.com>
…d as not failed (kunchenguid#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification
…livery (kunchenguid#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean
…3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout
…ndmate, and primary (kunchenguid#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>
…unchenguid#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>
…nguid#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes kunchenguid#3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner
…nguid#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership
…nchenguid#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks
* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior
…enguid#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass
… is set (kunchenguid#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>
)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
…ranch (kunchenguid#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy
@eduardstan
eduardstan merged commit 11a42a3 into main Sep 7, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.