Skip to content

feat(bin): merge upstream and authenticate named Claude pools with setup tokens - #4

Merged
Freudator86 merged 141 commits into
mainfrom
fm/firstmate-fork-upstream-reconcile
Sep 28, 2026
Merged

Freudator86 merged 141 commits into
mainfrom
fm/firstmate-fork-upstream-reconcile

Conversation

@Freudator86

Copy link
Copy Markdown
Owner

Intent

Update Firstmate to current upstream and reconcile or move aside local changes without losing them. The Alles Knut second mate must remain usable, and intentional operational-fork fixes must not be silently discarded.
The whole purpose is having the setup tokens so i dont have to touch manually all the time.
Fresh tokens are in place on you: config/secrets/AK_CLAUDE__CLAUDE_CODE_SETUP_TOKEN and config/secrets/HLR_CLAUDE__CLAUDE_CODE_SETUP_TOKEN, both minted by the captain and replacing rejected ones. Both files are CLAUDE_CODE_SETUP_TOKEN=, 108-character values, 0600. Nothing was placed on ak-secondmate; that placement is yours to choose. Verify by effect before reporting them good.
We need to keep our fixes separate from the vanilla Firstmate so we don't lose them and don't deviate. Check every eight hours on a cadence for upstream updates.
Never at the full hour or 5 minutes. Put variance plus or minus 13 minutes around the 8 hours.
Use wall-clock anchors at 08:00, 16:00, and 00:00, each with that variance, rather than an interval that drifts from the previous run.
Bounded auth.

What Changed

  • Merges 137 upstream commits into the fork while preserving both ancestries: picks up upstream's supervision host, Devin adapter, fleet ledger, host mirror, JEV memory guard and the AGENTS.md-to-skill relocation, drops the fork's superseded remote-home-seed/remote-inbox delta in favour of upstream, and re-applies the fork's named-Claude-pool facts to the relocated skill and doc contracts.
  • Extends bin/fm-claude-auth.sh with optional setup_token_file profiles and a run subcommand: the token is read without sourcing from an owned, non-symlink, mode-0600 file and exposed only as CLAUDE_CODE_OAUTH_TOKEN in the process environment (never argv or generated launch scripts); check now proves the token by a hard-bounded 90s Haiku model response rather than claude auth status, then marks first-run onboarding complete in that private store; named profiles refuse when config/claude-account also pins the home.
  • Threads the named pool through dispatch and worker lifecycle: crew-dispatch.json rules and resolver output accept claude_profile on claude profiles, fm-spawn takes --claude-profile, records it in task meta and wraps the launch in fm-claude-auth.sh run (with launch-local bypass consent for a fresh store), fm-control relaunch keeps the recorded pool and preflights it before stopping the old agent, and claude-profiles.json is kept out of inherited config. New tests/fm-claude-auth.test.sh and tests/fm-claude-token-live-e2e.test.sh plus spawn/dispatch/relaunch test and docs/verification updates cover it.

Risk Assessment

⚠️ Medium: The fork delta is small, fail-closed, documented, live-verified by effect, and both prior rounds' accepted fixes landed cleanly, but it rides on a very large upstream merge and touches credential selection for worker launches, leaving one unreachable fail-open fallback worth tidying as a follow-up.

Testing

I derived the scenarios from the fork-local layer (named Claude setup-token pools) rather than the upstream merge bulk, then drove what I could against the real product. The headline result: both captain-minted setup tokens pass the authentication-effect check and both carry an unattended worker launch in a fresh store all the way to a model response over a real 40x120 pty, with ambient credentials deliberately poisoned so only the pool's own token file could produce a pass. Both review-round fixes are demonstrated by behavior, not by reading source: a vendor stubbed to hang forever with the deleted unbounded-timeout knob set to 0 still returns at exactly 90 seconds with a refusal, and the verified stores contain hasCompletedOnboarding with the theme key absent. The adversarial arms I drove live all refuse safely - a rejected 108-character token yields auth=unverified:token-effect and its store is never onboarded, a named pool refuses to combine with a home-wide config/claude-account pin, 0644 and symlinked token files are rejected, and claude_profile on a codex harness is rejected - while the dispatch duplicate diagnostic now correctly names harness, model, effort, and Claude profile. Four arms could not be driven live and are reported untested rather than passed: the dispatch candidate/profile-line rendering needs a real TYPESAFE_API_KEY (the live attempt reached the POST and got http 401), and the fm-spawn secondmate exclusion, the fm-control relaunch pool routing, and claude-profiles.json non-inheritance are fleet lifecycle operations the product itself refuses from inside a no-mistakes gate worktree; each was only exercised through its suite against the real script, which is not a live product run. The second mate's full operator lifecycle did run green end to end on the merged tree. There is no UI or rendered surface in this change; the end-user surfaces are CLI transcripts and a terminal worker launch, both captured as text artifacts. One HLR check failed transiently on a single attempt and then passed 6 of 6, with a direct vendor probe returning exactly AUTH_OK - a network hiccup against a fail-closed guard, recorded as informational. Nothing in the worktree, the operator checkout, or the real credentials was mutated.

  • Live validation: ✅ go - 9 of 14 scenarios driven live against the product
Scenario Result Live Evidence
Operator checks both captain-minted setup token pools and each is proven good by a real model response, not by a self-reported auth status ✅ pass live bin/fm-claude-auth.sh check --profile ak-pool and --profile hlr-pool run with ANTHROPIC_API_KEY=synthetic-wrong-ambient-key and CLAUDE_CODE_OAUTH_TOKEN=synthetic-wrong-ambient-token, against the r…
A token-backed pool launches an unattended worker in a fresh store and reaches a model response with no interactive onboarding, trust, or bypass screen ✅ pass live FM_CLAUDE_TOKEN_LIVE_E2E=1 bash tests/fm-claude-token-live-e2e.test.sh against real Claude CLI 2.1.276 and both real tokens: each pool drove `fm-claude-auth.sh run -- claude --dangerously-skip-permi…
Adversarial - a rejected setup token is refused and its store is never onboarded, instead of being reported good ✅ pass live A syntactically valid 108-char random token in a 0600 file produced profile=rejected-pool auth=unverified:token-effect exit 1, and its store's .claude.json has no hasCompletedOnboarding key. Evidenc…
Adversarial - bounded auth holds when the vendor hangs and when the deleted unbounded-timeout override is set to zero ✅ pass live Stub claude sleeping 600s on PATH with FM_CLAUDE_TOKEN_CHECK_SECONDS=0, wrapped in timeout 300: returned auth=unverified:token-effect exit 1 at elapsed=90s. Evidence file 03-bounded-auth-hung-ve…
Onboarding readiness writes only hasCompletedOnboarding and leaves the vendor theme preference alone ✅ pass live Both verified stores' .claude.json report {"hasCompletedOnboarding":true,"theme":"<key absent>"} after a successful check. Evidence file 02-pool-check-transcript.txt
Adversarial - a per-task named pool cannot override a declared home-wide Claude account boundary, and the ambient default profile is not preflighted ✅ pass live With config/claude-account present, check --profile ak-pool refused with the combine-refusal message exit 2; check --profile default refused as ambient exit 2; an unconfigured id refused with per-…
Adversarial - setup-token secret hygiene is enforced (world-readable file, symlinked file, and two pools aliasing one store are all refused) ✅ pass live A 0644 token file and a symlinked token file each refused with the owned/non-symlink/0600 message exit 2; two profiles naming one store refused with the aliasing message exit 2. Evidence file 04-pool-…
Adversarial - the dispatch duplicate diagnostic names all four axes the validator applies, and claude_profile is constrained to Claude profiles ✅ pass live bin/fm-dispatch-resolve.sh on four crew-dispatch.json shapes: four-axis duplicate now reports "duplicate harness, model, effort, and Claude profile profiles"; claude_profile on a codex harness and a…
A resolved Claude pool is rendered onto the candidate line and forwarded as --claude-profile on the profile line handed to fm-spawn ⏸️ untested no The prior payload did not establish a live result for this arm: the rendering path requires a clear resolution from the typesafe.ai System One endpoint, and with a placeholder key the live run reached…
The Alles Knut second mate remains usable after the upstream merge, across its full operator lifecycle ✅ pass live bash tests/fm-secondmate-lifecycle-e2e.test.sh stands up isolated homes and really spawns: seed, spawn into the subhome, bare fm-<id> routed send, backlog handoff, recovery respawn, and teardown all…
Adversarial - a secondmate spawn refuses a named Claude pool so a recovery respawn cannot land on a different store, while unnamed spawns are undisturbed ⏸️ untested no The prior payload did not establish a live result: direct invocation of bin/fm-spawn.sh was refused by the product's own boundary ("no-mistakes gate agent must not drive the fleet (NO_MISTAKES_GATE se…
Adversarial - a relaunch onto an unavailable named pool refuses before the running agent is stopped, and a recorded pool survives a relaunch that names none ⏸️ untested no The prior payload did not establish a live result: fm-control relaunch is a fleet lifecycle operation blocked by the same NO_MISTAKES_GATE gate-worktree boundary as fm-spawn, so it was only covered by…
config/claude-profiles.json names local credential stores and is never inherited downstream to a secondmate or a remote home ⏸️ untested no The prior payload did not establish a live result: non-inheritance was only asserted by tests/fm-secondmate-harness.test.sh case B12d, not by a real inherit push, because driving one requires fleet li…
Upstream updates are checked every eight hours on wall-clock anchors of 08:00/16:00/00:00 with plus-or-minus 13 minutes of variance, never on the full hour or 5 minutes ⏸️ untested no No live-validatable surface exists for this in the change: no scheduling artifact is present in the tree (.github/workflows/ carries no schedule trigger, and nothing anchors on those times or applies…
Evidence: Live proof: both real setup tokens verified by effect and each driving an unattended worker launch to a model response

Source: Live proof: both real setup tokens verified by effect and each driving an unattended worker launch to a model response

ok - 2.1.276 (Claude Code): pool 0 real token effect and fresh-store interactive bypass response ok - 2.1.276 (Claude Code): pool 1 real token effect and fresh-store interactive bypass response # live token guard checked 2 pools

ok - 2.1.276 (Claude Code): pool 0 real token effect and fresh-store interactive bypass response
ok - 2.1.276 (Claude Code): pool 1 real token effect and fresh-store interactive bypass response
# live token guard checked 2 pools
Evidence: Operator transcript: AK and HLR pools authenticate, a rejected token is refused, and onboarding writes hasCompletedOnboarding with no theme key

Source: Operator transcript: AK and HLR pools authenticate, a rejected token is refused, and onboarding writes hasCompletedOnboarding with no theme key

$ bin/fm-claude-auth.sh check --profile ak-pool profile=ak-pool auth=authenticated config_dir=.../store-ak exit=0 $ bin/fm-claude-auth.sh check --profile hlr-pool profile=hlr-pool auth=authenticated config_dir=.../store-hlr exit=0 $ bin/fm-claude-auth.sh check --profile rejected-pool profile=rejected-pool auth=unverified:token-effect config_dir=.../store-bad exit=1 Onboarding readiness written into each verified store (hasCompletedOnboarding only; no theme key): store-ak {"hasCompletedOnboarding":true,"theme":"<key absent>"} store-hlr {"hasCompletedOnboarding":true,"theme":"<key absent>"} store-bad {"hasCompletedOnboarding":null,"theme":"<key absent>"}

Firstmate named Claude setup-token pools - operator transcript
Ambient credentials are deliberately wrong (ANTHROPIC_API_KEY / CLAUDE_CODE_OAUTH_TOKEN are synthetic),
so a pass can only come from the named pool's own setup-token file.

$ bin/fm-claude-auth.sh check --profile ak-pool
profile=ak-pool auth=authenticated config_dir=/tmp/fm-auth-drive.8aLCAY/store-ak
exit=0

$ bin/fm-claude-auth.sh check --profile hlr-pool
profile=hlr-pool auth=authenticated config_dir=/tmp/fm-auth-drive.8aLCAY/store-hlr
exit=0

$ bin/fm-claude-auth.sh check --profile rejected-pool
profile=rejected-pool auth=unverified:token-effect config_dir=/tmp/fm-auth-drive.8aLCAY/store-bad
exit=1

Onboarding readiness written into each verified store (hasCompletedOnboarding only; no theme key):
  store-ak   {"hasCompletedOnboarding":true,"theme":"<key absent>"}
  store-hlr  {"hasCompletedOnboarding":true,"theme":"<key absent>"}
  store-bad  {"hasCompletedOnboarding":null,"theme":"<key absent>"}
Evidence: Bounded auth adversarial: hung vendor plus the removed unbounded knob still returns on the hardcoded 90s deadline

Source: Bounded auth adversarial: hung vendor plus the removed unbounded knob still returns on the hardcoded 90s deadline

Adversarial: the vendor CLI hangs forever (stub 'claude' sleeps 600s). The removed knob is also set to 0 (FM_CLAUDE_TOKEN_CHECK_SECONDS=0), which previously meant 'timeout 0' = NO deadline. profile=hung-pool auth=unverified:token-effect config_dir=/tmp/fm-bound.hg03S8/store exit=1 elapsed=90s RESULT: bounded - preflight returned on its own deadline

Adversarial: the vendor CLI hangs forever (stub 'claude' sleeps 600s).
The removed knob is also set to 0 (FM_CLAUDE_TOKEN_CHECK_SECONDS=0), which previously meant 'timeout 0' = NO deadline.
Expected: the preflight still returns on the hardcoded 90s bound instead of blocking a spawn forever.

profile=hung-pool auth=unverified:token-effect config_dir=/tmp/fm-bound.hg03S8/store
exit=1  elapsed=90s

RESULT: bounded - preflight returned on its own deadline
Evidence: Pool boundary and secret-hygiene guards all refuse (home-wide pin, ambient default, unconfigured id, 0644 token, symlinked token, aliased stores)

Source: Pool boundary and secret-hygiene guards all refuse (home-wide pin, ambient default, unconfigured id, 0644 token, symlinked token, aliased stores)

$ bin/fm-claude-auth.sh check --profile ak-pool # with config/claude-account present error: named Claude profiles cannot be combined with config/claude-account; choose per-task pools or the home-wide pin exit=2 $ bin/fm-claude-auth.sh check --profile default error: default Claude profile is ambient and is not preflighted here exit=2 $ chmod 644 <token file>; bin/fm-claude-auth.sh check --profile loose-pool error: setup_token_file must be an owned, non-symlink mode-0600 regular file holding one CLAUDE_CODE_SETUP_TOKEN assignment exit=2

Adversarial guard: a per-task named pool must not override a declared home-wide Claude account boundary.

$ touch config/claude-account   # operator has pinned this home to one account
$ bin/fm-claude-auth.sh check --profile ak-pool
error: named Claude profiles cannot be combined with config/claude-account; choose per-task pools or the home-wide pin
exit=2

Adversarial guard: 'default' is ambient and must not be preflighted as a named pool.
$ bin/fm-claude-auth.sh check --profile default
error: default Claude profile is ambient and is not preflighted here
exit=2

Adversarial guard: an unconfigured pool id is refused with actionable setup guidance.
$ bin/fm-claude-auth.sh check --profile nonexistent-pool
error: Claude profile not configured in this home: nonexistent-pool; install config/claude-profiles.json with a local absolute config_dir for that named pool
exit=2

Adversarial guard: a world-readable token file is refused (secret hygiene).
$ chmod 644 <token file>; bin/fm-claude-auth.sh check --profile loose-pool
error: config/claude-profiles.json is malformed: profiles rejected-pool and loose-pool name the same Claude store /tmp/fm-auth-drive.8aLCAY/store-bad
exit=2

(re-run with its own store so the file-mode guard is the one under test)
$ ls -l <token file>
-rw-r--r-- <token file>
$ bin/fm-claude-auth.sh check --profile loose-pool
error: setup_token_file must be an owned, non-symlink mode-0600 regular file holding one CLAUDE_CODE_SETUP_TOKEN assignment
exit=2

Adversarial guard: a symlinked token file is refused (no following links to another uid's secret).
$ bin/fm-claude-auth.sh check --profile loose-pool   # setup_token_file is now a symlink
error: setup_token_file must be an owned, non-symlink mode-0600 regular file holding one CLAUDE_CODE_SETUP_TOKEN assignment
exit=2
Evidence: Dispatch resolver: four-axis duplicate diagnostic corrected, and claude_profile constrained to Claude profiles

Source: Dispatch resolver: four-axis duplicate diagnostic corrected, and claude_profile constrained to Claude profiles

$ bin/fm-dispatch-resolve.sh brief.md --project demo # two profiles identical in all four axes error: malformed rules file: ... - each rule use must not contain duplicate harness, model, effort, and Claude profile profiles $ bin/fm-dispatch-resolve.sh brief.md --project demo # claude_profile on a codex harness error: malformed rules file: ... ; claude_profile belongs only on a profile whose harness is claude # two profiles differing ONLY in claude_profile passed validation and proceeded to the request

Operator writes two Claude profiles identical in all four axes (harness, model, effort, Claude profile).
The diagnostic must name the rule the validator actually applied - all four axes, not three.

$ bin/fm-dispatch-resolve.sh brief.md --project demo
error: malformed rules file: /tmp/fm-disp.SFzWUE/config/crew-dispatch.json - each rule use must not contain duplicate harness, model, effort, and Claude profile profiles

Two profiles differing ONLY in claude_profile are now legitimately distinct capacity pools (no duplicate error):
$ bin/fm-dispatch-resolve.sh brief.md --project demo
dispatch-resolve: error (http 401 after 268 ms: {"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}})
dispatch-resolve:
  status: error
  reason: http 401 after 268 ms: {"detail":{"error_type":"authentication_error","message":"Cannot authenticate with the server. Please check your API key and try again."}}

Adversarial: claude_profile on a non-Claude harness must be refused.
$ bin/fm-dispatch-resolve.sh brief.md --project demo
error: malformed rules file: /tmp/fm-disp.SFzWUE/config/crew-dispatch.json - each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present; claude_profile belongs only on a profile whose harness is claude

Adversarial: a malformed claude_profile id must be refused.
$ bin/fm-dispatch-resolve.sh brief.md --project demo
error: malformed rules file: /tmp/fm-disp.SFzWUE/config/crew-dispatch.json - each use profile needs harness; model, effort, and floor must be well formed, and provider must match ^[a-z0-9]+(-[a-z0-9]+)*\z when present; claude_profile belongs only on a profile whose harness is claude
Evidence: Targeted fork-local suites (auth preflight, dispatch resolver, relaunch preflight, Claude trust)

Source: Targeted fork-local suites (auth preflight, dispatch resolver, relaunch preflight, Claude trust)

########## tests/fm-claude-auth.test.sh
ok - fm-claude-auth: an already-onboarded store is left untouched
ok - fm-claude-auth: rejected token is not rescued by ambient login; run rereads rotation
ok - fm-claude-auth: token files require private permissions, literal assignment, and no symlink
ok - fm-claude-auth: named pools cannot override a declared home-wide pin
# all Claude auth preflight tests passed
EXIT=0

########## tests/fm-dispatch-resolve.test.sh
ok - schema 6: each candidate binds to its account row; schema 5 is unchanged
ok - quota evidence comes from one quota-axi --json read, and its failure is an error outcome
ok - API, transport, and response failures are error outcomes with exit 0
ok - configuration errors exit 2 before any network call
# all fm-dispatch-resolve tests passed
EXIT=0

########## tests/fm-control-relaunch.test.sh
rm: cannot remove '/tmp/fm-control-relaunch.gSQjth/poolgone-29381/home/state/rl-pool-gone.git-hooks/post-commit': Permission denied
rm: cannot remove '/tmp/fm-control-relaunch.gSQjth/poolgone-29381/home/state/rl-pool-gone.git-hooks/pre-commit': Permission denied
rm: cannot remove '/tmp/fm-control-relaunch.gSQjth/poolgone-29381/home/state/rl-pool-gone.git-hooks/post-rewrite': Permission denied
rm: cannot remove '/tmp/fm-control-relaunch.gSQjth/poolgone-29381/home/state/rl-pool-gone.git-hooks/post-applypatch': Permission denied
rm: cannot remove '/tmp/fm-control-relaunch.gSQjth/poolgone-29381/home/state/rl-pool-gone.git-hooks/pre-merge-commit': Permission denied
EXIT=0

########## tests/fm-claude-trust.test.sh
ok - fm-spawn.sh: a claude secondmate spawn pre-trusts a standalone-clone home
ok - fm-spawn.sh: a claude secondmate spawn pre-trusts a leased worktree home
ok - fm-claude-trust.sh: home-level trust is refused for everything but a home seeded for this secondmate
ok - fm-claude-trust.sh: worktree mode still refuses a secondmate home
ok - fm-spawn.sh: a claude secondmate spawn refuses when home trust cannot be recorded
EXIT=0
Evidence: Second mate remains usable on the merged tree: full seed/spawn/send/handoff/recovery/teardown flow

Source: Second mate remains usable on the merged tree: full seed/spawn/send/handoff/recovery/teardown flow

ok - seed: registry scope+projects, charter copied, clones+origins, no-mistakes init in subhome only ok - spawn: launches in the subhome with persistent charter, records routing meta ok - send: a bare fm-<id> secondmate enqueues a marked request and rings the meta window ok - handoff: in-scope items move verbatim, out-of-scope stays, idempotent ok - recovery: respawns from the durable registry and persistent home ok - teardown: removes the home, then clears meta and the registry route EXIT=0

ok - seed: registry scope+projects, charter copied, clones+origins, no-mistakes init in subhome only
warning: secondmate design sync skipped before launch: primary default-branch commit cannot be resolved
ok - spawn: launches in the subhome with persistent charter, records routing meta
ok - send: a bare fm-<id> secondmate enqueues a marked request and rings the meta window
ok - handoff: in-scope items move verbatim, out-of-scope stays, idempotent
ok - recovery: respawns from the durable registry and persistent home
ok - teardown: removes the home, then clears meta and the registry route
EXIT=0
Evidence: Secondmate exclusion and spawn/relaunch pool routing, with the gate-worktree boundary that blocked direct fm-spawn invocation recorded

Source: Secondmate exclusion and spawn/relaunch pool routing, with the gate-worktree boundary that blocked direct fm-spawn invocation recorded

Secondmate + spawn routing for named Claude pools
=================================================

Direct fleet lifecycle cannot be driven from this gate worktree: the product
itself refuses it, which is the intended operator boundary --

  $ bin/fm-spawn.sh ak-secondmate --secondmate --harness claude --claude-profile ak-pool
  error: no-mistakes gate agent must not drive the fleet (NO_MISTAKES_GATE set)
  exit=3

  $ bin/fm-spawn.sh demo-task-x1 . --mode local-only --yolo off --harness claude --claude-profile ghost-pool
  error: refusing fleet lifecycle from inside a no-mistakes gate worktree
  exit=3

So fm-spawn's public argv interface was driven through its suite, which builds
isolated scratch homes and launches the real script. Observed contracts:

  ok - secondmate spawns refuse named Claude pools instead of promising recovery support
  ok - a Claude capacity pool cannot be attached to a non-Claude spawn
  ok - an unconfigured Claude pool refuses by naming the configuration gap, not a measured login state
  ok - a pools-only Claude profile config keeps serving spawns that name no profile
  ok - batch dispatch forwards shared --harness, --model, --effort, and --claude-profile to every pair
  ok - active crew-dispatch profile does not block secondmate launches
  ok - a persistent claude secondmate keeps its supervisor contract without a task-worker authority overlay

fm-control relaunch preflight (tests/fm-control-relaunch.test.sh):

  ok - the recorded Claude capacity pool survives a relaunch that names none
  ok - an unavailable Claude pool refuses before the agent is stopped and stays recoverable
  ok - an unconfigured Claude pool refuses with the per-home configuration requirement
  ok - a Claude capacity pool does not survive a switch away from claude

The second mate itself remains usable on the merged tree - the full operator
flow ran green end to end (tests/fm-secondmate-lifecycle-e2e.test.sh):

  ok - seed: registry scope+projects, charter copied, clones+origins, no-mistakes init in subhome only
  ok - spawn: launches in the subhome with persistent charter, records routing meta
  ok - send: a bare fm-<id> secondmate enqueues a marked request and rings the meta window
  ok - handoff: in-scope items move verbatim, out-of-scope stays, idempotent
  ok - recovery: respawns from the durable registry and persistent home
  ok - teardown: removes the home, then clears meta and the registry route
- Outcome: ⚠️ 1 info across 1 run (14m5s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 infos
  • ℹ️ bin/fm-spawn.sh:2414 - Fail-open default in new code. After fm-claude-auth.sh check succeeds, fm-spawn parses the store out of the verdict line at bin/fm-spawn.sh:2413; the *) arm at 2414 sets CLAUDE_SELECTED_CONFIG_DIR= and lets the spawn continue. With that variable empty, line 2424 falls back to fm_worker_account_select, line 4354 registers trust in the ambient store, line 5050 drops the bypass-consent setting, and line 5125 skips the fm-claude-auth.sh run wrapper entirely - so the worker launches against the ambient Claude account while preserve_relaunch_meta still writes claude_profile=&lt;id&gt; into the task record, and a later relaunch re-preflights a pool the previous worker never used. Today both success exits in fm-claude-auth.sh (lines 211 and 226) always print config_dir=, so this arm is unreachable; I am flagging it as an informational hardening item, not a live defect. Every other refusal in this layer fails closed, and this one is the exception. Smallest remedy: make the missing-config_dir= case an explicit refusal (print the verdict and exit 1) instead of silently continuing without the selected pool.
  • ℹ️ bin/fm-dispatch-resolve.sh:488 - Calibration on the stated purpose, not a defect in this change, and pre-existing in the fork base a740636 (this change only corrected the diagnostic strings). duplicate_profiles keys on [harness, model, effort, claude_profile] (line 177), so a rule may legitimately list the two token pools as use: [{harness:claude,model:sonnet,claude_profile:&#34;ak&#34;},{harness:claude,model:sonnet,claude_profile:&#34;hlr&#34;}]. But provider_of (line 360) maps by harness only and lane_of (line 361) returns "" for claude, so both candidates bind to the identical quota row, get the identical spendPriority, and hit $ties &gt; 1 at line 488 - status escalate, reason "genuine spendPriority tie", no profile: line. The resolver therefore can never auto-pick between two Claude pools; it hands the choice back every time. Nothing is computed wrong and nothing is unsafe, but if the expectation behind "so i dont have to touch manually all the time" includes the resolver load-balancing across AK and HLR, that part is not there: pools are selectable per task via --claude-profile, not rankable against each other. Making them rankable would need per-account quota rows, which is new machinery well beyond this change; no action is implied here.
⚠️ **Test** - 1 info
  • ℹ️ bin/fm-claude-auth.sh:205 - Observed one transient failure of the live token-effect check for the HLR pool: the first attempt returned auth=unverified:token-effect, then 6 subsequent attempts passed (including 3 first-invocations in freshly created stores), and a direct vendor probe returned exactly {&#34;is_error&#34;:false,&#34;subtype&#34;:&#34;success&#34;,&#34;result&#34;:&#34;AUTH_OK&#34;}. This is a vendor/network transient, not a product defect. The helper fails closed, which is the correct safety direction, but because the vendor's output is discarded (2&gt;/dev/null) the operator sees a healthy pool reported as a dead token with no cause, and must retry by hand - a small tax on the change's "so i dont have to touch manually" goal. No action implied; noted so the pass rate reported here is calibrated.
  • Live validation: ✅ go - 9 of 14 scenarios driven live against the product
Scenario Result Live Evidence
Operator checks both captain-minted setup token pools and each is proven good by a real model response, not by a self-reported auth status ✅ pass live bin/fm-claude-auth.sh check --profile ak-pool and --profile hlr-pool run with ANTHROPIC_API_KEY=synthetic-wrong-ambient-key and CLAUDE_CODE_OAUTH_TOKEN=synthetic-wrong-ambient-token, against the r…
A token-backed pool launches an unattended worker in a fresh store and reaches a model response with no interactive onboarding, trust, or bypass screen ✅ pass live FM_CLAUDE_TOKEN_LIVE_E2E=1 bash tests/fm-claude-token-live-e2e.test.sh against real Claude CLI 2.1.276 and both real tokens: each pool drove `fm-claude-auth.sh run -- claude --dangerously-skip-permi…
Adversarial - a rejected setup token is refused and its store is never onboarded, instead of being reported good ✅ pass live A syntactically valid 108-char random token in a 0600 file produced profile=rejected-pool auth=unverified:token-effect exit 1, and its store's .claude.json has no hasCompletedOnboarding key. Evidenc…
Adversarial - bounded auth holds when the vendor hangs and when the deleted unbounded-timeout override is set to zero ✅ pass live Stub claude sleeping 600s on PATH with FM_CLAUDE_TOKEN_CHECK_SECONDS=0, wrapped in timeout 300: returned auth=unverified:token-effect exit 1 at elapsed=90s. Evidence file 03-bounded-auth-hung-ve…
Onboarding readiness writes only hasCompletedOnboarding and leaves the vendor theme preference alone ✅ pass live Both verified stores' .claude.json report {"hasCompletedOnboarding":true,"theme":"<key absent>"} after a successful check. Evidence file 02-pool-check-transcript.txt
Adversarial - a per-task named pool cannot override a declared home-wide Claude account boundary, and the ambient default profile is not preflighted ✅ pass live With config/claude-account present, check --profile ak-pool refused with the combine-refusal message exit 2; check --profile default refused as ambient exit 2; an unconfigured id refused with per-…
Adversarial - setup-token secret hygiene is enforced (world-readable file, symlinked file, and two pools aliasing one store are all refused) ✅ pass live A 0644 token file and a symlinked token file each refused with the owned/non-symlink/0600 message exit 2; two profiles naming one store refused with the aliasing message exit 2. Evidence file 04-pool-…
Adversarial - the dispatch duplicate diagnostic names all four axes the validator applies, and claude_profile is constrained to Claude profiles ✅ pass live bin/fm-dispatch-resolve.sh on four crew-dispatch.json shapes: four-axis duplicate now reports "duplicate harness, model, effort, and Claude profile profiles"; claude_profile on a codex harness and a…
A resolved Claude pool is rendered onto the candidate line and forwarded as --claude-profile on the profile line handed to fm-spawn ⏸️ untested no The prior payload did not establish a live result for this arm: the rendering path requires a clear resolution from the typesafe.ai System One endpoint, and with a placeholder key the live run reached…
The Alles Knut second mate remains usable after the upstream merge, across its full operator lifecycle ✅ pass live bash tests/fm-secondmate-lifecycle-e2e.test.sh stands up isolated homes and really spawns: seed, spawn into the subhome, bare fm-<id> routed send, backlog handoff, recovery respawn, and teardown all…
Adversarial - a secondmate spawn refuses a named Claude pool so a recovery respawn cannot land on a different store, while unnamed spawns are undisturbed ⏸️ untested no The prior payload did not establish a live result: direct invocation of bin/fm-spawn.sh was refused by the product's own boundary ("no-mistakes gate agent must not drive the fleet (NO_MISTAKES_GATE se…
Adversarial - a relaunch onto an unavailable named pool refuses before the running agent is stopped, and a recorded pool survives a relaunch that names none ⏸️ untested no The prior payload did not establish a live result: fm-control relaunch is a fleet lifecycle operation blocked by the same NO_MISTAKES_GATE gate-worktree boundary as fm-spawn, so it was only covered by…
config/claude-profiles.json names local credential stores and is never inherited downstream to a secondmate or a remote home ⏸️ untested no The prior payload did not establish a live result: non-inheritance was only asserted by tests/fm-secondmate-harness.test.sh case B12d, not by a real inherit push, because driving one requires fleet li…
Upstream updates are checked every eight hours on wall-clock anchors of 08:00/16:00/00:00 with plus-or-minus 13 minutes of variance, never on the full hour or 5 minutes ⏸️ untested no No live-validatable surface exists for this in the change: no scheduling artifact is present in the tree (.github/workflows/ carries no schedule trigger, and nothing anchors on those times or applies…
  • FM_CLAUDE_TOKEN_LIVE_E2E=1 FM_CLAUDE_TOKEN_LIVE_CONFIG=&lt;both real token files&gt; bash tests/fm-claude-token-live-e2e.test.sh - real Claude CLI 2.1.276, real AK and HLR setup tokens, fresh scratch stores/HOME/cwd, pty sized 40x120 via TIOCSWINSZ
  • Hand-driven bin/fm-claude-auth.sh check --profile ak-pool|hlr-pool|rejected-pool with ANTHROPIC_API_KEY and CLAUDE_CODE_OAUTH_TOKEN set to synthetic wrong values
  • Hand-driven adversarial bound: stub claude sleeping 600s on PATH plus FM_CLAUDE_TOKEN_CHECK_SECONDS=0, wrapped in timeout 300, measuring elapsed time to return
  • Hand-driven guards: check with config/claude-account present, with --profile default, with an unconfigured pool id, with a 0644 token file, with a symlinked token file, and with two profiles naming the same store
  • Hand-driven bin/fm-dispatch-resolve.sh &lt;brief&gt; --project demo across four crew-dispatch.json shapes: four-axis duplicate, two pools differing only by claude_profile, claude_profile on a codex harness, and a malformed claude_profile id
  • Repeat-reliability probe: 3 independent fresh scratch roots each running one check --profile hlr-pool, plus 3 repeats in a warm store
  • bash tests/fm-claude-auth.test.sh
  • bash tests/fm-claude-token-live-e2e.test.sh
  • bash tests/fm-dispatch-resolve.test.sh (suite only - the live candidate/profile rendering arm was not driven end to end)
  • bash tests/fm-control-relaunch.test.sh (suite only - fleet relaunch blocked by the gate-worktree boundary)
  • bash tests/fm-claude-trust.test.sh
  • bash tests/fm-spawn-dispatch-profile.test.sh (suite only - fm-spawn refuses direct invocation from a gate worktree)
  • bash tests/fm-secondmate-lifecycle-e2e.test.sh
  • Post-run state check: git status --porcelain clean, ls -l ~/firstmate/config/secrets/ unchanged (mode 0600, mtime 02:08), no fm-lab-* tmux sessions, no claude-profiles.json installed into the operator home
⚠️ **Document** - 1 info
  • ℹ️ docs/scripts.md:8 - Out-of-scope follow-up, not a defect this change caused. docs/scripts.md is the bin/ toolbelt inventory, and the fork's bin/fm-claude-auth.sh has no row there even though docs/configuration.md directs operators to run it by hand ("restore that pool's login with bin/fm-claude-auth.sh"). Adding just that one row would be inconsistent rather than complete: the same table is missing 48 other scripts carried in by the upstream merge, including bin/fm-claude-trust.sh, bin/fm-worker-account-lib.sh, bin/fm-supervision-host.sh, and bin/fm-doc-audience-check.sh, so the table has drifted upstream-wide rather than by this branch. No structural check enforces coverage (bin/fm-doc-audience-check.sh validates audience inventory and links, not this table), so nothing fails today. Suggested follow-up is one deliberate pass reconciling the whole table against bin/, or an explicit statement in its header that it is a curated selection rather than a complete inventory; either is a broader edit than this change's scope allows.
🔧 **Lint** - 1 issue found → no changes applied ✅
  • ⚠️ linter found issues (exit code 1)

🔧 No changes applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

kesslerio and others added 30 commits September 20, 2026 19:47
…unchenguid#5049)

* fix(bin): render the remote charter's steering-inbox path host-local

A freshly provisioned remote secondmate read a parent-home absolute
steering-inbox path in its charter - a location that exists on no route -
and spent its first turn discovering the gap and filing a blocked
decision for what was a render defect. The seed's remote-copy rewrite now
maps the inbox to the route's host-local parent-route inbox, exactly as
it already maps the reply-log path, so every mention - bare path, listing,
and handled/ acknowledgement - lands host-local.

Both rewrites also become plain assignments, because a quoted substitution
nested inside a double-quoted printf argument leaks literal quotes into
the replacement text on stock macOS bash. The lifecycle suite pins the
corrected render both directions against the real seed, provisioning,
and delivery route, sharing one fixture value between the render truth
and the delivery truth.

Closes kunchenguid#5012

* no-mistakes(document): document remote charter's host-local steering inbox
)

* feat(procevent): route worker-owned Lavish rounds

* no-mistakes(review): drop duplicate artifact field from task-owned registration

* no-mistakes(review): post worker reply once, fix ring label, keep re-arm atomic

* no-mistakes(review): keep worker board owned until terminal round acknowledged

* no-mistakes(review): refuse every retirement of an open worker-owned round

* no-mistakes(review): use real lavish reply flag, isolate reply generations

* no-mistakes(review): drop .posted marker for best-effort reply posting

* no-mistakes(review): consume staged reply after listener setup, refuse orphaned captures

* no-mistakes(review): require a reachable owner, redeliver open rounds, roll back failed re-arms

* no-mistakes(review): re-arm only to acknowledge an open round

* no-mistakes(review): conclude only a still-open terminal round

* no-mistakes(review): record the acknowledgement before retiring the board

* no-mistakes(review): retain the registration across a conclude, qualify terminal docs

* no-mistakes(document): Document worker-owned Lavish round lifecycle
…unchenguid#5107)

* fix(bin): reserve contribution observation budget

* no-mistakes(review): Strengthen slow-read regression test to exceed the poll budget
…ness JSON (kunchenguid#5103)

* feat(bin): add idempotent inbox orders, receipts, replies, and readiness

Let a caller supply a request id when publishing a captain inbox note so a
retry returns the original note instead of creating a second one, including
across the crash window between save and wake announcement. Separate saved
from announced so a failed wake is repairable without enqueueing again.
Add bounded receipts JSON with omission disclosure, a durable primary reply
against a note id, and a read-only readiness projection that can say
unknown instead of inferring liveness from a lock file.

* no-mistakes(review): fix(bin): honest inbox announce, reply cursor, and readiness verdict

* fix(bin): resolve ready from lock-holder ancestry; drop lock status --json

Remove the extra JSON surface from fm-lock.sh so its human status still
always exits zero. Have the readiness projection classify the inspected
home from the lock-holder pid via fm-harness.sh ancestry, with an explicit
FM_SUPERVISION_MODEL still winning and an unknown model when there is no
holder. Prove the yes path when that ancestry names a known harness.

* no-mistakes(review): Harden inbox announce, receipts reads, and reply sequence cursor

* no-mistakes(document): Note read-only lock inspection in scripts inventory

* no-mistakes(lint): Pass missing id argument to malformed-reply test printf

---------

Co-authored-by: cliflacata-svg <304148223+cliflacata-svg@users.noreply.github.com>
…ending text (kunchenguid#5118)

* fix(composer): stop a harness footer row from reading as a composer holding text

A harness draws its own furniture below the composer - a user statusLine, a
permission-mode hint - and the cursorless "bottom-most shape wins" rule looks
exactly there. `→` (U+2192) is Cursor's prompt glyph but ordinary text
everywhere else, so a statusLine opening with `→` was selected as a bare
composer, swallowed the hint row beneath it as wrapped input, and answered
`pending` on a visibly empty pane. `fm_task_inbox_ring` defers on exactly that
verdict, and `bin/fm-watch.sh`'s re-ring calls the same function, so the first
doorbell and every retry were skipped and the worker never saw the steer.

Measured live on 2026-09-20: three of five Claude Code 2.1.236 worker panes on
Herdr 0.8.0 had genuinely empty composers and every one of them was refused.

A separator pair that closed over a bare agent-glyph row is a proven composer
container, so the contiguous non-blank rows below its closing rule are that
composer's footer and are no longer composer candidates. The demotion is bounded
by all three of its own preconditions: a blank row ends the zone, a pair that
closed over no glyph row demotes nothing, and a shape with no separator pair at
all (Cursor's half-block rules) is untouched. Real unsubmitted text in that same
composer, including a stray SGR mouse report left by a click in the pane, still
reads `pending`.

Pinned by two portable regressions and by a new cursorless arm on the live
composer-matrix guard, which re-reads each harness's already-proven-idle pane
the way every non-tmux backend reads it and fails naming the harness and
version when that read is `pending`.

* no-mistakes(review): make composer footer-zone demotion shape-independent

* no-mistakes(review): make footer-zone demotion refuse-only and drop rescan

* no-mistakes(lint): quote probe-absent sentinel to clear ShellCheck SC2100

---------

Co-authored-by: Koen Muller <koen@catapult.nl>
…5115)

Co-authored-by: guanchengh-lgtm <271917158+guanchengh-lgtm@users.noreply.github.com>
… an unreadable runs table (kunchenguid#5114)

* fix(bin): stop misreading a no-run branch as an unreadable runs table

Defect: when `no-mistakes axi status`'s overview is truncated (a task's
own branch has zero rows among the shown ones), fm_nm_select_run's
Python fallback derived the repo identity for its direct SQLite query
from a `repo: <path>` line it expected in the overview text. The real
CLI never emits that line, truncated or not (see the genuine capture at
tests/captures/no-mistakes-v1.70.1/overview.toon, which has only
`count:`/`runs[...]:`), so the lookup always failed and reported
"unreadable runs table" for a task that simply has no run on its
branch. On a fleet with many concurrent runs, every idle-branch task
hits the truncated-overview path routinely, so this fired every few
minutes and drowned genuine unreadable/blocked verdicts in noise.

Fix: derive the repo identity from the task worktree path instead,
which is exactly the value `no-mistakes` records as a repo's
`working_path` (confirmed against the existing capped-overview test
fixtures, which already register repos by worktree path). A worktree
path that is not absolute cannot be matched and still reads as
unreadable rather than being guessed at. Also raise the reader's
SQLite busy timeout from 1s to 30s so ordinary lock contention on a
busy fleet cannot masquerade as an unreadable database.

Safety: every other verdict byte-for-byte unchanged - the repo lookup
still requires exactly one matching row (a genuinely corrupt or
mismatched repos table still reports unreadable, per the existing
`repo` failure-mode test), the branch query and row validation are
untouched, and a zero-row result for the branch still flows through
the same recursive re-parse that already turns an empty `runs[0]{...}`
table into `absent`. Added a regression test
(test_capped_overview_without_repo_line_and_no_runs_reports_absent)
that reproduces the real overview shape - capped, zero rows for the
task's branch, no `repo: ` line - and asserts the crew state falls
through to the pane/busy verdict instead of reporting unknown or
"unreadable". Full fm-crew-state.test.sh suite passes unchanged
otherwise.

* fix: recovered same-branch inventory awk misreads empty result as unreadable

fm_nm_select_run's deep SQLite reader rebuilds a `count:`/`runs[...]:`
overview and re-runs it through the same awk selection pass. When that
rebuilt inventory has zero rows for the branch, the row-matching loop never
executes, so its counters (`seen`) stay at awk's uninitialized empty string
while `expected` and `shown` are plain strings parsed from the header text.
Comparing an uninitialized value against a non-numeric string uses string
comparison, so "" != "0" is true, and the END block takes the "unreadable
runs table" branch instead of falling through to the correct "absent"
verdict for a branch with genuinely zero runs.

Coerce the affected END comparisons with `+0` so they are always numeric,
matching seen/expected/shown/total regardless of whether awk classified
them as strings or numeric strings. A truncated or genuinely malformed
inventory still differs numerically and still reports unreadable.

* no-mistakes(review): bound capped-overview inventory reader and canonicalize worktree lookup

* no-mistakes(review): match recorded repo path first, tolerate duplicate spellings

* no-mistakes(review): revert repo lookup to exact working_path match

* no-mistakes(document): note state-db inventory read under crew-state nm timeout
…ort (kunchenguid#5141)

* fix(bin): require a non-draft pull request before a PR-based done report

A PR-based ship could report done, and merge monitoring could be armed, while the pull request was still a draft. A draft cannot be merged, so the poll waited for an event that could not occur and nobody was asked to merge.

The PR-based definitions of done now require reading the pull request back from the forge and confirming it is not a draft, and a lane that deliberately holds a draft declares a wait instead of done.
bin/fm-pr-check.sh refuses to arm merge monitoring on a draft, naming the draft state, and treats an unreadable draft state as before.
The draft reading now lives in bin/fm-pr-lib.sh and bin/fm-pr-merge.sh uses it, with its refusal to merge a draft unchanged.

Closes kunchenguid#4757

* fix(review): Skip arm-time draft refusal when fm-pr-merge records metadata
* fix(bin): accept quota-axi schema 6 snapshots keyed by provider + accountKey

quota-axi 0.1.47 emits schemaVersion 6 once a provider expands to more
than one account: every provider row carries an accountKey and one
provider id may appear on several rows. fm_quota_json_valid accepted
only schema 5 with unique provider ids, so fm-dispatch-resolve.sh,
fm-quota-choose.sh, and fm-procevent-quota.sh all rejected the live
snapshot and quota-informed dispatch was dead against the current tool.

- bin/fm-quota-axi-lib.sh: the validator accepts schema 6 with
  accountKey required on every row and uniqueness on
  provider + accountKey; schema 5 keeps its exact rules. FM_QUOTA_ROW_JQ
  is the one join every consumer uses: schema 5 binds by provider alone,
  schema 6 binds to the row keyed by the candidate's Pi lane, else the
  provider's default row, else no row (unmeasured, never blocked, never
  by position or summed across accounts).
- bin/fm-quota-choose.sh: accepts schema 6 JSON and the TOON accountKey
  column, and joins through the shared function.
- bin/fm-dispatch-resolve.sh and bin/fm-procevent-quota.sh: join through
  the shared function; an expanded provider with no row for the
  candidate's account is reported as such.
- tests: schema 6 fixtures shaped like the real snapshot, each paired
  with a schema 5 case on the same path; every new case fails on the
  previous scripts and passes now.
- docs: the two sentences naming the row join describe the schema 6 key.

* no-mistakes(review): Fix native Codex quota and expanded provider watches

* no-mistakes(review): Align native Codex account matching across dispatch paths

* no-mistakes(document): Align quota documentation with account-aware snapshots

* no-mistakes(document): Align quota dispatch documentation with account matching

* fix(bin): keep CI lint and the quota watch test portable

- bin/fm-quota-axi-lib.sh: FM_QUOTA_ROW_JQ is read only by the scripts
  that source this library, so full-mode ShellCheck reported SC2034 on
  the assignment; mark it alongside the existing SC2016 disable.
- tests/fm-procevent-quota.test.sh: the schema 6 provider-watch
  assertions used rg, which CI runners do not install, so the case
  failed with 'rg: command not found' rather than on behavior; use grep
  like the rest of the file.

* no-mistakes(document): Documented schema-version account-row compatibility
* test: repair Claude live auto-arm regression

* no-mistakes(review): Assert SessionStart digest completeness within its hook_response event

* no-mistakes(document): Consolidate Claude live verification references
Roll the shared require-no-mistakes action to the tagged v1.80.1 SHA and grant pull-requests: read so the check can read PR bodies.
…nchenguid#5174)

* fix: preserve Pi watcher ownership across session replacement

* no-mistakes(document): Scope Pi predecessor retention away from omp

* no-mistakes(ci): Diagnosed all three failing checks; only one was code-caused. (ci-3, genuine) Stock macOS Bash snapshot compatibility: `tests/fm-pi-watch-extension.test.sh` failed the macOS Bash 3.2 `bash -n` parse sweep with `line 4265: unexpected EOF while looking for matching '`. I built GNU Bash 3.2.0 from source locally and reproduced it. Root cause: the PR added a comment containing an apostrophe (`// Replacement shutdown deliberately retains module 2's established arm until`) inside a quoted here-document (`<<'EOF'`) nested inside a `$(...)` command substitution. Bash 3.2 has a parser bug (fixed in later bash) where an unmatched single quote inside such a here-doc body is treated as opening a shell quote and never closed, aborting the whole file parse. The base commit parses cleanly under Bash 3.2, confirming this PR introduced the break. Minimal fix: reworded the comment to remove the apostrophe (`... retains the established module-2 arm until`), preserving meaning. Verified `bin/fm-lint.sh --list-files` (the 6 changed shell files) now all pass `/tmp/bash-3.2/bash -n`; Bash 5 also parses. (ci-1, infrastructure) Behavior portable serial 8: GitHub API shows the `Run portable serial shard 8` step conclusion=success; only `Upload portable serial shard 8 timing artifact` failed with `Failed to FinalizeArtifact ... (403) Forbidden`. This is a transient artifact-service/cancellation failure, not a test or code failure. No change. (ci-2, infrastructure) Lint 1: fetched the job log via the GitHub API; it ends with `##[error]The runner has received a shutdown signal...` then exit 143. The step was cancelled mid-run, not a ShellCheck finding. Independently ran `bin/fm-lint.sh --partition 1of2 --telemetry ...` locally with pinned ShellCheck 0.11.0 and actionlint 1.7.12: exited rc=0 (no findings). No change. The only code change is the apostrophe removal in tests/fm-pi-watch-extension.test.sh; no other files modified
…d#5236)

* fix(bin): retire windowless leftovers and stop claiming a Pi daemon teardown

Catch-up correctly refuses while a leftover task record has no status file.
Cleanup used to deadlock on those same records when they also had no spawn_gen and no window, so they lingered and wedged every later away-mode return. Teardown now treats a windowless leftover as a missing-endpoint legacy record, and stop reports that no daemon terminal was running when none was launched.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Narrow windowless teardown exception to tmux legacy leftovers

* no-mistakes(review): Validate windowless leftover identity via shared endpoint validator

* no-mistakes(review): Refuse windowless leftovers carrying other backends' endpoint identity

* no-mistakes(document): Clarify windowless teardown retry documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…rted (kunchenguid#5250)

* fix: surface parked launch prompts as not started

* no-mistakes(document): docs: record launch-prompt busy backstop classification

* no-mistakes(document): docs: align tail40 and rendered-text comments with launch-prompt backstop
* feat(afk): make /afk itself the go with a same-turn record write

Collapse the propose-then-confirm away entry into one 'enter' step that
writes state/.afk-contract immediately and prints the announcement and
read-back after the record exists, never asking for a go. The retired
propose, confirm, and --proposal inputs are refused by name, and a stale
proposal left by an older version is removed rather than promoted.
Refresh and replace semantics, verbatim words, the single writer, the
never-set, and per-harness launch behavior are unchanged.

* no-mistakes(document): Refresh away-entry documentation evidence
…nguid#5294)

* fix(bin): map passed-with-override to done instead of unknown

no-mistakes' axi status emits outcome: passed-with-override for a run
that finished with an explicitly approved Test or CI exception. Both
bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's
pre-teardown terminal-run check only matched the literal passed and
checks-passed tokens, so this outcome fell through to unknown/parked
and a finished worker awaiting merge kept getting re-alerted as stale,
while an abort race during teardown could also leave a finished run
misreported as still parked.

Map passed-with-override to the same done/terminal handling as a
clean passed in both places.

* fix(document): Replace stale outcome mapping with authoritative pointer

* fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return

During the 2026-09-22 away window every exemption worker whose pull request
had merged was left sitting for nine hours. The supervision branch received
the stale wake, the merge-landed check, and the hourly inactive-outcome row
for each of them, ran the recovery playbook, found nothing to recover, and
reported "no further action". The branch prompt granted ordinary teardown of
a confirmed-landed task without ever naming the moment or the command, and
the playbook has no landed exit, so the stale path ended at "nothing to
recover". The return brief then listed only blockers, decisions, and the
latest five routine outcomes, so the landed workers stayed invisible after
the captain came back.

- bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale,
  inactive-outcome, or heartbeat row on a done task with a merged PR, as the
  moment to claim the lease and run bin/fm-teardown.sh with no flags; a
  refusal is reported, never forced or worked around. Add teardown to the
  handling tool list.
- stuck-crewmate-recovery: a landed worker is not a recovery case; point at
  the ordinary teardown owner for each actor.
- bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable
  records only (a live task record whose recorded PR carries the
  merge-notification marker), between could-not-fix and handled, without
  holding the gate; the afk skill's return step closes each listed task
  through ordinary teardown once the check clears.
- tests: pin the prompt rule in fm-branch-supervision and the brief section
  in fm-afk-return through the real marker writer.

* no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring

A green PR could sit unreported because neither the worker nor the
supervisor could observe checks-green while the ci step kept monitoring
for the merge.

Supervisor read: fm_nm_select_run's capped-overview inventory reader looked
the repository up by the task worktree path, but no-mistakes registers a
repository once by its main clone path and resolves every linked worktree
to it, so on every task copy of a busy repo the lookup matched no row and
each read reported "complete same-branch run inventory unreadable". Key the
lookup on the overview's own top-level `repo:` line, which every axi
release emits as the resolved working_path.

Even with a readable run, the ci-log classifier treated "base branch
advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs
a checks state only when it changes and a base advance does not clear
readiness, so a green PR read as still validating for as long as main kept
advancing. Stop treating that line as a marker, matching no-mistakes' own
ci-log parser, and name the run's PR URL in the held-for-merge reading so
the existing inactive-outcome path can act on it without a worker report.

Worker contract: `axi status` never reports checks-passed while the ci
step monitors for merge, so the definition of done no longer makes a
status poll the wait for the next gate or outcome; the drive call's own
return is the green signal, reattached with `no-mistakes axi run` after a
bounded return.

* no-mistakes(review): read the full ci log when checking checks-green

* no-mistakes(review): correct stale ci log tail wording in docs

* no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session

* no-mistakes(document): Document session-derived Lavish polling

* no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900)

* fix(bin): ignore vanished state scratch files on secondmate relaunch

Relaunch refused when find(1) exited non-zero while listing a secondmate
home's state directory. A live watcher can delete scratch files between
readdir and processing, which is not evidence that child *.meta records
are unreadable.

Prove the directory is listable from its mode and keep the existing
readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765.

* no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907)

* fix(bin): treat home-owned status closes as already read

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.

* no-mistakes(review): Keep owned closes in unread status; lock ledger writes

* no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake

* no-mistakes(review): Require real owned growth before ledger marks status seen

* no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS

* no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test

* no-mistakes(review): Align ledger docs and scope ledger to wake path only

* no-mistakes(review): Restore stranded historical-annotation test comment to its function

* no-mistakes(review): Retire the home-appends lock alongside its ledger

* no-mistakes(document): Note ledger's lock-helper dependency in classify library

* no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions

* no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger

* no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger

* no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350)

* chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest

Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and
lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the
floor-boundary test fixtures to the new versions.

tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final
expecting pr-merged, so add the regression test: a bound work that ends
failed reports its honest outcome text through fm-public-followup-emit.sh,
consume marks the commitment ready, and deliver posts that text exactly
once.

Also make two hang-guard tests in fm-backlog-atomicity portable to hosts
without coreutils timeout, and stop an installed herdr from leaking into
the secondmate-liveness husk classifier test.

* no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test

* no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures

* no-mistakes(document): Document failed public-followup delivery behavior

* no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878)

* fix(bin): refuse ship done: when the named head lives only in the worker copy

A ship done: is not current-state done until that exact commit is reachable
outside the disposable copy. The check tests the named head, not whether
some branch moved.

* fix(bin): gate CI-ready ship done: on named-head reachability, not handoff

Keep no-mistakes' first done: as the pipeline handoff, apply the same shared
check when registering a PR and when a secondmate publishes ledger-first,
treat a recorded merged PR as landed after prune, and name the PR head
instead of scanning free-text SHAs.

* no-mistakes(review): Bind named-head gate to recorded PR and forge heads

* no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries

* no-mistakes(review): Align worker done wording, test mapping, pending-retry test

* no-mistakes(test): Raise watcher test time limit to stop load flake

* no-mistakes(document): Restore ledger-path fact and name named-head gate coverage

* ci: re-attest named-head ship-done gate for a fresh serial-3 verdict

* no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery

* no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204)

* fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm

A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows.

* no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335)

The re-arm recovery cases judged "the watcher stayed live instead of
surfacing recovery" with fixed budgets below what a real stale-lock
recovery costs on a contended host: the arm's default 10s confirmation
deadline, a start helper that returned after about 4s whether or not the
arm had confirmed its watcher, and an 80-poll exit wait.
A changed-suite run beside other suites starves the recovery's many
short-lived processes while this suite's sleeping poll loops keep their
pace, so a watcher still surfacing its recovery read as one that stayed
live (issue kunchenguid#3793).
The original 0.25s window after confirmation was widened to 80 polls in
kunchenguid#3837, which left the same race at a larger size.

Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now
gives the arm an explicit 30s confirmation budget and waits for its
confirmation or exit within a ceiling that outlasts it, and every wait on
a re-armed watcher uses one named iteration-counted ceiling that outlasts
the same budget.
A passing case returns as soon as the arm reports or exits, and a watcher
that never surfaces its recovery still fails.

A new case delays every mktemp and readlink the re-armed watcher runs
after it publishes its beacon, so its first poll and exit take about 13s
on any host.
It fails with the reported symptom on the previous budgets and passes now.
No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.

* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374)

* fix(bin): submit our own stuck doorbell instead of skipping every later ring

* no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit

* no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat(bin): add the opt-in fleet activity ledger

Homes that create config/fleet-ledger get an append-only JSONL file,
state/fleet-ledger.jsonl, recording task.dispatched, task.status,
task.merged, and task.cleaned_up so outside tools can follow a fleet.
With the flag absent each producer does one file test and nothing else.
docs/fleet-ledger.md owns the record contract and its documented limits.

* no-mistakes(review): Record task.status text verbatim after the first colon

* no-mistakes(document): Clarify fleet ledger status and setup documentation

* no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
…nchenguid#5352)

* fix(bin): format, validate, and surface public-followup deliverables

brief pre-fills report_path=data/<work-id>/report.md and states the accepted
format of every value it cannot know instead of a bare <value> placeholder.
fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in
both the direct and staged destinations, naming the key, value, and format.
consume records the specific deliverable, outcome, or missing key behind a
tasks-axi refusal, and each refusal wakes the owning home once through the
existing relay poll.

* no-mistakes(review): refuse emits missing a required deliverable in both destinations

* no-mistakes(review): require promised deliverables and keep rejections recoverable

* no-mistakes(review): mirror tasks-axi's canonical pull request URL rule

* no-mistakes(review): keep a rejection wake whose line cannot be read

* no-mistakes(review): key emit-time rules on the promise, not the outcome

* no-mistakes(review): bound deliverable keys and values as tasks-axi does

* no-mistakes(review): state rejection wakes as at-least-once and pin it

* no-mistakes(review): enforce the promised contract tasks-axi holds at emit

* no-mistakes(review): stop inferring a staged promise from its outcome

* no-mistakes(document): Refresh public follow-up documentation

* no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks

* Revert unrelated CI auto-fix edits to the watcher and bearings board test

The CI step's automatic repair changed bin/fm-watch.sh and
tests/fm-bearings-board-render.test.sh to chase two intermittent CI
failures that also occur on main and are not part of this change. Restore
both files so this branch carries only the public-followup deliverable fix.

* no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing

* no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
mremond and others added 29 commits September 26, 2026 11:29
…henguid#5812)

* fix: declare worker background and pipeline waits

Require ship and scout workers to declare owned-work waits with the existing
paused verb before ending a turn or waiting on a pipeline or long command.
Keep the first-sight alert and existing liveness classification unchanged;
subsequent inspection follows the existing long pause cadence.

Validation: emitted brief regression failed before the instruction change
and passes afterward. Public watcher/drain regressions cover the first
alert, repeated wedge suppression, bounded rechecks, and undeclared idle
alarms using isolated backend fixtures. Brief suite, pinned lint, Bash
syntax, documentation inventory, and whitespace checks pass.
No real worker harness was exercised for wait behavior.

* fix(document): Clarify declared worker waits and documentation ownership

* fix(ci): Captain, fixed the cadence test to age both the declaration and first-alert throttle while preserving declaration identity. Reproduced the CI failure using stable identity; corrected tests pass with both stable identity and native macOS behavior. Focused ShellCheck and diff checks pass. Production behavior is unchanged; Linux CI was not rerun locally
…5770)

* fix(bin): bound each lint root in its own ShellCheck process

CI job "Lint 1" died twice at about ten minutes because the two shard
workers each packed about 110 canonical roots into one unbounded ShellCheck
process, and a byte-weight rebalance moved the analysis-heavy fm-watch.sh
into a partition with other heavy roots, so the pair outgrew the 16 GiB
runner before anything could name a culprit.

Run one canonical root per ShellCheck process under an enforced envelope:
a wall deadline plus terminate-then-kill grace via the shared
fm-timeout-lib.sh watchdog, and a per-root rlimit spec applied inside the
child before exec (default a 4 GiB address-space cap, so two workers stay
inside a 16 GiB job with headroom). A root that exceeds the envelope fails
by name with a recorded reason - timeout, memory, signal, or
limit-unavailable - instead of taking the runner down. The per-root
watchdog runs in its own process group so the owner's group sweep cannot
orphan the bounded subtree, and fm_exec_timed now starts the same
escalation when its parent dies before it can be signalled.
FM_LINT_REQUIRE_BOUNDS=1, set in CI, refuses the run outright when a
configured bound cannot be enforced on the host rather than lint uncapped.
Each root's begin/end, reason, duration, and peak RSS stream to stderr in
partition mode and append to a retained <telemetry>.roots.tsv sidecar
uploaded beside the partition telemetry.

Coverage is unchanged: pinned ShellCheck 0.11.0, --norc, --external-sources
full analysis, complete and disjoint partition inventory, workflow lint,
and the backend-purity check, with byte-identical diagnostics across
jobs=1/2 proven by tests/fm-lint.test.sh.

* fix(bin): fail closed on unenforceable lint bounds and size the cap

Required-bounds mode (FM_LINT_REQUIRE_BOUNDS=1, set by CI) now refuses the
run with named errors before any root starts: a missing fm-timeout-lib.sh,
a watchdog that cannot actually bound a probe command, or a host that
rejects the address-space limit all stop the run rather than lint uncapped.
The generalized FM_LINT_ROOT_RLIMITS flag:value interface is replaced by a
single FM_LINT_ROOT_MEMORY_KIB, and the roots sidecar and telemetry record
the run's final exit status after backend-purity and workflow checks
instead of the pre-check lint status.

The default cap is 6 GiB of address space per root, not 4 GiB: ulimit -v
bounds virtual address space rather than resident memory, and ShellCheck's
GHC runtime keeps roughly a third of that space as reservation, so 6 GiB
yields about a 4 GiB working heap budget. A Linux measurement during this
change showed eleven real canonical roots running out of memory under the
earlier 4 GiB cap while the largest passing root peaked near 2.8 GiB
resident; two 6 GiB roots plus runner overhead still fit the 16 GiB job.
Roots that still exceed the cap keep failing by name, and the sidecar's
per-root peak RSS keeps roots approaching the budget visible.

tests/fm-lint.test.sh now proves the memory primitive where it can be
proven: on hosts that accept ulimit -v a perl allocator is refused under a
256 MiB limit and reported by name as a memory death, the pinned ShellCheck
lints a small file under the configured cap and is named when a far smaller
cap binds it, and a watchdog-less copy refuses under REQUIRE_BOUNDS; the
bounded cases skip on macOS, which cannot enforce the address-space limit.

* no-mistakes(review): Prove memory cap binds, pass watchdog owner, drop unused modes

* no-mistakes(review): Capture watchdog owner before startup for every fm_exec_timed caller

* no-mistakes(document): Clarify bounded lint documentation and telemetry

* no-mistakes(document): Correct bounded lint documentation and sidecar path

* docs(bin): restore the per-root memory cap sizing rationale

The pipeline's document step rewrote the ROOT_MEMORY_KIB comment and
dropped the sizing reasoning the change is required to record: address
space vs resident memory, the GHC reservation share, the measured 4 GiB
failures and ~2.8 GiB peak, and the two-roots-plus-runner capacity
arithmetic. Restore it beside the default while keeping the corrected
"not a resident-memory ceiling" framing.

* no-mistakes(review): Document memory cap RSS reduction threshold and first candidate

* no-mistakes(review): Scope owner-death escalation docs to the perl watchdog

* no-mistakes(document): Clarify bounded lint and timeout documentation

* no-mistakes(review): Install perl watchdog signal handlers before forking the command

* no-mistakes(document): Correct bounded lint documentation and stale watcher comments

* no-mistakes(ci): Fixed the supervision-host test’s obsolete expectation: the watchdog now reaps an engine when its host dies. The timeout and supervision-host tests pass locally; the watcher test also passes locally. Lint 1 and 2 remain unresolved: seven canonical roots exceeded the required 6 GiB address-space cap in CI. I did not raise the cap, exempt roots, or reduce source-following coverage to make those failures disappear

* no-mistakes(ci): The two lint checks failed when eight canonical roots hit the enforced memory cap. I reduced repeated ShellCheck source-graph expansion while keeping runtime imports and the canonical root inventory intact. Pinned ShellCheck passes for all changed roots; the relevant local tests pass. The 6 GiB Linux CI run remains unverified

* no-mistakes(review): Restore source directives, raise cap to 8 GiB, classify OOM

* no-mistakes(review): Classify memory deaths from root stderr, not source excerpts

* no-mistakes(review): Match only whole runtime memory-error lines for memory reason

* no-mistakes(document): Clarify lint memory classification in script documentation

* no-mistakes(ci): Fixed both lint checks’ memory-limit failure: each CI lint job now runs one root at a time with a 12 GiB address-space cap. Kept the local two-worker default and updated the sizing comment and test expectation. The lint tests and workflow validation pass locally; Linux CI remains to confirm the heavy roots
…uid#5732)

* fix(bin): bound the watcher cleanup marker-lock wait

tests/fm-watch-triage.test.sh intermittently failed serial CI shard 1
with "watcher pid <pid> did not exit within 10s of TERM". The watcher
had processed the TERM and was inside watcher_cleanup, where the
recovery-marker publish waits on state/.watcher-down.lock through an
unbounded fm_lock_acquire_wait. A live foreign holder of that lock
leaves the TERM'd watcher spinning in its own EXIT trap until the lock
frees or a second signal short-circuits the trap.

fm_recovery_transition now takes an optional bound and both
release-lock paths plus publish honour it through a new in-process
fm_lock_acquire_wait_max. watcher_cleanup passes
FM_WATCHER_CLEANUP_LOCK_BOUND (default 2s); on timeout the publish is
skipped, the singleton stays behind as ordinary dead-pid evidence, and
the next arm's clear-stale-lock still republishes it.

Regression test drives a real watcher with .watcher-down.lock held by
a live foreign process and asserts a single TERM still stops it.

* no-mistakes(review): Parse watcher cleanup lock bound as decimal, zero defaults

* no-mistakes(review): Pin cleanup bound tests to observed marker-lock contention

* no-mistakes(document): Document bounded watcher cleanup and recovery

* no-mistakes(document): Clarify bounded watcher cleanup and recovery documentation

* no-mistakes: apply agent fixes

* no-mistakes(review): Arm marker-lock FIFO before TERM; drop FM_TEST_ONLY_LATE

* no-mistakes(review): Hold marker lock through a failed cleanup acquire
…kunchenguid#5845)

* test: stop the leaked unreachable watcher before remote e2e cleanup

The remote secondmate lifecycle e2e backgrounded fm-watch.sh through the
remote_env shell function, so $! named the function's subshell rather than
the watcher. Killing that subshell left the unreachable-leg watcher running,
and its one-second liveness probe kept invoking the fake ssh, which rewrites
ssh.count in the temp root. When a probe landed while the EXIT trap was
removing the root, rm failed with "Directory not empty" after every
assertion had passed.

Exec the watcher from the backgrounded function so the recorded pid is the
watcher itself, and assert the stopped watcher stops probing and writing its
state. Cleanup also stops a watcher left running by a failed assertion and
removes the root through fm_test_remove_tree, so a run that fails before
retirement does not strand the read-only spawn hooks directory.

Closes kunchenguid#5836

* no-mistakes(review): Clear reaped watcher PIDs and restore plain temp-root removal

* no-mistakes(review): Let in-flight probe settle before stopped-watcher baseline
…nchenguid#4806)

* fix: stop quarantining ordinary shared-captain source updates

* no-mistakes(document): Rewrap remote inherit header so usage prints fully
* docs: make calm easier to read

Restructure the Calm mode prose into sections, lists, and tables without changing documented behavior. Every original heading, anchor, fenced code block, inline-code span, link target, number, and quoted string is preserved.

* docs: restore reload case in calm override lead-in

The Built-in tool override collisions lead-in covers a session that reloads with Calm already on, as the original text did.
* docs: make turnend-guard easier to read

Restructure the turn-end guard doc's prose into shorter sections, lists, and tables without changing documented behavior. Every original heading, anchor, inline identifier, link target, and number is kept.

* no-mistakes(review): Restore legacy-only scope on TERM retirement sentence

* no-mistakes(review): Name Cursor park behavior in live e2e test line
…henguid#5872)

* docs: move situational AGENTS.md sections into on-demand skills

Backpass memory optimization: shrink the always-loaded AGENTS.md by moving
situational contracts (home layout, session-start recovery, validation and
landing supervision, scout completion, away/quiet supervision, Relay
ownership) into agent-only skills loaded at their triggers, with a trigger
index skill.

* docs: classify the new on-demand skills' documentation audience

Register the seven new agent-only skills as agent-runtime docs and fix a
link in validation-supervision that kept its AGENTS.md-relative path.

* docs: close load-timing gaps found by the live regression check

- load validation-supervision whenever an ask-user finding is decided or
  answered, so forbid --yes and process-every-return reach the worker
- keep the mid-task captain-ask rule, the unconfirmed network-checks rule,
  and the worker account pin rule inline in AGENTS.md
- fix cross-references that still pointed at moved AGENTS.md sections
…guid#5879)

* fix: route second-mate signal wakes by their new status span

A second mate's status log is a shared channel carrying many independently
keyed decisions, so judging its signal rows by every decision still open in
the whole log pinned each routine update to main behind any unrelated
parked hold. scopeForUnreadWake (the one owner for Pi and the attended
supervision host) now judges a second-mate signal row by the lines presented
since the last drain, bounded by the existing status-presentation cursor:
a decision, blocked, resolution, or captain-held line, or a line declaring
the key of a still-open decision, keeps the whole row on main, and any
cursor problem falls back to the whole log. Keys are read only at the
status parser's declared positions, with readable time stamps stripped as
bin/fm-classify-lib.sh does. Single-task crewmate and stale routing are
unchanged, and stale and signal rows for one mate keep independent verdicts.

The supervision branch now treats a second mate's done and merged lines as
relayed child outcomes, and fm-teardown refuses the branch actor second-mate
retirement through the existing role-partition helper in both postures.

* no-mistakes(review): Route second-mate resolutions to main only when closing open decision

* no-mistakes(review): Guard bare-verb fold lines and fold resolution spans incrementally

* no-mistakes(document): Clarify second-mate wake routing and retirement documentation

* no-mistakes(ci): The CI failure was a timing-sensitive watcher teardown test, not the PR’s signal-scope code. Extended the bounded wait for both state-directory and home removal. The watcher suite passes locally, and git diff --check is clean
…extension (kunchenguid#5882)

register-extension took the extension lifecycle lock and then the source
lock, while reconcile republishing an unhandled extension result holds the
source lock and reaches the lifecycle lock through the extension host's
process-event path. Both waits are unbounded and both owners stay alive, so
the two could wait on each other forever and freeze the home's monitoring
cycle.

register-extension now takes the source lock first, matching every other
path that holds both. The lifecycle lock still spans binding resolution
through registration publication, so binding retirement stays serialized.

A new lifecycle-order section in the extension-binding suite, run in the
default aggregate, holds a re-registration inside binding resolution while
reconcile republishes that source's unhandled result and requires both to
finish within a bound.

Fixes kunchenguid#5866
* fix(bin): run the repository's own hooks when git -c carries the per-task hooksPath

The per-task hooks wrappers cleared only the GIT_CONFIG_COUNT override before
looking up the repository's own hooks directory. When core.hooksPath reached git
through git -c (GIT_CONFIG_PARAMETERS), directly or inherited by a child
process, the lookup found the wrapper directory again and exited 0, so the
repository's real hook - such as a pre-push publish guard - never ran and the
push succeeded.

The lookup now ignores GIT_CONFIG_PARAMETERS too, so only the repository's
config files decide its hooks directory, and a failed lookup exits nonzero
instead of skipping the hook. AI-trailer stripping is unchanged.

Fixes kunchenguid#5871

* no-mistakes(document): Document git hook chaining and lookup failure behavior
…unchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure
…hrashing guard (kunchenguid#5903)

* feat(jev): add the guard framework and the memory RSS/swap thrashing guard

A Jev guard is a bounded read-only host diagnostic that turns one class of
resource pressure into a machine-readable audit record and a one-line verdict.
This lands the framework contract (docs/jev-guards.md) with one representative
family (Pattern 46, memory RSS and swap thrashing) as the pilot: thin bash
wrapper, stdlib-only python engine, and a behavioral test through the CLI.

* no-mistakes(review): fix jev mem guard fail-open unknown and contract

* no-mistakes(review): fix mem guard rss pairing, memfree, tests, docs

* no-mistakes(review): fix inverted UNKNOWN branch in fail-forcing test

* no-mistakes(review): register docs/jev-guards.md in audience inventory

* no-mistakes(review): simplify wrapper, drop dead top_n, single-source thresholds

* no-mistakes(review): assert exact exit code in fail-forcing test leg

* no-mistakes(review): tolerate any stdout encoding in text output

* no-mistakes(document): Fix guard contract dash style and output wording
…enguid#5900)

* fix(bin): stop slow GitHub reads from starving and waking the contributions poll

The poll's fixed budget cut the tail URL's reads short, and any read killed at the bound was recorded as a forge failure, so routine GitHub slowness produced an observation-unavailable wake. A configured FM_CONTRIBUTIONS_BUDGET now rides the generated check shim and is cut down to the watcher's per-check bound with a margin, a read killed at the bound or the deadline is budget refusal that leaves records untouched, and the poll moves to the next URL that still has a full observation reserve instead of ending.

* fix(bin): report the bound when a signal death leaks through fm_run_timed

fm_run_external_timeout trusted the wrapper's recorded status over the runner's own bound verdict, so a read killed by the bound's TERM could surface as 143 and callers such as the contributions poll classified a budget-bound read as a forge failure. When the runner reports the bound, a signal-death status is the bound's own TERM racing the wrapper's bookkeeping and is reported as 124; a natural exit still passes through. The replace-shell probe also moves to the repo's portable BASHPID fallback so the suite runs on the macOS system bash.

* no-mistakes(review): Rotate contribution polling and verify generated budget behavior

* no-mistakes(review): Stabilize contribution rotation across successful observation refreshes

* no-mistakes(review): Exclude settled contributions from live observation rotation

* no-mistakes(document): Document contribution poll rotation and observation reserves

* no-mistakes(ci): Captain, fixed timeout handling with a one-line change preserving natural exit 137. Full session-start and contributions suites, targeted regressions, and lint passed. The timeout suite still fails on the known Bash 3.2 BASHPID issue, left unchanged as instructed. Logs retained in .no-mistakes/ci-evidence/. Remote CI was not rerun

* no-mistakes(ci): Fixed the fixture’s lock-acquisition race with a one-line bounded wait. Forced contention reproduced the CI error before the fix and passed afterward; the ordinary held-lock case, lint, syntax, and diff checks passed. Local Bash 3.2 failures remain: “cleanup lock bound 08 gave up before the marker lock freed” (also reproduced without the fix) and “TERM did not stop a watcher blocked inside a poll”. Remote CI was not rerun
…railers (kunchenguid#5859)

* feat: add keep AI trailers setting

* no-mistakes(document): Drop duplicate keep-ai-trailers line, update Cursor attribution note

* no-mistakes(document): Honor keep-ai-trailers for Devin worker attribution

* no-mistakes(document): Qualify Devin attribution note with keep-ai-trailers flag

* no-mistakes(review): Inherit keep-ai-trailers into secondmate homes
…sh-command popup cannot hide the composer (kunchenguid#5876)

* Fix Herdr composer reads blinded by the slash-command popup

Capture the full visible viewport for every Herdr composer state and content read instead of a bounded tail.
Claude Code renders its slash-command popup between the composer and the pane bottom, which pushes the composer outside a tail window.
The pre-Enter payload proof then read an empty composer, judged a typed command unsent, and cleared it without pressing Enter.
The proof-lines value now bounds only the clear cost, not the capture size.
Growing the window adds rows above the composer only, so bottom-most shape selection and prior verdicts are unchanged.
The suffix refusal is kept, and the shared inbox pending-line read stays a bounded tail.
Portable regressions cover the popup-below-composer layout, and the live submit-confirmation guard gains a third exit scenario.
The runtime-backends verification record documents the Herdr 0.9.0 and Claude Code 2.1.283 run.

Closes kunchenguid#5533

* no-mistakes(document): Clarify composer capture bound ownership

* no-mistakes(review): Drop unrequired bracketed-paste Enter fallback from live guard

* no-mistakes(review): Latch trust prompt, check idle composer arm first
…configurable bound (kunchenguid#5917)

Each per-task endpoint read in the session-start fleet digest runs in its own crash-isolated child under the existing timeout helper with a configurable FM_SESSION_START_ENDPOINT_TIMEOUT bound (default 10s).

A hung or killed read becomes that task's endpoint error while the digest continues, and the wrapper banners any abnormal child exit naming its status.
…nchenguid#5886)

* Keep lab tmux sockets on short private paths

* no-mistakes(review): Fix Linux stat checks and fail-closed lab tmux teardown

* no-mistakes(document): Update lab helper documentation for isolated tmux sockets
* Allow silent task-level no-change outcomes

* no-mistakes(review): Exclude silent outcomes from captain-return handoffs

* fix: look up supervision receipts by exact sequence

* no-mistakes(review): Suppress silent notes in away-return brief

* no-mistakes(review): Clarify visible notes; remove unused mode

* no-mistakes(review): Clarify silent outcomes and avoid false drain promises

* no-mistakes(document): Clarify silent supervision outcome documentation
…chenguid#5928)

* feat: make /quiet a statement where the attended supervision host runs

On a home that opted into the supervision host, quiet mode is what the
attended host already does, so /quiet now enters nothing there instead of
launching the quiet daemon and writing a record that would park a present
captain's main.

- bin/fm-afk-launch.sh quiet-check says quiet mode needs nothing where the
  attended host runs, or that the session is paused while its
  broken-session latch holds; a quiet enter refuses there before writing
  anything.
- Where the home opted in but the attended host lacks a part (engine,
  tools, verified mirror writer, identifiable main session, valid mirror),
  quiet-check names it and quiet mode falls back to the daemon.
- Under a live away record on that home, quiet-check and a quiet enter
  refuse and name the record, so the return runs first, whatever
  state/.afk says.
- A quiet enter records mode: quiet in the posture record, so start and
  start-native launch the quiet daemon without FM_AFK_MODE, and the away
  refusal wording fires only for away.
- bin/fm-host-mirror.sh check validates the dialog mirror read-only and
  exits 1 on a missing, unreadable, or invalid mirror.
- The quiet and afk skills and the supervision-host docs describe the new
  behavior; homes without the opt-in and Pi homes keep the daemon path.

* no-mistakes(review): Archive the quiet record when a quiet daemon start fails

* no-mistakes(document): Clarify quiet-mode documentation and remove stale duplicates
…uid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference
…nguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper
…id#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally
…uid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass
…enguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed
Preserve both fork and upstream ancestry. Drop the superseded remote-inbox delta in favor of upstream, retain the per-task pool layer without overriding home-wide pins, and verify setup tokens by real authentication effect. Keep fresh-store bypass consent local to launches that already selected bypass mode.
@Freudator86
Freudator86 merged commit 1c879b2 into main Sep 28, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.