Skip to content

Merge upstream firstmate into the fork - #46

Merged
cloud-practitioner merged 14 commits into
mainfrom
fm/fm-merge-upstream-4
Oct 7, 2026
Merged

cloud-practitioner merged 14 commits into
mainfrom
fm/fm-merge-upstream-4

Conversation

@cloud-practitioner

Copy link
Copy Markdown
Owner

Intent

fm-maint: merge upstream into the fork.

This answers the recommendation to merge upstream now: one PR to cloud-practitioner/firstmate that brings kunchenguid/firstmate main into the fork under the same rules as the earlier upstream-merge PRs 35 and 39 - a true merge of upstream main into fork main carrying upstream changes only (no fork-made changes and no pipeline auto-fix content), landed as a merge commit, never squashed or rebased, so fork history stays intact; title "Merge upstream firstmate into the fork".
The upstream changes are the 13 upstream PRs merged since fork PR 39 (upstream tip 237e1cf at the time of the recommendation): kunchenguid#6654 (a repeat missed second-mate reply escalates again after a resolve), kunchenguid#6733 (an exited worker no longer reads unknown forever), kunchenguid#6442 (tasks can start from a named base branch), kunchenguid#5354 (opt-in per-task pipeline spend record), kunchenguid#6725 (opt-in daily startup-file growth check, plus a Pi Calm test wait fix), kunchenguid#6666 (refuse a confirming Enter into Claude's background-task exit picker), kunchenguid#6647 (fallback path to the coding-guidelines skill), kunchenguid#6699 (project init runs in a subshell), kunchenguid#6655 (retire a contribution whose forge object is gone), kunchenguid#6708 (a later remote-reply break reopens after a resolve), kunchenguid#6639 (remote-reply recapture gets up to three waits), kunchenguid#6484 (fake ps PID labelling test fix), kunchenguid#6028 (watchdog no longer reads BASHPID).
Resolve the five conflicts keeping both sides:

  1. bin/fm-backend.sh fm_backend_send_text_submit: keep the fork's endpoint-ready guard first (as fm_backend_endpoint_ready "$backend" "$target" "$label" || return 1), then upstream fix: refuse a confirming Enter on the Claude background-task exit picker kunchenguid/firstmate#6666's block unchanged.
  2. bin/fm-spawn.sh relaunch key list: keep both herdr_process_identity (fork) and base_branch (feat(bin): start tasks from a named base branch kunchenguid/firstmate#6442).
  3. bin/fm-test-run.sh pr-forge family: list both fm-pr-bitbucket.test.sh and fm-pipeline-spend.test.sh.
  4. tests/fm-remote-reply.test.sh end of file: upstream fix(bin): reopen the remote-reply continuity decision on a later break kunchenguid/firstmate#6708's new cases first, then the fork's stop-every-fixture-worker block last.
  5. tests/fm-review-diff.test.sh: keep both the fork's Bitbucket cases and feat(bin): start tasks from a named base branch kunchenguid/firstmate#6442's base-branch case and both call lists, re-lettering upstream's header case, without losing the closing brace of the last fork function.
    List each conflict and its resolution in the PR, plus these follow-ups that are not part of this PR: optionally drop the fork's own remote-reply wait raise (fork PR 27) now that test: extend remote-reply whole-log recapture waits kunchenguid/firstmate#6639 landed; rebase the open upstream Bitbucket offer feat(bin): support Bitbucket Cloud pull requests in the PR scripts kunchenguid/firstmate#6327, which now conflicts; the fork's zsh fallback (fork PR 17) has no upstream path since upstream fix(bin): resolve backend adapter sibling libs when sourced under zsh kunchenguid/firstmate#6279 closed unmerged.

What Changed

Conflict Resolutions

  1. bin/fm-backend.sh: keep fm_backend_endpoint_ready "$backend" "$target" "$label" || return 1 before upstream fix: refuse a confirming Enter on the Claude background-task exit picker kunchenguid/firstmate#6666’s unchanged dialog-check block.
  2. bin/fm-spawn.sh: retain both the fork’s herdr_process_identity and upstream’s base_branch in the relaunch key list.
  3. bin/fm-test-run.sh: retain both fm-pr-bitbucket.test.sh and fm-pipeline-spend.test.sh in the pr-forge family.
  4. tests/fm-remote-reply.test.sh: place upstream fix(bin): reopen the remote-reply continuity decision on a later break kunchenguid/firstmate#6708’s new cases before the fork’s final stop-every-fixture-worker block.
  5. tests/fm-review-diff.test.sh: retain the fork’s Bitbucket cases and upstream’s base-branch case, including both call lists; re-letter the upstream header case to (j) and preserve the last fork function’s closing brace.

Follow-ups Outside This PR

  1. Optionally remove the fork’s remote-reply wait increase from fork PR test: stop leaked supervision-host cycles per case and widen the remote-reply recapture wait #27 now that upstream test: extend remote-reply whole-log recapture waits kunchenguid/firstmate#6639 is included.
  2. Rebase the now-conflicting upstream Bitbucket offer: feat(bin): support Bitbucket Cloud pull requests in the PR scripts kunchenguid/firstmate#6327.
  3. The fork’s zsh fallback from fork PR fix(bin): load backend sibling libraries when fm-backend.sh is sourced from zsh #17 has no upstream path because upstream fix(bin): resolve backend adapter sibling libs when sourced under zsh kunchenguid/firstmate#6279 closed unmerged.

Risk Assessment

⚠️ Medium: The broad upstream merge changes worker lifecycle and reply-recovery behavior and adds optional monitoring and accounting, but the required conflict resolutions are preserved and static review found no substantiated blockers.

Testing

All requested syntax, coverage, and four targeted suite checks passed without gate skips. Captured fixture-backed CLI output and refusal-state evidence, removed all temporary fixtures, and left no source changes. The operator reported an earlier nine-suite pass on this exact commit; it was not independently rerun. No live harness, Herdr lab, or rendered UI was exercised, per the operator’s narrowed scope. These automated observations do not establish live results for any of the five scenarios, so all five are untested under the live-validation contract.

  • Live validation: ⚠️ inconclusive - 0 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Review a task based on a named branch and see only the task’s changes ⏸️ untested no The prior payload did not establish a live result. tests/fm-review-diff.test.sh passed the recorded-base-branch case and stale-head and corrupt-branch checks, but executed the review CLI against dis…
Review a Bitbucket PR using its current head, with explicit fallback warnings when unavailable ⏸️ untested no The prior payload did not establish a live result. All three retained Bitbucket cases in tests/fm-review-diff.test.sh passed using a stubbed Bitbucket response; no live forge was contacted. Live Bit…
Start a task from its named base branch and refuse missing or contradictory base selections ⏸️ untested no The prior payload did not establish a live result. Named-base, brief-mismatch, and missing/local-only-base cases in tests/fm-spawn-pool-base-freshen.test.sh passed using disposable Git state and a f…
Exit Claude without sending a confirming Enter into its background-task exit picker ⏸️ untested no The prior payload did not establish a live result. Already-open, submit-opened, and delayed-render picker cases in tests/fm-control.test.sh passed against modeled terminal transitions. A real Claude…
Resolve a missed secondmate reply, then reopen the decision when the same failure recurs ⏸️ untested no The prior payload did not establish a live result. The same-kind escalation after operator close case in tests/fm-pending-reply.test.sh passed, including suppression of duplicate escalation while op…
Evidence: Conflict-focused suite transcript, including actual spawn CLI output, persisted metadata, and refusal-state evidence; terminal and forge services are fixtures, not live

Source: Conflict-focused suite transcript, including actual spawn CLI output, persisted metadata, and refusal-state evidence; terminal and forge services are fixtures, not live

FM_TEST_BEGIN 2026-10-07T07:50:11Z tests/fm-review-diff.test.sh family=pr-forge expected_gate_skip=none
ok - fm-review-diff falls back to recorded pr_head when pull head cannot be fetched
ok - fm-review-diff fetches refs/pull/<n>/head when pr_head= is absent
ok - fm-review-diff prefers freshly fetched PR head over a stale recorded pr_head=
ok - fm-review-diff without pr= keeps the worktree-branch diff
ok - fm-review-diff falls back to local branch with a warning when PR head is unreachable
ok - fm-review-diff reviews the meta-recorded ship branch even when the worktree HEAD moved off it
ok - fm-review-diff refuses a corrupt recorded ship branch instead of reviewing the wrong content
ok - fm-review-diff reviews a Bitbucket pull request's live head over a stale recorded pr_head=
ok - fm-review-diff falls back to a Bitbucket task's recorded pr_head= only with a warning
ok - fm-review-diff claims a Bitbucket recorded-head fallback only when it uses that head
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no live watcher process holds this home lock (last beat: 0s ago).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair a missing or failed watcher cycle with the Pi tool fm_watch_arm_pi, or restart Pi with -e ~/.no-mistakes/worktrees/450411b3e67c/01M4AK0KYBPP2H30CZ1HR0HKJ4/.pi/extensions/fm-primary-turnend-guard.ts -e ~/.no-mistakes/worktrees/450411b3e67c/01M4AK0KYBPP2H30CZ1HR0HKJ4/.pi/extensions/fm-primary-pi-watch.ts if the extensions are not loaded.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - fm-review-diff compares a task against its recorded base branch
FM_TEST_END 2026-10-07T07:50:25Z tests/fm-review-diff.test.sh exit=0 duration_ms=14148 gate_skip=false
FM_TEST_BEGIN 2026-10-07T07:50:25Z tests/fm-spawn-pool-base-freshen.test.sh family=standalone expected_gate_skip=none
FM_TEST_BEGIN 2026-10-07T07:50:25Z tests/fm-pending-reply.test.sh family=standalone expected_gate_skip=none
ok - normal correlated reply resolves once (idempotent)
ok - completed turn with no report triggers exactly one recovery
ok - recovery never pokes a mate waiting on its own decision, and runs once it closes
ok - recovery sends during an open decision when config/wait-no-turns is absent
ok - recovery grace is measured from the request turn's completion, not delivery
ok - one fresh status read immediately before firing catches a just-landed reply, any verb
ok - a resolve that fails after committing resolved still blocks repost and escalation
ok - missed-report escalation grace is measured from the recovery turn's completion
ok - recovery attempts reconcile without reinjection
ok - recovery reply resolves the original expectation
ok - second missed turn escalates once and remains durable
ok - escalations and replies wake; the home's own escalation close stays quiet
ok - failed escalation publication remains retryable and publishes once
ok - legacy escalation closes under the shared default key
ok - legacy escalation cannot close an unrelated default-key decision
ok - foreign correlated blocker cannot impersonate a pending-reply escalation
ok - concurrent resolution closes one keyed escalation exactly once
ok - concurrent escalation yields to a late correlated reply
ok - transport success cannot masquerade as reply success
ok - undelivered records remain immutable across scan paths
ok - delivery confirmation fallback reconciles durably
ok - delivery confirmation serializes with reconciliation
ok - unrelated events and stale correlation ids cannot resolve
ok - restart preserves expectation and exact parent destination
ok - wrong-home reports are detected but do not silently acknowledge
ok - direct unmarked captain input creates no expectation
ok - fm-send marked secondmate path creates pending and embeds corr
ok - status-pointed document resolves the expectation
ok - optional helper report resolves without being required for correctness
ok - backend busy/idle observation covers Pi/Claude paths without conversation scrape
ok - tmux and zellij unknown states use bounded capture fallback
ok - pending replies scope Kimi capture fallback by recorded harness
ok - tick skips terminal records and reuses target observations
ok - the tick leaves settled records alone and still does the selected records' work
ok - correlations are reused only for matching open task records
ok - tick end-to-end: miss -> one recovery -> escalate -> durable
ok - failed transport discards undelivered expectation only
ok - a remote repost waits for the reply channel and still fires on a real miss
ok - a mirrored correlated remote reply resolves without any repost
ok - same-basename self-home corr= is restated onto the parent channel and resolves
ok - same-basename reply resolves at the recovery failure boundary
ok - a child-file mate-home sighting is not copied and still escalates
ok - mechanical helper writes the parent channel from verb, corr, and note
ok - remote parent-replies.status is not classified as wrong-home
ok - local parent-replies.status remains wrong-home evidence
ok - an escalated correlation stays retryable only while undelivered
ok - a same-kind escalation after an operator close opens the decision again
ok - all pending-reply tests passed
FM_TEST_END 2026-10-07T07:51:47Z tests/fm-pending-reply.test.sh exit=0 duration_ms=81840 gate_skip=false
# remote-seeded Treehouse spawn command
$ FM_HOME=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/remote-seeded/home bin/fm-spawn.sh pool-remote-seeded-r13 /tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/remote-seeded/project --scout
spawned pool-remote-seeded-r13 harness=codex kind=scout window=firstmate:fm-pool-remote-seeded-r13 worktree=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/remote-seeded/pool
exit=0
published worktree=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/remote-seeded/pool
resolved project lock=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/remote-seeded/home/state/.treehouse-project-038e0021d39a55495b905fcdbc146ddcf1c57373.lock
ok - a remote-seeded secondmate home allocates and launches from its Treehouse pool
ok - a Treehouse slot claim names the launched task, refuses when unclaimable, and is dropped by a locked abort
# evidence begin: linked-home spawn, returned=primary
$ bin/fm-spawn.sh pool-linked-primary-r12 /tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/linked-primary/secondmate --scout
error: treehouse get did not enter an isolated worktree within 60s (last seen '/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/linked-primary/project': it is the repository's primary checkout (its git dir is the spawning project's common git dir); spawning project '/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/linked-primary/secondmate'); inspect window firstmate:fm-pool-linked-primary-r12
exit=1
primary HEAD before=5045d2a8c96be4c4e3d0f54a6ab697a53678233d after=5045d2a8c96be4c4e3d0f54a6ab697a53678233d
primary reflog before:
5045d2a HEAD@{0}: commit (initial): initial
primary reflog after:
5045d2a HEAD@{0}: commit (initial): initial
FETCH_HEAD absent
task metadata absent
# evidence end
ok - linked spawning home: primary preserves the primary before any refresh
# evidence begin: linked-home spawn, returned=primary-alias
$ bin/fm-spawn.sh pool-linked-primary-alias-r12 /tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA

... [3898 bytes truncated] ...

ew branch is created
ok - --base-branch picks the copy's starting point and is recorded; a missing or local-only base refuses
ok - a spawn refuses when --base-branch and the brief's Base branch lines disagree
ok - a Base branch line in the captain's intent neither blocks nor redirects a spawn
ok - a based scout on a forge=gerrit project is refused at spawn
ok - a stale pooled worktree resolves and refreshes a non-main default branch
# observed direct-pr spawn: spawned pool-direct-pr-r3 harness=codex kind=ship mode=direct-PR yolo=off window=firstmate:fm-pool-direct-pr-r3 worktree=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/direct-pr/pool
# observed scout spawn: spawned pool-scout-r3 harness=codex kind=scout window=firstmate:fm-pool-scout-r3 worktree=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/scout/pool
ok - direct-PR ships and scouts both refresh stale pooled worktrees before launch
# observed dirty refusal: error: pooled worktree '/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/dirty-refusal/pool' is not clean; refusing to discard uncommitted work while refreshing its base; preserved=keep this local work
ok - a dirty pooled worktree is refused without discarding its local work
# observed unresolved-default refusal: error: could not resolve origin's current default branch for pooled worktree '/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/unresolved-default/pool'; refusing to launch from a potentially stale base
ok - an unresolved remote default branch refuses the pooled worktree
# observed unreachable-origin refusal: error: could not fetch origin for pooled worktree '/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/unreachable-origin/pool'; refusing to launch from a potentially stale base
ok - an unreachable origin refuses a potentially stale pooled worktree
# observed origin-less launch: spawned pool-originless-r6 harness=codex kind=ship mode=no-mistakes yolo=off window=firstmate:fm-pool-originless-r6 worktree=/tmp/fm-merge-checks.MBIQal/fm-test-run.N0fy4c/w2/tmp/fm-spawn-pool-base-freshen.SpkTA7/originless/pool
ok - an origin-less pooled worktree launches as-is, skipping the freshness gate
ok - a dirty origin-less pooled worktree is refused without discarding its local work
ok - an origin configuration without a URL refuses the pooled worktree
ok - an empty origin configuration section refuses the pooled worktree
ok - an empty-only included origin section documents the accepted detection boundary
ok - an inactive conditional origin include leaves the pooled worktree origin-less
# observed stale-pin refusal: error: submodule 'ui' is checked out at a7e45cd5ebafe166cd4b96bc2e2acff43d253ae1, but this base records a2ba0ffea7ff1515c7f86813918b8434bf8edefc
ok - an origin-less pool with a stale submodule pin refuses while naming both pins and no remedy
ok - an unpushed submodule commit keeps the uncommitted-work refusal and survives it
ok - work inside a submodule is still refused as uncommitted work, not called stale
ok - a stale pin carrying real work is refused conservatively, never called stale
ok - a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line
# all fm-spawn-pool-base-freshen tests passed
FM_TEST_END 2026-10-07T07:52:22Z tests/fm-spawn-pool-base-freshen.test.sh exit=0 duration_ms=116537 gate_skip=false
FM_TEST_BEGIN 2026-10-07T07:52:22Z tests/fm-control.test.sh family=backend-dispatch expected_gate_skip=none
ok - fm-control exit: every verified harness gets its own verified exit command
ok - fm-control interrupt: every verified harness gets its own verified key and repeat count
ok - fm-control Devin interrupt: second press only after the armed hint, then busy invalidated without fabricating idle
ok - fm-control Devin interrupt: an idle agent gets one Escape and reports not-running
ok - fm-control Devin exit: a busy record whose turn already ended never opens the revert picker
ok - fm-control Devin interrupt: a revert picker a press opened is closed with Escape, never Enter
ok - fm-control Devin: an open revert picker refuses every typed command
ok - fm-control interrupt: opencode needs a double Escape, claude a single one
ok - fm-control: a harness with no verified control mechanics is refused, not guessed at
ok - fm-control-lib: a recorded harness resolves to its verified adapter without guessing
ok - fm-control-lib: only a runtime's own recorded session has a relaunch resume form
ok - fm-control: prefixed recorded harnesses reach interrupt and exit mechanics
ok - fm-control-lib: the backend key matrix matches each adapter's real send-key surface
ok - fm-control-lib: adapter capability is per task kind, not per adapter alone
ok - fm-control interrupt: a backend that cannot deliver the harness's key refuses instead of sending another
ok - fm-control: a backend that cannot prove an agent stopped refuses exit and relaunch
ok - fm-control-lib: stop-proving verbs are gated on the backends that really classify agent state
ok - fm-control: a legacy window label is refused and the exact task id is named
ok - fm-control: an explicit backend endpoint is never a control target
ok - fm-control: an unrecorded task id is refused
ok - fm-control: a record whose endpoint identity names another task is refused
ok - fm-control: a remotely placed secondmate is refused by placement, not by a metadata complaint
ok - fm-control: interrupt and exit lock before task-state resolution
ok - fm-control: the verb list is closed - no raw keys, arbitrary text, or clear verb
ok - fm-control: resume is refused with the determinism reason and the alternative
ok - fm-control: profile and note flags belong to relaunch only
ok - fm-control exit: an already-stopped agent is idempotent success with no bytes sent
ok - fm-control exit: an unprovable tmux endpoint refuses instead of claiming the agent stopped
ok - fm-control interrupt: refuses when no agent is running rather than keying a shell
ok - fm-control exit: an endpoint whose process cannot be attributed refuses
ok - fm-control exit: a busy agent receives interrupt delivery before the exit command
ok - fm-control exit: an idle agent goes straight to its exit command
ok - fm-control exit: retiring a codex incarnation drops busy_gen with the sidecar
ok - fm-control exit: an already-open background-task picker is not typed into
ok - fm-control exit: the Enter that opens the background-task picker is not followed by a confirming Enter
ok - fm-control exit: a picker that renders after the submit returned is named when the stop wait times out
ok - fm-control interrupt: unconfirmed delivery preserves observed busy state
ok - fm-control interrupt: muse confirms cancellation from its session log
ok - fm-control interrupt: postconditions are revalidated after acknowledgement polling
ok - fm-control exit: an interrupt-stopped agent satisfies the gone-state postcondition
ok - fm-control exit: a stubborn agent reports delivered input and an unconfirmed exit
ok - fm-control interrupt: grok reports delivery without claiming cancellation
ok - fm-control interrupt: grok's idle footer does not confirm cancellation
ok - fm-control: a lifecycle command to a secondmate is unmarked and opens no reply expectation
ok - fm-control's arrival leaves fm-send's from-firstmate marking untouched
FM_TEST_END 2026-10-07T07:53:40Z tests/fm-control.test.sh exit=0 duration_ms=77991 gate_skip=false
FM_TEST_SUMMARY total=4 failed=0 skipped_gate=0 duration_ms=208965
FM_TEST_SUMMARY_FAMILY family=backend-dispatch count=1 duration_ms=77991 failed=0
FM_TEST_SUMMARY_FAMILY family=pr-forge count=1 duration_ms=14148 failed=0
FM_TEST_SUMMARY_FAMILY family=standalone count=2 duration_ms=198377 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-spawn-pool-base-freshen.test.sh duration_ms=116537
FM_TEST_SLOWEST rank=2 script=tests/fm-pending-reply.test.sh duration_ms=81840
FM_TEST_SLOWEST rank=3 script=tests/fm-control.test.sh duration_ms=77991
FM_TEST_SLOWEST rank=4 script=tests/fm-review-diff.test.sh duration_ms=14148
- Outcome: ⚠️ 1 warning across 2 runs (35m45s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - medium risk

✅ No issues found.

⚠️ **Test** - 1 warning
  • ⚠️ The Test agent did not finish within its invocation budget. Reported: agent run tests timed out after 30m0s: agent last produced output 4s ago (332 observed); agent reported: pi exited: exit status 143. This is a budget or provider-slowness cut, not a code failure. Re-running the same request costs another full budget, so no further attempt is made automatically. If this repository's targeted tests or evidence gathering routinely approach the default 30m0s, raise test_agent_timeout in global config. Respond with fix to spend another budget: a repair turn runs only for selected findings other than this budget cut, then validation re-runs. Or abort and retry after raising the budget.
  • 🚨 Approval is refused: the run worktree at ~/.no-mistakes/worktrees/450411b3e67c/01M4AK0KYBPP2H30CZ1HR0HKJ4 holds work no Test turn validated, and the steps after Test would commit and publish it. It holds uncommitted changes to .test-phase/ (inspect with git -C ~/.no-mistakes/worktrees/450411b3e67c/01M4AK0KYBPP2H30CZ1HR0HKJ4 status and git -C ~/.no-mistakes/worktrees/450411b3e67c/01M4AK0KYBPP2H30CZ1HR0HKJ4 diff). Respond with fix to validate it, or abort.

🔧 No changes applied.
1 warning still open:

  • ⚠️ live validation verdict: inconclusive (0 of 5 scenarios were driven live against the product); untested: Review a task based on a named branch and see only the task’s changes, Review a Bitbucket PR using its current head, with explicit fallback warnings when unavailable, Start a task from its named base branch and refuse missing or contradictory base selections, Exit Claude without sending a confirming Enter into its background-task exit picker, Resolve a missed secondmate reply, then reopen the decision when the same failure recurs
  • Live validation: ⚠️ inconclusive - 0 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Review a task based on a named branch and see only the task’s changes ⏸️ untested no The prior payload did not establish a live result. tests/fm-review-diff.test.sh passed the recorded-base-branch case and stale-head and corrupt-branch checks, but executed the review CLI against dis…
Review a Bitbucket PR using its current head, with explicit fallback warnings when unavailable ⏸️ untested no The prior payload did not establish a live result. All three retained Bitbucket cases in tests/fm-review-diff.test.sh passed using a stubbed Bitbucket response; no live forge was contacted. Live Bit…
Start a task from its named base branch and refuse missing or contradictory base selections ⏸️ untested no The prior payload did not establish a live result. Named-base, brief-mismatch, and missing/local-only-base cases in tests/fm-spawn-pool-base-freshen.test.sh passed using disposable Git state and a f…
Exit Claude without sending a confirming Enter into its background-task exit picker ⏸️ untested no The prior payload did not establish a live result. Already-open, submit-opened, and delayed-render picker cases in tests/fm-control.test.sh passed against modeled terminal transitions. A real Claude…
Resolve a missed secondmate reply, then reopen the decision when the same failure recurs ⏸️ untested no The prior payload did not establish a live result. The same-kind escalation after operator close case in tests/fm-pending-reply.test.sh passed, including suppression of duplicate escalation while op…
  • Inspected git log -1 --format=&#39;%H%n%P%n%s&#39;: HEAD is 9afdf3f56016db613a7df70bd7f4a1149b470199, with the required fork-base and upstream-tip parents.
  • Removed untracked .test-phase/, adjusting only its read-only fixture-directory permissions as needed.
  • Ran bash -n separately on bin/fm-backend.sh, bin/fm-spawn.sh, bin/fm-test-run.sh, tests/fm-remote-reply.test.sh, and tests/fm-review-diff.test.sh.
  • Ran bin/fm-test-run.sh --check-coverage.
  • Ran FM_TEST_EVIDENCE=1 bin/fm-test-run.sh tests/fm-review-diff.test.sh tests/fm-spawn-pool-base-freshen.test.sh tests/fm-control.test.sh tests/fm-pending-reply.test.sh with disposable temporary fixtures and inherited fleet-path overrides unset.
  • Verified temporary-directory removal and clean status using git status --short --untracked-files=all, git diff --exit-code, and git diff --cached --exit-code.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

tiago-peixoto and others added 14 commits October 5, 2026 07:26
…t-in (kunchenguid#5354)

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* feat(bin): record each task's no-mistakes pipeline spend at cleanup

no-mistakes keeps every pipeline agent invocation's token usage only in its
local agent_invocations records, and cold pipeline agents (review, test,
document) leave no session log, so a task's review-loop cost never reached
Firstmate's records and could not be attributed after cleanup.

bin/fm-pipeline-spend.sh attributes a task's runs by the repository
no-mistakes resolves for the task copy, the task branch, and the branch's
creation (a relaunch mints a new spawn_gen while the branch keeps
validating), then sums every invocation, failed and cancelled included.
Token fields use no-mistakes' per-round deltas so resumed review rounds are
not counted twice, and unrecorded values stay unknown rather than zero.
show prints the record read-only; record appends it once per task
incarnation to data/pipeline-spend.jsonl. Teardown records it for every
ship task it cleans up, before deleting the branch and the task record.

fm_nm_state_db becomes the one owner of where no-mistakes' state database
lives, shared with the capped run-inventory reader.

* no-mistakes(review): Drop spend show command and timeout override

* style(bin): rewrap fm-pipeline-spend header comment

* Make pipeline spend recording opt-in

* no-mistakes(document): Note pipeline-spend opt-in gate in teardown comments

* no-mistakes(ci): I fixed all three Greptile findings following your decision. Both test suites pass: tests/fm-teardown.test.sh (105 ok, exit 0) and tests/fm-pipeline-spend.test.sh (8 ok). I didn't run either new test against the old code, so the claim that they fail before the fix is from reading the code, not a run. Nothing was committed or pushed. **ci-1 / ci-2 (bin/fm-teardown.sh):** The rule is that every owned ship task in a home that has opted in gets a durable spend record before its task record is deleted. The `[ -d "$WT" ]` check skipped the recorder when the worktree folder was already gone, so no record was written. I removed that check and kept the other conditions: the task must be a ship task, teardown must own its worktree, and `config/pipeline-spend` must exist. `bin/fm-pipeline-spend.sh` already writes an unavailable-source line when the worktree is missing. The new test `test_teardown_records_unavailable_spend_for_a_gone_worktree` in tests/fm-teardown.test.sh sets up an owned ship task whose worktree is missing, with recording turned on. It checks that teardown succeeds, writes a line with `source=="unavailable"`, `total==null` and a reason saying the task copy is gone, and still removes the task record. Under the old check no ledger would have been written. **ci-3 (bin/fm-nm-run-lib.sh):** When `NM_HOME` and `HOME` were both unset, `fm_nm_state_db` fell back to `${HOME:-}/.no-mistakes`, which resolves to `/.no-mistakes`. It now uses `root=~/.no-mistakes`. Bash expands `~` to the account's home directory even when `HOME` is unset, which matches the previous Python `Path.home()` lookup. The new test `test_state_db_without_nm_home_or_home_uses_the_account_home` in tests/fm-pipeline-spend.test.sh runs the function with both variables unset and compares the result to the home directory from the system's password database. The old code would have returned `/.no-mistakes/state.sqlite`. shellcheck reports nothing new on the changed files. I added one `disable=SC2016` comment in the new test, where the single-quoted `$1` is meant to expand in the child shell
* feat(bin): start tasks from a named base branch

Spawns always reset a task's pooled copy to origin's default branch,
so work that belongs on a feature, integration, or release branch
started from the wrong code and opened its PR against the default.

fm-brief.sh --base-branch records the base in the brief, fm-spawn.sh
resets the copy to origin/<base> and records base_branch= in task
meta, and the worker targets its PR at that branch. Review diffs,
the cleanup content check, and scout promotion read the recorded
base. local-only and Gerrit deliveries refuse a named base.

* no-mistakes(review): Anchor brief base-branch parse on scaffold Setup line

* no-mistakes(review): Require explicit --base-branch on spawn, matching brief lines

* no-mistakes(document): Note base-branch reset in fm-spawn freshness header

* no-mistakes(ci): I fixed all four Greptile findings the way you specified. The affected suites pass locally: fm-brief, fm-dod-lib, fm-spawn-pool-base-freshen, fm-review-diff, fm-teardown and fm-task-delivery. shellcheck on the changed files reports only info-level notes, no warnings or errors. I didn't revert the code to watch each new test fail before the fix. - **ci-1 (brief text blocked or redirected spawns):** In bin/fm-dod-lib.sh, `fm_brief_base_branches` now only accepts a `Base branch: <value>` line that comes directly after the base-variant Setup sentence fm-brief.sh writes. Any other `Base branch:` line is ignored. `--base-branch` is still the only authority, and the existing checks in fm-spawn.sh are unchanged. With the flag, the spawn is refused unless a matching pair exists and no pair names a different branch. Without the flag, it is refused only if such a pair exists. The header comment in fm-spawn.sh is updated to match. - The `brief_with_base` test helper now writes the real two-line pair. - The decoy refusal case now uses a full decoy pair in the captain's intent. - New test `test_prose_base_branch_line_is_ignored`: a brief with no base whose intent quotes `Base branch: release/1.2` spawns normally with no `base_branch=` recorded. A spawn with `--base-branch feature/hub` whose intent names release/1.2 starts from `origin/feature/hub` and records it. - **ci-2 (scout base not checked against the forge):** fm-spawn.sh now looks up the project's registered forge with `bin/fm-project-mode.sh --forge`, like the ship path does, before `fm_base_branch_valid` runs. Previously it passed `none`. New test `test_scout_base_branch_refused_on_gerrit_forge`: a scout with a base on a forge=gerrit project is refused at spawn and no meta file is written. - **ci-3 (base not shell-quoted in worker commands):** `fm_dod_block` now quotes the base with `printf %q` in the `--base` and `--base-branch` arguments, and the branch name in the prose stays readable. The test in fm-brief.test.sh uses the base `release/$HOTFIX` and checks that both direct-PR and no-mistakes briefs render `release/\$HOTFIX` in the commands. - **ci-4 (ambiguous PR sentence):** The direct-PR sentence now reads "open a PR with `gh-axi` that is ready for review, not a draft, against the base branch `X` (`--base X`), not the repository default." A brief with no base renders exactly as before. The existing assertion in fm-brief.test.sh is updated

* no-mistakes(ci): Both real failures on PR 6442 (run 37064756768) come from CI tooling and a Pi version bump, not from the base-branch feature. Fixes for both already exist on upstream main, so I applied those two commits to the working tree with `git cherry-pick --no-commit`. Nothing is committed or pushed, and no feature file changed. - **Lint 2:** ShellCheck ran out of memory on `bin/fm-teardown.sh` (rss about 8 GB, exit 251, "shellcheck: out of memory"). The rule that broke is that the lint gate must finish on any valid shell root without exceeding the runner's memory ceiling. Upstream commit d719ef3 (kunchenguid#6443) fixes this in `bin/fm-lint.sh`: a root that hits the memory ceiling is retried without `--external-sources`. It also updates `tests/fm-lint.test.sh` and `docs/fm-test-portable-shards.md`. - **Behavior portable serial 6:** `tests/fm-calm-pi-extension.test.sh` failed with "grep disappeared from /export calm.html HTML while calm mode was on". Pi 1.0.1 renamed its HTML export renderer lookup, and upstream commit fede619 (kunchenguid#6530) updates the test to accept the new name. This change is test-only. Both commits applied cleanly. The four touched files are `bin/fm-lint.sh`, `docs/fm-test-portable-shards.md`, `tests/fm-lint.test.sh` and `tests/fm-calm-pi-extension.test.sh`. None of them is in the feature diff, so the feature contract is unchanged. **Local checks:** - `tests/fm-lint.test.sh` passes. - `bin/fm-lint.sh` on `bin/fm-teardown.sh`, `bin/fm-spawn.sh` and `bin/fm-dod-lib.sh` exits 0 with full ShellCheck analysis. - `tests/fm-calm-pi-extension.test.sh` exits 0 with 7 ok and 0 not ok. One case skipped because the Pi package isn't installed locally, so the Pi 1.0.1 export path that failed in CI couldn't be reproduced here. CI needs to confirm it. - I didn't reproduce the 8 GB OOM locally. The PR has to be updated through the pipeline with a normal push: no force push, no replacement PR, no merge
…ole (kunchenguid#6647)

* fix(bin): point project workers at the Firstmate skill file

The Skill tool cannot resolve a Firstmate skill from another project's worktree, so the launch role names the readable skill file instead.

* no-mistakes(review): fix(bin): name Firstmate skill file as fallback only
…ne (kunchenguid#6655)

* feat(bin): retire a contribution whose forge object is permanently gone

Add fm-contributions.sh retire <task> <url> <captain|fleet> <reason>,
which records actor, reason and time on the saved record and removes
that task/url pair from known, poll rotation and coverage even while a
backlog link remains. Repeating a retire keeps the first provenance;
an unrecorded pair, unknown actor, empty reason or unacknowledged
pending signal is refused.

* no-mistakes(review): Keep retirement per task when settling final owners

* no-mistakes(ci): Made the three Greptile fixes the captain chose. ci-1, retirement needs the captain's word: `retire` now accepts only the actor `captain`. Any other actor is refused, including `fleet`. I removed `fleet` from the usage line, the script header (`bin/fm-contributions.sh`), the check in `bin/fm-contributions.sh` and the retired-record validation in `bin/fm-contributions.jq`, which now requires `.actor == "captain"`. The script header now says a retirement always records the captain's word, because the script cannot verify who runs it and the authority to retire is the captain's. The Bearings skill line in `.agents/skills/bearings/SKILL.md` now says `retire` records the captain's word. In the tests, a `fleet` retire is now a refusal case and must leave the record unretired. Every other retire call in the tests uses `captain`. The PR description has not been changed yet: the ruling asks for the same captain's-word statement there, and that belongs to the PR phase. ci-2, blank reasons: the reason check is now `[ -n "${5//[[:space:]]/}" ]`, so a reason made only of spaces or tabs is refused. I added a refusal test that passes a space-tab-space reason. The header also lists a blank reason among the refusals. ci-3, the gone-object test: the test forge wrapper has a new `not-found` fault that returns HTTP 404 for every `api repos/o/r/...` read. The happy-path test `test_retire_ends_observation_of_a_gone_contribution` now uses it in place of the 502 `down` fault, so it models a permanently gone repository rather than an outage. Verification: `bash tests/fm-contributions.test.sh` passed with exit 0. The last lines of its output include the three retire tests passing and no failures. `shellcheck` flagged only SC2034 (`FM_WAKE_QUEUE` unused) and SC1091 (an unfollowed source of `bin/fm-path-lib.sh`), neither on a line this round changed
* test: give the remote-reply whole-log recapture a longer wait

* no-mistakes(review): Extend both recapture waits and simplify retry handling
…nguid#6654)

* fix(bin): reopen a pending-reply escalation after its resolve

A retry of an open decision still appends nothing. A later same-kind escalation after a resolved line for that key appends again, so the next loss is visible.

* no-mistakes(ci): Both Greptile findings are fixed in the working tree; the four directly affected test suites pass, and a wider run of related suites had not finished when I returned this result. Invariant: a line already in the status file is recorded, and only a caller with its own evidence of a new episode may append it again. It must hold for every caller of status_event_recorded: the continuity break and the other call in bin/fm-procevent-remote-reply.sh, fm_parent_channel_append_once, and the pending-reply escalation. Changes: - bin/fm-classify-lib.sh: status_event_recorded is back to the idempotent retry check. It returns at the first matching line, and a later resolved line no longer makes that line look new. This fixes ci-1 for every caller and restores the early exit for ci-2, with no separate scan optimization. - bin/fm-pending-reply-lib.sh: _fm_pending_reply_maybe_escalate_locked now owns the reopen. It appends a new blocked line when the record already has an escalated_epoch (it escalated before and was reset) and the keyed decision is no longer open. Otherwise it uses status_event_recorded, so a retry or a resend that leaves the decision open appends nothing. - Comments on status_event_recorded and the pending-reply header describe this scope. Tests: - tests/fm-pending-reply.test.sh: the regression for escalate, operator resolve, second escalate is unchanged and passes. - tests/fm-remote-reply.test.sh: new regression drives the real reader. After an operator resolve of remote-reply-continuity-ios, a repeated handle and ingest of the same break appends nothing and leaves the decision closed. I placed it beside the existing continuity-break test, not in the pending-reply file, because only that file has the reader harness. - tests/fm-classify-corr-token.test.sh and tests/fm-classify-decision-key.test.sh: the cases that expected a blocked line to append again after a resolve now assert it stays recorded. Verification: - fm-classify-decision-key, fm-classify-corr-token, fm-pending-reply and fm-remote-reply test suites all exit 0 with the fix. - With the two bin files reverted to the commit under review, the new continuity regression and the updated parent-publisher case both fail, so they reproduce ci-1. - shellcheck -x on the changed files reports nothing. - Not finished: a background run of every other suite that mentions the parent channel or pending replies had produced no output, pass or fail, when I returned. One known gap: the escalation writes escalated_epoch before phase. If the process dies between those two writes and an operator resolves the key before the next tick, that tick appends a blocked line for the same episode. I left it, because closing it needs new state. Issue 4755 and other escalation behavior are untouched. Nothing is committed
…ker (kunchenguid#6666)

* fix: refuse a confirming Enter on the Claude exit picker

The background-task picker still classifies as pending, so a retried Enter confirms "Exit and stop tasks".
Stop after the Enter that opened it, and raise the existing stale wake with the dialog name.

A non-paused secondmate still skips that wake, so a secondmate parked on the picker still looks idle.
Model-downgrade confirmation, MCP approval, and any other Claude exit confirmation are not covered, because there is no recorded screen for them.

* no-mistakes(review): fix: anchor exit picker match and wake second mates

* fix: refuse a typed submit while the Claude exit picker is open

A pane that already shows the picker must not receive the message or a confirming Enter.
The match requires the recorded heading and footer lines, and it is skipped when no dialog name is being recorded.

* no-mistakes(review): restore watcher to main behaviour, drop dialog wake

* no-mistakes(test): test: align Herdr picker fixtures with the preflight read

* no-mistakes(document): document exit refusal on a recognised dialog

* fix: remove the dialog file when exit runs in a subshell

do_exit is invoked from a command substitution, so the parent cleanup never saw the path it is supposed to delete.

* no-mistakes(review): fix: set the dialog file path after the control lock

* no-mistakes(document): document why the dialog file path follows the lock

* no-mistakes(ci): The exit cleanup in bin/fm-control.sh now deletes the dialog file first and releases the control lock second. The changes are in the worktree and are not committed. Invariant: only the process that holds the control lock for a task may write or delete that task's dialog file ($STATE/<id>.composer-dialog, the file that carries the picker name from the composer read to the Enter checks). The old cleanup released the lock and then deleted the file, so the delete ran outside the lock. Places where this invariant must hold: control_cleanup is the only place that deletes the file, and the single assignment after CONTROL_LOCK_HELD=1 is the only place that names it. do_exit only truncates the file, and it runs while the parent holds the lock. A process that loses the lock never sets the path, so it deletes nothing (earlier fix, unchanged). No sibling site needed a change. Fix: in control_cleanup, the rm of the dialog file moved above the fm_lock_release call, with a two-line comment that gives the reason. No dialog recognition or supervision code changed. Test: added test_exit_removes_the_dialog_file_before_releasing_the_lock in tests/fm-control-relaunch.test.sh. It runs one `fm-control exit` with a recording rm on PATH. The lock release removes paths at or under the control lock with rm, so the recording rm writes whether the dialog file exists at that moment. The test asserts the last record is "absent". It starts no second command. Verification, all run locally: - Before the fix, the new test failed with "the dialog file must be gone when the control lock is released, got: present". - After the fix, tests/fm-control-relaunch.test.sh exits 0 with 75 ok lines and no "not ok" line, including the new test and the existing test that the file is gone after exit and after relaunch. - tests/fm-control.test.sh exits 0 with 44 ok lines. - shellcheck -x on both changed files exits 0 with no output. One thing I noticed and did not change: tests/fm-control-relaunch.test.sh prints 615 "rm: cannot remove ... Permission denied" lines during its temp-directory teardown. The count is the same with my change reverted, so this change does not cause it, and the suite still exits 0. I could not read the Greptile check log (the check run was not found), so I worked from the finding text and the user instructions
…henguid#6028)

* fix: support stock macOS Bash in timeout watchdog

* no-mistakes(ci): Fixed ci-1 in tests/fm-timeout-lib.test.sh: both startup-owner failure paths now kill the bounded command group before killing the watchdog. Fault injection reproduced both leaks before the fix and confirmed cleanup afterward. Bash 3.2 suite, syntax check, ShellCheck, and diff checks passed; GNU timeout coverage skipped because the binary is unavailable. Production code unchanged
…nchenguid#6699)

Wrap the cd command in a subshell to comply with the cd-guard policy
that blocks persistent top-level directory changes in the primary
firstmate checkout. The subshell form (cd projects/<name> && ...) is
accepted by the policy as documented in issue kunchenguid#6502.

Fixes kunchenguid#6502
* Add daily startup growth check

* no-mistakes(review): watch printed startup memory, retain growth baselines, drop knobs

* no-mistakes(review): delegate budget verdict, report before publish, pin shim home

* no-mistakes(review): baseline first sightings silently, drop mtime, spare secondmates

* no-mistakes(review): cap the wake line, validate budget verdict fields

* no-mistakes(review): guard record schema, check appends, tighten assertions

* no-mistakes(review): drop arbitrary tracked library, make re-arm idempotent

* no-mistakes(review): correct tracked-set wording, restore shim guards, register artifacts

* no-mistakes(review): exit on signal instead of publishing partial record

* no-mistakes(document): correct startup-growth record removal cost in state registry

* no-mistakes(test): sweep orphaned empty startup-growth temp records on due evaluation

* no-mistakes(document): note watcher need for armed startup growth check

* no-mistakes(ci): Diagnosed "Behavior portable serial 5": the shard failed on tests/fm-contributions.test.sh in test_arm_plumbs_a_configured_budget_into_the_check_shim (inherited mode) with "generated check did not attempt a read" and a missing forge/calls file. Root cause: bin/fm-contributions.sh:379-382 computes DEADLINE=$(date +%s)+BUDGET and then gates each read on DEADLINE-$(date +%s) >= OBSERVATION_RESERVE (= min(BUDGET,15)). At the one-second budget this test uses, that gate demands zero elapsed whole seconds, so a wall-clock second boundary crossing between the two date calls makes the poll loop break before any forge read. The test calls wrap_forge (which installs a fake date reading $FORGE/clock only when that file exists) but never seeded forge/clock, so it ran against the real clock. The repo already documents this exact hazard at tests/fm-contributions.test.sh:675-677 and freezes the clock in every other budget-constrained test: test_budget_exhaustion_keeps_prior_record (both modes), test_genuine_failure_near_deadline_is_unavailable, test_reservation_defers_later_url_when_fifteen_seconds_do_not_remain, test_slow_read_deadline_kill_is_budget_refusal, test_budget_is_cut_down_to_the_watcher_check_bound. test_arm_plumbs was the only omission; its single loop body covers both the configured and inherited modes. The remaining wrap_forge users run the default 20-second budget (5 seconds of slack above the reserve) and are not reachable by this race. Fix (smallest, matching the sibling idiom): added the clock freeze `/bin/date +%s > "$home/forge/clock"` plus a two-line comment inside that test's loop, before the poll. Three added lines in tests/fm-contributions.test.sh; no production code and no other file touched. Verification: reproduced the exact CI failure message by running a scratch copy whose fake date advances one second per read (worst case of the unfrozen clock) - missing forge/calls, "not ok - generated check did not attempt a read". With the fix, the single test passes 10/10 standalone and the full tests/fm-contributions.test.sh passes all 46 assertions with exit 0. tests/fm-startup-growth-check.test.sh still exits 0. bash -n and shellcheck -S warning are clean on the edited file; scratch repro files were removed, leaving only the intended three-line change. Note for the author: this flake is not caused by this branch - tests/fm-contributions.test.sh is untouched by the PR and the race is pre-existing, surfacing only on a loaded runner (the shard's neighbouring assertions show 80+ second gaps). It was fixed because it is genuine nondeterminism with an established in-repo remedy and the check cannot otherwise go green. No work was done on the unselected greptile findings (ci-2..ci-5) or the cancelled "Behavior portable serial 2" check (ci-6)

* no-mistakes(ci): Diagnosed "Behavior portable serial 6": the omitted log section (fetched with gh api) shows the shard failed on tests/fm-calm-pi-extension.test.sh in test_hidden_block_geometry_e2e with "Pi Calm hidden-block geometry E2E did not complete the /reload viewport transition" (exit=1, duration_ms=23437; the family line pure-contract-unit count=3 duration_ms=51193 failed=1 matches devin-harness 3203 + calm-pi 23437 + task-delivery 24553). Root cause: tests/fm-calm-pi-extension.test.sh:2831 wait_for_geometry_transition detected the /reload by SAMPLING a single intermediate frame - it polled tmux capture-pane for Pi's transient "Reloading keybindings, extensions, skills, prompts, themes, and context files..." box and only accepted the rebuilt transcript afterwards (elif gated on saw_transient). Pi shows that box only while the reload runs and replaces it via dismissReloadBox when done, so on a loaded runner the box can live entirely between two polls; saw_transient then stays 0 forever and the 600-attempt wait times out although the reload fully succeeded. Reproduced locally under CPU oversubscription with CI's Pi 1.0.4: 1 failure in 8 runs, diagnostics proving the mechanism ("DIAG: transition timeout saw_transient=0 attempt=600") while the captured viewport already showed the completed reload status row and the intact collapsed transcript. Not a Pi-version regression: 1.0.4's handleReloadCommand and both banner strings are identical to the installed 0.85.1, and shard 6 of the previous pipeline run (job 112490351105) passed this same script on Pi 1.0.4. Invariant violated: a TUI E2E assertion must key off a durable state the UI retains, never off one intermediate frame polling may miss. Sibling enumeration: the helper was defined once and called once; grep for transient/saw_/two-stage waits across tests/, bin/ and .pi/ found no other one-shot-frame wait - every other wait in the file keys off durable text presence/absence or off continuously repeating animation frames (the working-ship motion checks) where sampling eventually connects; no doc or other test referenced the helper. Fix (smallest, one helper): renamed it wait_for_geometry_reload and made it wait for the durable status row Pi appends once the reload completed and the chat was rebuilt ("Reloaded keybindings, extensions, skills, prompts, themes, and context files", a persistent chat child via showStatus, present in both 0.85.1 and 1.0.4) together with CALM_GEOMETRY_FINAL. That marker is absent before the reload and never appears if the reload throws, so the check is strictly stronger than before (a failed reload now fails the test instead of passing on transient-then-final). One file changed: tests/fm-calm-pi-extension.test.sh, 2 hunks, no production code. Verification: deterministic before/after - with polls spaced 0.4s apart (the loaded-runner worst case) the pre-fix helper fails with exactly the CI message and the post-fix helper passes; 12/12 consecutive passes of the isolated test under load with Pi 1.0.4; full file 15/15 ok exit 0 under LANG=C.utf8 (matching the runner locale); bash -n, shellcheck -S warning and bin/fm-lint.sh clean; scratch repro files and the temp Pi 1.0.4 install removed, git status shows only the one intended modification. Notes: (1) the flake is not caused by this branch - that test file is untouched by the PR and the race is pre-existing; fixed because it is genuine nondeterminism and the check cannot otherwise go green. (2) Under far harsher load than CI applies a separate unrelated flake appeared in the same test: Pi's "Warning: tmux extended-keys is off..." row occasionally lands between the collapsed [skill] ahoy row and the final response, making assert_geometry_gap see 5 rows instead of 2. Different cause, never seen in CI, and its plausible remedies would change key delivery for the other TUI tests in the file, so it was left alone rather than growing this fix

---------

Co-authored-by: Quartermaster <quartermaster@users.noreply.github.com>
kunchenguid#6708)

* fix: reopen a remote-reply continuity break after repair

A later break for the same route and reason was swallowed after the
operator resolved the first one, because the status line matched for
the life of the log. The continuity ingest now appends again when the
cursor has moved or retirement has reset that episode, and an unchanged
re-read still appends nothing.

status_event_recorded is unchanged. Its other callers are the
pending-reply escalation, which already decides its own episode, the
parent-channel note append, and the remote document transfer note.

* no-mistakes(review): seed continuity episode for already recorded break line

* no-mistakes(document): document when a remote-reply continuity break reopens

* no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone

* fix: record a continuity break's reader position on its status line

A later break at another cursor is then a different line, so the existing
duplicate check appends it and reopens the decision. An unchanged re-read
builds the same line and appends nothing.

* fix: reopen a continuity break after an identical restore

A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct.

* no-mistakes(review): remove continuity match for full-prefix line without retirement count

* no-mistakes(document): clarify what a continuity break status line records

* fix: remove the reply cursor before recording retirement

A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.
…is retired (kunchenguid#6733)

* fix(control): drop busy_gen when an incarnation is retired

A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable.

* no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch

* no-mistakes(review): correct lock comment to name the deadlock

* no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it
…#6484)

* test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe

Closes kunchenguid#6456

* no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes

---------

Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
# Conflicts:
#	bin/fm-backend.sh
#	bin/fm-spawn.sh
#	bin/fm-test-run.sh
#	tests/fm-remote-reply.test.sh
#	tests/fm-review-diff.test.sh
@cloud-practitioner cloud-practitioner changed the title feat: merge upstream firstmate into the fork Merge upstream firstmate into the fork Oct 7, 2026
@cloud-practitioner
cloud-practitioner merged commit 5bff92a into main Oct 7, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants