Conversation
…nguid#5294) * fix(bin): map passed-with-override to done instead of unknown no-mistakes' axi status emits outcome: passed-with-override for a run that finished with an explicitly approved Test or CI exception. Both bin/fm-crew-state.sh's outcome resolver and bin/fm-teardown.sh's pre-teardown terminal-run check only matched the literal passed and checks-passed tokens, so this outcome fell through to unknown/parked and a finished worker awaiting merge kept getting re-alerted as stale, while an abort race during teardown could also leave a finished run misreported as still parked. Map passed-with-override to the same done/terminal handling as a clean passed in both places. * fix(document): Replace stale outcome mapping with authoritative pointer * fix(ci): Fixed a pre-existing mock-clock race in tests/fm-contributions.test.sh by advancing time only during the serial issue read. Reproduced the exact CI failure before fixing it. Forced-race replay, all 38 contribution scenarios, scoped ShellCheck, Bash syntax, and diff checks pass. Only the test fixture changed; CI rerun remains with the outer executor
* fix: close landed workers from supervision in both postures and at return During the 2026-09-22 away window every exemption worker whose pull request had merged was left sitting for nine hours. The supervision branch received the stale wake, the merge-landed check, and the hourly inactive-outcome row for each of them, ran the recovery playbook, found nothing to recover, and reported "no further action". The branch prompt granted ordinary teardown of a confirmed-landed task without ever naming the moment or the command, and the playbook has no landed exit, so the stale path ended at "nothing to recover". The return brief then listed only blockers, decisions, and the latest five routine outcomes, so the landed workers stayed invisible after the captain came back. - bin/fm-branch-prompt.sh: name the merge-landed wake, and any later stale, inactive-outcome, or heartbeat row on a done task with a merged PR, as the moment to claim the lease and run bin/fm-teardown.sh with no flags; a refusal is reported, never forced or worked around. Add teardown to the handling tool list. - stuck-crewmate-recovery: a landed worker is not a recovery case; point at the ordinary teardown owner for each actor. - bin/fm-afk-return.sh: render a "Landed, cleanup due" section from durable records only (a live task record whose recorded PR carries the merge-notification marker), between could-not-fix and handled, without holding the gate; the afk skill's return step closes each listed task through ordinary teardown once the check clears. - tests: pin the prompt rule in fm-branch-supervision and the brief section in fm-afk-return through the real marker writer. * no-mistakes(document): Document landed-task cleanup ownership
* fix(bin): surface a green no-mistakes PR still in ci merge monitoring A green PR could sit unreported because neither the worker nor the supervisor could observe checks-green while the ci step kept monitoring for the merge. Supervisor read: fm_nm_select_run's capped-overview inventory reader looked the repository up by the task worktree path, but no-mistakes registers a repository once by its main clone path and resolves every linked worktree to it, so on every task copy of a busy repo the lookup matched no row and each read reported "complete same-branch run inventory unreadable". Key the lookup on the overview's own top-level `repo:` line, which every axi release emits as the resolved working_path. Even with a readable run, the ci-log classifier treated "base branch advanced ..., re-arming CI monitor timeout" as not-ready. The monitor logs a checks state only when it changes and a base advance does not clear readiness, so a green PR read as still validating for as long as main kept advancing. Stop treating that line as a marker, matching no-mistakes' own ci-log parser, and name the run's PR URL in the held-for-merge reading so the existing inactive-outcome path can act on it without a worker report. Worker contract: `axi status` never reports checks-passed while the ci step monitors for merge, so the definition of done no longer makes a status poll the wait for the next gate or outcome; the drive call's own return is the green signal, reattached with `no-mistakes axi run` after a bounded return. * no-mistakes(review): read the full ci log when checking checks-green * no-mistakes(review): correct stale ci log tail wording in docs * no-mistakes(document): Document checks-green supervisor fallback
* fix: derive Lavish polling server from its board session * no-mistakes(document): Document session-derived Lavish polling * no-mistakes(document): Correct Lavish routing verification claims
… vanish (kunchenguid#4900) * fix(bin): ignore vanished state scratch files on secondmate relaunch Relaunch refused when find(1) exited non-zero while listing a secondmate home's state directory. A live watcher can delete scratch files between readdir and processing, which is not evidence that child *.meta records are unreadable. Prove the directory is listable from its mode and keep the existing readable-meta loop as the child-record guarantee. Fixes kunchenguid#4765. * no-mistakes(review): Skip chmod-000 unlistable-state relaunch test when running as root
…d#4907) * fix(bin): treat home-owned status closes as already read Self-announced bookkeeping appends now record their exact byte ranges. Later drains and signal scans skip those ranges, so two distinct --resolve-key answers after an OPEN DECISIONS fold do not each wake the supervisor. Worker-authored lines outside that ledger still signal. * no-mistakes(review): Keep owned closes in unread status; lock ledger writes * no-mistakes(review): Drop fold-lag wake suppression so folded worker decisions still wake * no-mistakes(review): Require real owned growth before ledger marks status seen * no-mistakes(document): Clarify home-appends ledger scope versus UNREAD STATUS * no-mistakes(review): Restore fold-lag path, drop owned-range filters, fix test * no-mistakes(review): Align ledger docs and scope ledger to wake path only * no-mistakes(review): Restore stranded historical-annotation test comment to its function * no-mistakes(review): Retire the home-appends lock alongside its ledger * no-mistakes(document): Note ledger's lock-helper dependency in classify library * no-mistakes(review): Append-and-coalesce home-appends ledger; fix stamped-line assertions * no-mistakes(review): Drop redundant empty-span branch; make owned test pin ledger * no-mistakes(document): Document covers' ascending-order dependency on home-appends ledger * no-mistakes(document): Note owned-append skip in watcher signal-scan comment
…nguid#5350) * chore(bin): raise tasks-axi, quota-axi, and lavish-axi floors to latest Raise the minimum versions to tasks-axi 0.2.6, quota-axi 0.1.50, and lavish-axi 0.1.77, pin CI's tasks-axi install to 0.2.6, and move the floor-boundary test fixtures to the new versions. tasks-axi 0.2.6 makes a failed relation deliverable for a promised-final expecting pr-merged, so add the regression test: a bound work that ends failed reports its honest outcome text through fm-public-followup-emit.sh, consume marks the commitment ready, and deliver posts that text exactly once. Also make two hang-guard tests in fm-backlog-atomicity portable to hosts without coreutils timeout, and stop an installed herdr from leaking into the secondmate-liveness husk classifier test. * no-mistakes(review): drop out-of-scope bounded_run hang-guard helper from atomicity test * no-mistakes(review): pin quota-axi floor at 0.1.49 across fixtures * no-mistakes(document): Document failed public-followup delivery behavior * no-mistakes(ci): Updated quota-axi floor and all 0.1.49 fixtures to 0.1.51, corrected bootstrap boundaries to 0.1.51/0.1.52/0.1.50, and bumped the bearings lavish-axi stub to 0.1.77. Bearings, quota procevent, quota chooser, startup budget, and bootstrap floor coverage passed; the full bootstrap suite exceeded the 240-second local command limit after relevant checks passed. git diff --check passed
…rker copy (kunchenguid#4878) * fix(bin): refuse ship done: when the named head lives only in the worker copy A ship done: is not current-state done until that exact commit is reachable outside the disposable copy. The check tests the named head, not whether some branch moved. * fix(bin): gate CI-ready ship done: on named-head reachability, not handoff Keep no-mistakes' first done: as the pipeline handoff, apply the same shared check when registering a PR and when a secondmate publishes ledger-first, treat a recorded merged PR as landed after prune, and name the PR head instead of scanning free-text SHAs. * no-mistakes(review): Bind named-head gate to recorded PR and forge heads * no-mistakes(review): Gate direct-PR forge heads and keep pending ledger deliveries * no-mistakes(review): Align worker done wording, test mapping, pending-retry test * no-mistakes(test): Raise watcher test time limit to stop load flake * no-mistakes(document): Restore ledger-path fact and name named-head gate coverage * ci: re-attest named-head ship-done gate for a fresh serial-3 verdict * no-mistakes(review): Simplify local-only gate, gate keyed done lines, document recovery * no-mistakes(document): Name fm-crew-state among named-head gate callers
…all alarm (kunchenguid#5204) * fix(bin): ring a proven-idle secondmate before a wake-loop stall alarm A leftover foreign-queue row on an idle, alive, ring-safe mate is still drainable in that home. Ring once, reset the observation interval, and keep the parent alarm for unknown, busy, or still-frozen rows. * no-mistakes(review): Mark drain steer with from-firstmate fire-and-forget carrier
…unchenguid#5335) The re-arm recovery cases judged "the watcher stayed live instead of surfacing recovery" with fixed budgets below what a real stale-lock recovery costs on a contended host: the arm's default 10s confirmation deadline, a start helper that returned after about 4s whether or not the arm had confirmed its watcher, and an 80-poll exit wait. A changed-suite run beside other suites starves the recovery's many short-lived processes while this suite's sleeping poll loops keep their pace, so a watcher still surfacing its recovery read as one that stayed live (issue kunchenguid#3793). The original 0.25s window after confirmation was widened to 80 polls in kunchenguid#3837, which left the same race at a larger size. Following the CONTRIBUTING.md fixture-budget rule, the re-arm helper now gives the arm an explicit 30s confirmation budget and waits for its confirmation or exit within a ceiling that outlasts it, and every wait on a re-armed watcher uses one named iteration-counted ceiling that outlasts the same budget. A passing case returns as soon as the arm reports or exits, and a watcher that never surfaces its recovery still fails. A new case delays every mktemp and readlink the re-armed watcher runs after it publishes its beacon, so its first poll and exit take about 13s on any host. It fails with the reported symptom on the previous budgets and passes now. No bin/ change.
* fix(bin): let one TERM always stop the watcher on bash 5.2
Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped silently, consuming the signal. The
watcher's `trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and
keep polling while its stopper waited: the triage suite's reap waited
forever (CI jobs cancelled at 30 minutes), and the arm's signal path and
the away-mode daemon's shutdown wait for the watcher the same way.
Bash 5.3 fixed the parser; 5.2 is the stock bash on Ubuntu 24.04.
HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits on bash 3.2, 5.2, and 5.3. INT keeps
its trap because bash ignores a direct SIGINT while a child runs. The
check-spawn deferral window no longer contains a command substitution.
The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging the job, and a new regression test
proves TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.
* no-mistakes(document): Clarify watcher stop-signal documentation
…id#5374) * fix(bin): submit our own stuck doorbell instead of skipping every later ring * no-mistakes(review): Confirm and retry Enter once on stuck-doorbell submit * no-mistakes(document): Clarify doorbell retry and pending-composer documentation
* feat(bin): add the opt-in fleet activity ledger Homes that create config/fleet-ledger get an append-only JSONL file, state/fleet-ledger.jsonl, recording task.dispatched, task.status, task.merged, and task.cleaned_up so outside tools can follow a fleet. With the flag absent each producer does one file test and nothing else. docs/fleet-ledger.md owns the record contract and its documented limits. * no-mistakes(review): Record task.status text verbatim after the first colon * no-mistakes(document): Clarify fleet ledger status and setup documentation * no-mistakes(ci): Fixed a timing race in tests/fm-pi-branch-extension.test.sh: the replacement-wake test now waits for the prompt to start before releasing it. The focused test passed twice, and git diff --check passed
…nchenguid#5352) * fix(bin): format, validate, and surface public-followup deliverables brief pre-fills report_path=data/<work-id>/report.md and states the accepted format of every value it cannot know instead of a bare <value> placeholder. fm-public-followup-emit.sh refuses a deliverable tasks-axi would refuse, in both the direct and staged destinations, naming the key, value, and format. consume records the specific deliverable, outcome, or missing key behind a tasks-axi refusal, and each refusal wakes the owning home once through the existing relay poll. * no-mistakes(review): refuse emits missing a required deliverable in both destinations * no-mistakes(review): require promised deliverables and keep rejections recoverable * no-mistakes(review): mirror tasks-axi's canonical pull request URL rule * no-mistakes(review): keep a rejection wake whose line cannot be read * no-mistakes(review): key emit-time rules on the promise, not the outcome * no-mistakes(review): bound deliverable keys and values as tasks-axi does * no-mistakes(review): state rejection wakes as at-least-once and pin it * no-mistakes(review): enforce the promised contract tasks-axi holds at emit * no-mistakes(review): stop inferring a staged promise from its outcome * no-mistakes(document): Refresh public follow-up documentation * no-mistakes(ci): Fixed both CI flakes. Watcher cleanup is now installed before singleton acquisition, preventing timeout races from leaving stale locks while preserving recovery-failure evidence. Bearings render fixtures now publish a valid isolated Lavish session store and retire each listener after rendering, eliminating false unowned-source races. Verified with checkpoint stress, fm-watch-checkpoint, fm-watcher-lock, repeated fm-bearings-board-render runs, project lint, syntax checks, and git diff checks * Revert unrelated CI auto-fix edits to the watcher and bearings board test The CI step's automatic repair changed bin/fm-watch.sh and tests/fm-bearings-board-render.test.sh to chase two intermittent CI failures that also occur on main and are not part of this change. Restore both files so this branch carries only the public-followup deliverable fix. * no-mistakes(review): Refuse a repeated --deliverable key at emit argument parsing * no-mistakes(document): Clarify public-followup validation and rejection-wake documentation
* Add verified Devin CLI worker adapter * no-mistakes(review): Drop Devin resolver refusal and launch marker * no-mistakes(review): Verify devin in bootstrap, fold kind rule, update docs * no-mistakes(document): Document Devin sidecar, resume, and worker-only facts * no-mistakes(document): Document Devin interrupt, liveness anchor, composer signals * fix(control): never pair Devin interrupt presses on an idle agent A fast double Escape on an idle Devin opens its /revert picker, where Enter reverts file changes. fm-control now sends the second press only after the first renders Devin's 'esc again to interrupt' armed hint, never sooner than 0.5 s, closes a revert picker a mistimed press opened with one Escape, and refuses to type the exit command while that picker is open. An unarmed interrupt reports cancel=not-running and leaves the busy record untouched. * fix(devin): disable Claude hook import and commit attribution for workers The per-task Devin config now forces read_config_from.claude=false, so a worker no longer runs the user's or project's Claude Code hooks (including Herdr's Claude agent-state hook), and attribution=false, so Devin adds no Co-Authored-By trailer or Generated-with line to commits and PRs. * test(devin): extend live guard and record Herdr and revert-picker evidence The credentialed live guard now fails if an imported Claude Code hook runs, if the worker's commit carries Devin attribution, if an idle interrupt sends more than one press or opens the revert picker, or if an open picker lets exit through or is closed with a revert. The Devin reference, agent-control doc, and verification records carry the 2026-09-22 tmux and Herdr lab results, including the Herdr exit refusal. * no-mistakes(document): Correct Devin documentation links and lifecycle guidance --------- Co-authored-by: Denis Beliaev <battler73@yandex.ru>
…id#5322) fm-crew-state classifies the no-mistakes outcome 'passed-with-skips' as unknown, so a finished worker awaiting merge is re-alerted as stale. The same blind spot lets fm-teardown's pre-teardown terminal-run check refuse a legitimate abort race that lands on this outcome. Map passed-with-skips to done in crew-state resolution, keeping the skipped publication/CI verification visible in the detail rather than reporting a clean pass, and recognize it as terminal during teardown.
…nguid#5382) * fix: refuse missing backend adapter before source * no-mistakes(review): Gate backend precheck under stock Bash * no-mistakes(document): Clarify adapter precheck docs * no-mistakes(lint): Suppress intentional child Bash ShellCheck warning
…kunchenguid#5338) * fix(test): repair tmux liveness and calm follow-up loaded_off regressions Both self-tests fail on untouched main on a host whose coreutils are a multicall binary and whose Chrome has no pre-warmed profile, and each failure masks the other's file. tests/fm-tmux-agent-liveness.test.sh - the stand-in harness processes were symlinks to the host's `sleep`. A single-purpose `sleep` runs happily under another name, but a multicall coreutils binary (uutils or busybox) resolves its applet from argv[0]: `claude-link -> sleep` invoked under the harness name runs the wrong applet and exits immediately, so no foreground process exists and every positive case reads not-alive ("last verdict for liveness:agent was missing (expected alive); title=sh comms=[sh ]"). Build a dedicated spinner as the stand-in target, exactly the way the version-string case already builds its executable, and require the fallback target to demonstrably survive the rename before using it. Every assertion is untouched; the stand-in identity signal is unchanged (the kernel still records the symlink name as the executable identity). tests/fm-calm-pi-extension.test.sh - render_export_dom pinned a brand-new `--user-data-dir` per attempt. On Google Chrome for Testing 151.0.7922.34 that pristine profile makes Chrome's first-run initialization never complete: the browser and its renderers start, but --dump-dom never returns, so all three bounded attempts end exit=0 timed_out=yes bytes=0 and the DOM assertions never run ("could not render calm-mode HTML export DOM"). Chrome's own profile creation under a fresh HOME renders the same document in about a second, so the helper now gives Chrome a private per-attempt HOME instead of the explicit profile flag. Each attempt still gets an isolated profile, and every DOM assertion is unchanged. Root-cause evidence: a pristine --user-data-dir with `--headless=new --dump-dom` had not returned after 150s, while the same command with an empty HOME and no --user-data-dir returned the full DOM in ~1s, and reusing an already-populated profile also returned it in ~1s. The render failure masked the rest of the file: with it repaired, the Pi follow-up loaded_off case passes unmodified against an installed @earendil-works/pi-coding-agent package. These two failures block downstream validation of every lane on hosts with multicall coreutils or a fresh Chrome profile. Verification: - timeout 300 bash tests/fm-tmux-agent-liveness.test.sh -> exit 0, 16 assertions ok - timeout 700 bash tests/fm-calm-pi-extension.test.sh -> exit 0, 13 assertions ok, including the Pi operational follow-up loaded_off case - bash -n and shellcheck clean on both touched files - rest of tests/: bin/fm-test-run.sh --all bounded by timeout 900 completed 17 files with 0 failures (fm-afk-contract.test.sh through fm-backend-herdr-launcher-workspace-e2e.test.sh), then the bound cut off the 18th (fm-backend-herdr-presentation-e2e.test.sh, a real-herdr-gated lab test) with no failure recorded * fix(test): give wake-queue observation checkpoints the alerting ceiling tests/fm-wake-queue.test.sh's secondmate stall case runs bounded foreground watcher checkpoints whose job is to record an observation, with the alerting checkpoint that follows asserting the stall. A checkpoint's exit publishes a downtime marker, and the next checkpoint consumes it only by reaching the end of the watcher's poll loop, where the recovery surfacing runs after the stall tick; the observation itself is recorded by that same stall tick. On a loaded host a 1s ceiling sits under the cost of that iteration (which includes a pane capture in the active-turn gate), so the observation was never recorded, the downtime marker stayed pending, and the alerting checkpoint surfaced `check: rearm-resurface` instead of the stall it asserts: not ok - a foreign queue with no progress did not alert: check: rearm-resurface not ok - a frozen reprovisioned queue generation was hidden: check: rearm-resurface Give the observation checkpoints that feed a later alert the same 4s ceiling the file already documents for alerting checkpoints. The ceiling is only a bound - a checkpoint still returns on its first actionable wake - so no assertion is weakened, and the quiet windows get longer, not shorter. * no-mistakes(document): docs: correct export-DOM Chrome render root cause * no-mistakes(review): Isolate Chrome profile on macOS, dedupe tmux CC_BIN lookup * chore: re-trigger fork workflow approval for triage --------- Co-authored-by: Captain <blackxwhite88@users.noreply.github.com> Co-authored-by: kunchenguid <kunchenguid@users.noreply.github.com>
…chenguid#5383) * fix(bin): classify a status span without re-folding the whole log A watcher poll could take minutes, so its liveness beacon aged past the guard's 300s grace and the Stop auto-arm reported the watcher down. On the main home, cycles ended with beacon_age 91-235s while healthy and 534-706s while the laptop was CPU-starved. Cause: whenever a newly appended status span held a keyed needs-decision or blocked line, status_span_first_actionable_record re-read and re-folded the ENTIRE log to decide whether that opening was still live, forking several subshells per line. On a remote second mate's mirrored parent channel (1.2MB, ~2300 lines) that is 13-20k subshells, about 17s per log per classification when idle, paid by every signal and heartbeat scan. Nothing regressed recently: subshell counts per classification were 20,272 from kunchenguid#3268 (2026-08-29, which introduced the whole-log fold) and 13,188 from kunchenguid#3753 onward through HEAD. The cost grew with log size, since parent-channel logs only grow. Fix: fold only the captured span. An accepted opening does not depend on earlier lines and only later lines close or supersede it, and every later line lies inside the span, so the span fold names the same live openings at a cost bounded by the span. Old and new classification outputs are byte-identical across 51 span offsets of real-shaped secondmate and ship logs. A real-watcher regression test records every read the classification makes through the span-reader seam and asserts none reaches before the classified offset; it fails on the old code (5,157 bytes read from offset 0 to classify an 84-byte span). * no-mistakes(document): Clarify span classification and watcher regression coverage
kunchenguid#5362 and kunchenguid#4878) (kunchenguid#5381) * test: fix watcher timing flakes in fm-pr-check-security The bounded watcher's hang guard now counts only the watcher's own time: a case marks the intervals where it holds the watcher on injected work or makes it wait on concurrent work, and those no longer count against its budget. The budget itself stays at main's sixty seconds. The helper also stops forcing a one-second per-check timeout, which killed a correct merged poll whenever that poll took longer than a second, so the watcher only retried it or exited on a later check's wake without the merge. The concurrent-publication case pauses the guard while its arming is in flight, and its task now sorts before the contributions observer the arming also registers, so the watcher stops on the poll under test before running that unrelated fleet snapshot. The case also prints the watcher's stderr when it fails. The replacement case pauses the guard while the re-arm runs inside the watcher, runs that injected arming with the fixture root every other arming here uses, and waits on the replacement merge's process instead of a two-second cap. Merged-poll runs retire the contributions observer before the watcher starts, since no case here exercises it. The returned-descendant case no longer races a four-second sleep or a TERM landing at an arbitrary point in the watcher's idle loop: its descendant holds until killed, and a second check in the same cycle witnesses that it was drained and stops the watcher. * no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed * Revert "no-mistakes(ci): Reproduced the intermittent board-render failure. Its Lavish stub listed an open session but omitted the session-state record required by the listener, so the build could race the listener’s exit. Added matching fixture state; the affected suite passed three consecutive runs, and shell syntax and diff checks passed" This reverts commit 6a59859.
…enguid#5385) * feat: record task.pr_ready in the fleet ledger when a task PR is registered * feat: record worker status lines in the fleet ledger as they are written * no-mistakes(review): Keep worker status append failures and pass the resolved config to the ledger * no-mistakes(review): Resolve relative config override before embedding in worker command * no-mistakes(document): Clarify fleet ledger status capture timing
…unchenguid#5386) * test: synchronize foreign secondmate stall legs on the watcher's recorded observation Each leg of test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once ran the watcher under a 1s or 4s wall-clock checkpoint, but every later leg depends on the progress observation the previous leg's watcher recorded. Under load the watcher was killed before its first stall tick, the observation was never written, and the next leg treated its own sighting as the first one, so the stall alert never fired. Run the watcher directly and end each leg on its observable outcome: the progress marker recording the expected observation, or the watcher's own first wake. Also move a comment orphaned above this test back to the drain liveness test it describes. * no-mistakes(review): Wait for full stall reset before stopping watcher leg
kunchenguid#5391) The listener resolves its server from that store before it polls. Without a session for this board, it exits in the gap after the build has already sampled a live claim.
…#5390) * fix(bin): prune a torn-down task's wake rows at teardown Prune pending durable wake rows (.wake-queue) for a task when it is torn down, clearing stale wakes for its target window, signal wakes for its status or turn-ended files, and task-specific check wakes. Fixes kunchenguid#3419. Adjacent to kunchenguid#5252. - bin/fm-wake-lib.sh: add fm_wake_queue_prune_task - bin/fm-teardown.sh: call fm_wake_queue_prune_task in cleanup_firstmate_home_children and main teardown - tests/fm-wake-queue.test.sh: add test_wake_queue_prune_task * no-mistakes(document): docs: note teardown prunes a task's wake rows --------- Co-authored-by: Captain <blackxwhite88@users.noreply.github.com>
…ixture readiness (kunchenguid#5392) * Make portable tests match resolved host paths Summary: - Match macOS full Node command paths by basename in the Gemini behavior test. - Mirror symlink-resolved Nix PATH behavior and give the loaded-host race bounded headroom. Testing: - bin/fm-lint.sh - bin/fm-test-run.sh tests/fm-on.test.sh tests/fm-gemini-harness.test.sh tests/fm-procevent.test.sh Related: - None * no-mistakes(review): Mirror production PATH helper rules per directory group in tests * no-mistakes(test): Wait for orphan runner start marker instead of fixed sleep * no-mistakes(document): Clarify gemini ancestry test comment for versioned node comm --------- Co-authored-by: Sandeep Salwan <salwansa@amazon.com>
…ns (kunchenguid#5389) The sibling secondmate stall cases in tests/fm-wake-queue.test.sh now wait for the watcher's recorded observation instead of a one-second wall-clock checkpoint, so they can neither fail nor pass vacuously under load. Deterministic proof with a 5s watcher launch delay: before the fix 4 cases passed vacuously and 6 failed; after it all 10 pass on the recorded observation. Also includes a CI flake fix from validation: fm_control_harness_supported in bin/fm-control-lib.sh finishes reading the harness allowlist before returning, removing intermittent broken-pipe diagnostics. Behavior is unchanged.
… tail (kunchenguid#5336) * fix(bin): refuse a Herdr submit that would send only a message tail A long typed payload can sit in the composer as a suffix, or as a paste placeholder plus a remainder, and the following Enter was still reported as delivered. Prove the selected composer holds the payload before Enter, and report failure when it does not. * no-mistakes(review): Scope Herdr payload proof to Claude, clear composer on refusal * no-mistakes(test): Clear refused Herdr composer drafts one wrapped row per press * no-mistakes(test): Accept Claude's multi-line paste placeholder in Herdr submit proof * no-mistakes(review): Accept Claude read-back that drops U+2063 in Herdr proof * no-mistakes(document): Document Herdr proof ignoring U+2063 operational mark * no-mistakes(ci): I made a one-line test change. The failing check comes from a timing race in an existing test that this PR doesn't touch. **What failed:** `tests/fm-procevent.test.sh` failed at "the superseded paced runner invoked its stale command" (line ~3313). The PR only changes the Herdr files and their tests, and the same shard passed on main at the base commit. **Why it can fail:** the fixture starts a second runner with a 3-second launch floor (the minimum wait since the source's last launch). That runner sleeps for the rest of the floor and only then checks whether its registration was replaced (`fm_procevent_launch_floor_wait` in `bin/fm-procevent-lib.sh`). The test then waits for the claim and re-registers the source. If that takes longer than about 3 seconds after the first launch, the old runner wakes up, finds its registration still current, and runs the stale command. That produces the second log line the test reports. The CI shard was slow (this one test took 160 s). **Fix:** in `tests/fm-procevent.test.sh` I raised the superseded runner's floor from 3 to 15 seconds and added a comment explaining why. The floor now outlasts the fixture setup even on a loaded runner. Nothing else changed: the first launch and the later fresh-registration start still use a 3-second floor, and no product code changed. **Verification:** - The full test file can't give a reliable result on this machine (load average about 64 on 8 cores). It failed earlier, at the reconcile assertion around line 1680, before it reached this section. - I ran the changed section by itself (file setup plus the pacing-race block) five times with the fix: all passed, in about 9-13 s each. - The original code also passed five out of five, so the race didn't reproduce locally. The diagnosis rests on the code path and the CI log. - I haven't seen the full file or the CI shard pass with the fix yet * no-mistakes(test): Accept Claude folder-trust prompt via down+enter in live e2e * no-mistakes(document): Note unreadable Claude composer refusal in Herdr docs
…unchenguid#5427) Speaking as Kun's firstmate: squash-merging — opt-in (forge=gerrit registry-gated; default project-mode stdout restored to two words), attestation MATCH, CI+NM green, safe review, MERGEABLE.
…uid#5358) * feat(bin): add an opt-in per-home worker account pin A home that mixes work and personal accounts for one runner had no way to say which account its workers launch on: Claude workers inherited whatever CLAUDE_CONFIG_DIR the supervising process had, Pi workers the pane's ambient root, and an ambient API key outranked both, with no signal at launch. config/claude-account and config/pi-account now pin that choice per home. With neither file every launch is unchanged. With one, every launch of that runner from the home (ship, scout, local secondmate, raw Claude command, and relaunch) runs under the declared root, and the spawn refuses before any endpoint exists when the file is malformed or the runner's own check (claude auth status, pi auth check with a model-listing fallback) says the pinned account is not signed in. The check runs in a cleared environment so an ambient credential cannot answer for an empty root. A pinned Claude launch sheds the environment credentials Claude ranks above a stored login; a pinned Pi launch needs an explicit <provider>/<id> model for a declared provider and also carries --provider. The chosen account is printed on the spawned line and recorded in the task record, and relaunch checks the pin before stopping the running agent. * test(secondmate): give the concurrent config-push wait room for a slow host test_config_reread_serializes_concurrent_pushes waited about two seconds for the first fm-config-push.sh to reach its first send-keys. On a slower host that push takes four to five seconds, so the test failed on main before the push ever got there. The loop still leaves as soon as the marker appears, so the larger bound costs nothing where the push is fast. * no-mistakes(review): Refuse raw Claude account overrides under a pin
…ts (kunchenguid#5470) * fix(bin): keep the Herdr lab session option before a -- delimiter fm-herdr-lab.sh run appended --session <lab> after every argument, so a command with a passthrough delimiter such as agent start ... -- <agent args> handed the session flag to the agent and Herdr routed the call by the caller's ambient socket instead of the lab. The helper now inserts --session <lab> immediately before the first -- delimiter and keeps the trailing form otherwise. * no-mistakes(document): Clarify Herdr lab session option placement
…nguid#5554) * fix(bin): bound the away digest and log why a delivery failed The away daemon joined every buffered escalation into one unbounded digest. A start-up catch-all span can exceed what one transport argument carries (tmux rejects the send-keys command; Linux refuses to exec any argument above 131,071 bytes, which is how herdr receives it), so the initial send failed on every housekeeping pass and was logged as an unconfirmed Enter with text possibly in the composer. escalate_flush now builds the injected digest under a fixed byte budget: each event is cut at a UTF-8 boundary with an omitted-bytes marker, the joined events stop with a "+K more event(s)" tail, and a bounded digest names a state/.subsuper-digests/ file that keeps every buffered event verbatim. The buffer itself is untouched, so the return catch-up stays complete. The tmux submit core and the herdr literal send now replay the transport's stderr on failure, and inject_msg logs the failing stage (initial send versus Enter confirmation) with the byte count and that stderr. The wedge alarm line and marker carry the last failure reason. Fixes kunchenguid#4382 * no-mistakes(review): Drop digest pruning; label send-failed as send-or-Enter stage * no-mistakes(review): Keep digest full text once submit ran; reuse on retry * no-mistakes(lint): Count digest files with find instead of ls --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…henguid#5638) * feat(tests): add FM_TEST_SEAM launch seam and gate lab-primary recipe Part 1 of the kunchenguid#5615 split: the pieces that let the no-mistakes pipeline live-validate firstmate changes, without the gate-refusal rescoping. - bin/fm-afk-launch.sh: FM_TEST_HARNESS pins the detected harness only alongside the FM_TEST_SEAM=1 marker test suites set, so a leaked variable in a real primary's environment stays inert and unknown tokens fall through to real detection. - tests/lib.sh: export FM_TEST_SEAM=1 for every suite. - .no-mistakes.yaml: per-harness recipe for running a real fixture primary from a gate run - a plain mktemp lab FM_HOME on a private tmux socket, with FM_GATE_REFUSE_BYPASS=1 scoped to it and NO_MISTAKES_GATE scrubbed. - tests/fm-wake-queue.test.sh: stop the owned watcher fixture with KILL and clear its lifecycle state so the next leg starts clean; TERM could leave bash waiting in a child on some runners. - tests/fm-remote-secondmate-lifecycle-e2e.test.sh: wait for the liveness lock holder's post-acquire marker instead of the lock dir, which is published before the claim finishes. * no-mistakes(review): Scrub lab home overrides and require FM_TEST_SEAM separately * no-mistakes(document): Clarify test seam and disposable lab bypass documentation * no-mistakes(document): Clarify lab isolation and test-seam documentation * no-mistakes(ci): Fixed the CI failure: test cleanup killed the remote worker child but left its supervisor able to restart it during fixture removal. Cleanup now stops the worker tree. The lifecycle test passed locally; ShellCheck and diff checks passed
… home is gone (kunchenguid#5552) * fix(bin): refuse watchers from disposable checkouts and exit when the home is gone Fixes kunchenguid#321 Fixes kunchenguid#4760 A watcher armed from a disposable no-mistakes validation checkout under .no-mistakes/worktrees/ outlived the validation step and kept writing the real home's state, and a running watcher never noticed when its home, state directory, or code root disappeared. The arm now refuses from such a checkout with the typed failure line, the watcher checks once per poll that its home, state directory (or its own lock holder record), and bin directory still exist and exits with a logged reason scoped to itself, and the shared test helpers reap every watcher a suite armed for a temporary home through the home-scoped stop. * no-mistakes(lint): fix SC1007 by assigning empty string in watch-arm test * no-mistakes(ci): Found and fixed a genuine, reproducible hang introduced by this branch's test-watcher reaper, which is what killed both CI checks (serial-2 cancelled at the 30-min cap; Lint 2 exit 143 = the suite's own TERM-trap code). Root cause: test_drain_asserts_watcher_liveness (tests/fm-wake-queue.test.sh) fabricates a .watch.lock whose pid is the test runner's own $$ with the runner's real identity, to make the drain believe a live watcher exists. The new make_case tracking registers that state dir for reaping, so at fm_test_cleanup the new fm_test_reap_watchers drives fm-watch-arm.sh --stop; its identity check matches (the fixture recorded the runner's identity) and it kill -TERMs the test runner. tests/lib.sh:231 is `trap 'fm_test_cleanup; exit 143' TERM`, so the TERM re-enters cleanup -> reap -> kills $$ again -> infinite loop until the runner cap. I reproduced this locally: the suite ran all tests then looped forever in cleanup spawning fm-watch-arm.sh --stop against a lock naming its own PID. Fix (tests/lib.sh, +5 lines): in fm_test_reap_watchers, skip any tracked lock whose pid equals our own $$ before driving --stop. This is the single shared reap boundary; seven $$-self-lock fixtures across four test files are all covered by the one guard, and real armed watchers (pid != $$) are still reaped. Invariant: the test reaper must only signal real armed watcher processes, never the test runner itself. Verified locally: tests/fm-wake-queue.test.sh -> EXIT 0 (63 ok, no hang); tests/fm-watch-arm.test.sh -> EXIT 0 (21 ok, including test_reaper_stops_a_tracked_watcher, confirming the guard does not over-skip). Lint 2's exit 143 was the same shard/cap signature; a fresh CI run on this new commit will re-evaluate it --------- Co-authored-by: firstmate-oss <firstmate@kunchenguid.local>
…m as silence (kunchenguid#5588) * fix(bin): surface an unrecognized status prefix instead of dropping it A parked or holding declaration, and a verb whose correlation token did not parse, never became an event, so the supervisor still saw the earlier line. * no-mistakes(review): Require verb-shaped unrecognized status prefixes, add continuation tests * no-mistakes(document): Document unrecognized status prefix escalation in afk skill * no-mistakes(ci): I reproduced the "Behavior portable serial 6" failure locally and fixed it by changing the test data in one test. No product code changed. **What failed:** `tests/fm-session-start.test.sh`, in `test_orphan_status_logs_are_printed`, with "matched status log was printed 2 times". **Why:** the test writes status lines with made-up prefixes, `matched: surfaced once` and `orphan: step N`. The test only uses them as placeholder text. It checks that the session-start digest prints each task's status tail exactly once. This PR (kunchenguid#4763) deliberately makes an unrecognized one-word lowercase prefix a status event. So those lines now surface as captain-relevant events, and the wake queue's STATUS OUTCOME BACKSTOP section prints them a second time. The code under review is behaving as the issue asks. Only the test's placeholder data had become meaningful. **Rule the test depends on:** its status lines must not be captain-relevant, so the digest is the only place they are printed. Both lines in this test broke that rule. The orphan line would have failed the same count check right after the matched line did. **Fix:** in that test only, I switched both lines to the recognized, non-captain verb `working:`: `working: surfaced once` and `working: orphan step 1..6`. I updated the matching assertions and counts to use the new text. What the test checks is unchanged: orphan logs are labelled, the tail is bounded, the log path is printed, and each tail appears once. **Verification:** before the fix, the test failed locally the same way as in CI. After it, `bash tests/fm-session-start.test.sh` reports "all assertions passed * no-mistakes(review): Detect unrecognized prefixes on unstamped lines; share verb list --------- Co-authored-by: Kun's firstmate <kunchenguid+firstmate@users.noreply.github.com>
…henguid#5658) Fixes kunchenguid#5295 Session start now reports a remote inheritance failure using the push's own error line instead of the first unchanged item that happened to print before it, and the shared captain preferences header check now names the first required phrase it did not find, on both the local and remote inheritance paths.
…yloads (kunchenguid#5657) * fix(bin): stand down the Claude Stop auto-arm on pi-code-delivered payloads pi-code loads the tracked Claude settings but has no asyncRewake, so it awaits every Stop hook; without a stand-down the auto-arm runs synchronously inside Pi's turn end and holds it open for the declared multi-hour timeout. Stand down when the payload's transcript_path contains a /.pi/ path component, the same discriminator the closed-but- unmerged fix in kunchenguid#3352 used, with an explicit string-type check on the jq filter. Fixes kunchenguid#3343 * no-mistakes(document): document pi-code stand-down in harness integrations reference
…kunchenguid#5659) * fix(bin): match whole multi-word project names in the registry lookup bin/fm-project-mode.sh matched a registered project name against only the first whitespace-delimited token of a registry row, so a name containing a space never matched, silently defaulting the project to no-mistakes off instead of its declared posture. The lookup now matches the whole registered name against the raw line text, so a name is compared literally (never as a regex) and a name that is a leading prefix of another registered name still resolves to its own row. * no-mistakes(document): docs already accurate for multiword registry name match * chore: drop accidental empty err file Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
…d#5546) * fix(bin): classify the stdin program of `bash -s` with operands in the arm policy With -s, sh/bash/zsh read the program from stdin even when operands follow; the operands are only positional parameters. The arm policy treated the first operand as a script path, so heredoc and here-string payloads were never classified and a hidden bin/fm-watch.sh execution was allowed. A protected path in the operand position still fails closed as before. Fixes kunchenguid#1489 * no-mistakes(document): Clarify stdin shell operand documentation * no-mistakes(ci): Captain, fixed `shellInvocation` so `bash -- -s` treats `-s` as a script name, and updated R21 to test the exact command. The targeted policy suite, lint, documentation check, and diff check pass. Both hosted workflows show `action_required` before any jobs ran; that external approval state remains unresolved * fix(bin): keep main's handling of words after a leading `--` Revert the pipeline CI-step change that made the first word after a leading `--` always a script. It turned forms that main denies today into allow (for example `bash -- -c 'bin/fm-watch.sh'`), which is outside kunchenguid#1489 and loosens a fail-closed policy. `--` after `-s` still ends option parsing.
The changed-mode exclusion test wrote tests/fm-lint-local-exclude-fixture.test.sh into the checkout, so fm-test-run.test.sh running concurrently in the same proven-isolated lane could enumerate it and fail its serial-shard coverage check. Name the fixture *.sh so it stays inside lint's tests/*.sh set but outside the tests/*.test.sh inventory.
…hang bounds Upstream added three fixed iteration-count waits (a remote watcher's auto-relaunch exit, the remote liveness lock holder, and the bootstrap liveness lock holder) after the fork moved the same suites onto fm_wait_for_marker's liveness-aware hang bound. Bring the new waits onto that shape so a loaded host is not read as a hang.
The readable-sibling check iterated an unquoted string, which zsh never word-splits, so every adapter looked unreadable and fm_backend_source failed when sourced from zsh. Its local named path also shadowed zsh's PATH-tied array while the adapter was sourced. Use an array and a distinct variable name; tests/fm-backend.test.sh's zsh case covers it.
…nchenguid#5695) * fix(bin): strip AI co-author trailers from fleet-launched commits Cursor and other non-Claude runtimes append the trailer after the typed message. A per-task commit-msg hook removes it and leaves human co-authors and the author identity untouched. * no-mistakes(review): Export pane hooksPath override and drop generated-with stripping * no-mistakes(ci): This PR caused all three CI failures, and the fix is test-only: 4 test files change, no product code. **Cause.** `fm-spawn.sh` now installs the AI-trailer strip hooks for every spawn, secondmates included. The installer refuses a worktree that is not a git repository, and the PR deliberately keeps that fail-closed rule because real secondmate homes are firstmate clones. Four test fixtures still gave secondmates a plain directory as their home, so each spawn failed with "not a git worktree ... could not install the AI-trailer strip hooks": - serial 5: `tests/fm-backlog-atomicity.test.sh` ("secondmate spawn failed"). - serial 8: `tests/fm-secondmate-harness.test.sh` ("split: no meta written"). - Herdr: `tests/fm-backend-herdr-launcher-workspace-e2e.test.sh` and `tests/fm-backend-herdr-workspace-per-home-e2e.test.sh`. **Rule that must hold.** Every home a test spawns as a secondmate must be a git worktree. I checked the other places in the changed area: the only secondmate spawns in these tests are the ones listed. The earlier rounds already fixed the other fixtures (`fm-secondmate-liveness`, `fm-secondmate-safety`) the same way. **Fix.** - Each of those four secondmate homes now gets the same `.gitignore` plus `git init -q -b main` that the liveness and safety tests already use. - The two Herdr tests clean up with their own plain `rm -rf "$TMP_ROOT"`, not the shared `tests/lib.sh` helper. Because the installer leaves each `state/<id>.git-hooks` directory read-only, that cleanup printed "Permission denied" and left the directories behind. Both cleanups now restore the owner's write bit on every directory before removing (`find ... -exec chmod u+rwx`), which is what `fm_test_remove_tree` in `tests/lib.sh` does. **Verification.** - `tests/fm-secondmate-harness.test.sh` passes. - `tests/fm-backlog-atomicity.test.sh` passes (99 ok, exit 0). - shellcheck is clean on all four files. - I could not run the two real-Herdr tests locally: the Herdr lab on this host refuses to start because it needs exactly one running default session, and I did not change the host's Herdr state to get around that. Instead I checked their two changed steps directly: the installer succeeds on a home set up the new way, and the new cleanup removes the read-only hooks directory completely. Those two tests will only be proven on CI
…id#5683) * fix(bin): treat Pi's dollar-first cost footer as furniture An idle Pi status row opening with $0.000 was read as a dead-shell prompt, so exit and relaunch refused on an empty composer. * test: wait for the draining holder to exec sleep before reading its identity The procevent drain fixture read fm_pid_identity immediately after backgrounding setsid sleep, racing the child's exec chain. Mid-exec the cmdline can read empty, failing the fixture on a loaded CI runner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…5534) * fix(bin): refuse a merge when a required check never reported fm-pr-merge.sh built its GitHub refusals only from checks present in statusCheckRollup, so a required check that never ran was simply absent and the merge proceeded on the subset that reported, contradicting its own "every required check green" claim. The GitHub verify now reads the base branch's required contexts from the forge itself - the classic branch protection summary on GET repos/{o}/{r}/branches/{b} and the active ruleset rules on GET repos/{o}/{r}/rules/branches/{b} - and refuses when a required context has no entry in the same rollup, at the same head, that the merge is bound to. Absence reads as unknown, never green. The required-set read joins the existing refusal list, so a draft, a red check, and an unreported required check are all reported together. Could not read vs nothing required: both endpoints need only repository read access. The admin-only GET .../branches/{b}/protection endpoint is deliberately not used: it answers a non-admin token with the same 404 an unprotected branch gets (observed live on kunchenguid/firstmate main with this token), which would read a missing permission as "nothing required". Any failed or malformed read of either source (auth, missing fine-grained permission, rate limit, network, 404, unexpected shape) refuses the merge with a line naming the unreadable source. The one exception is GitHub's plan-gated 403 on the rules endpoint ("Upgrade to GitHub Pro or make this repository public"), which already means "this repository has no branch rules" for the merge-queue reader; that check moves into one shared helper and the classic summary still decides for such a repository. Attended waiver: --allow-missing <check-name> is the twin of --allow-red and follows the same design and recording path: once, separate name argument, waives only that exact unreported required check, still requires every other required check reported and every check green, never waives an unreadable required set, refused while the away-posture record exists, and refused on GitLab. Merge-state BLOCKED policy is unchanged. How this differs from the withdrawn kunchenguid#5353 (read from its diff): - kunchenguid#5353 read the admin-only branches/{b}/protection endpoint and treated its 404 as "no required checks", so for any non-admin token the required set silently read as empty; this change reads the read-access branch summary and treats every failure as unreadable. - kunchenguid#5353 ignored rulesets; this change also reads required_status_checks rules from the effective branch rules. - kunchenguid#5353 made separate per-head REST reads of statuses and check-runs capped at per_page=100 with no pagination; this change checks presence in the same statusCheckRollup view the red-check gate already reads at the verified head. - kunchenguid#5353 stopped at the first unreadable read; this change reports it as one refusal among all the others. - kunchenguid#5353 also claimed kunchenguid#5345 (lock stealing) and changed 39 files, most unrelated; this change is kunchenguid#5344 only. Live proof, read-only (a gh wrapper refused every merge and mutating call): - cli/cli#14474 (trunk requires 3 classic build contexts, none ran): refused, naming build (macos-latest), build (ubuntu-latest), build (windows-latest); with --allow-missing "build (macos-latest)" it still refused, naming the other two. - cli/cli#13665 (ran build (ubuntu-24.04-firewall) instead): refused, naming build (ubuntu-latest). - hashicorp/terraform#39262 (ruleset-required checks absent): refused, naming Code Consistency Checks, End-to-end Tests, Race Tests, Unit Tests. - cli/cli#14485 (all required reported and green): verified; the wrapper blocked the merge call and the pull request read back open. Fixes kunchenguid#5344 * fix(review): Preserve required-check producers and aggregate independent read failures * fix(document): Clarify required-check verification and waiver documentation * fix(bin): match an app-bound required commit status by name The producer-identity check resolved an app-bound required context only against check runs, so a required context that the required app reports as a commit status could never match and always read as "has not reported". A commit status carries no app id to compare, so an app-bound requirement that arrives as a status now matches by name, as before producer binding; check runs keep requiring the configured producer app. Live, read-only: hashicorp/terraform#39262 requires license/cla from integration 865473, reported green as a commit status by the CLA app. The previous head refused it as unreported; this head no longer does, while still naming the four required check runs that never ran there. Refs kunchenguid#5344 * fix(document): Clarify accepted commit-status producer verification limitation --------- Co-authored-by: firstmate-oss <firstmate-oss@kunchenguid.local>
Upstream's pre-Enter payload proof reads a short tail of the pane, but a typed /exit opens a Claude completion popup of about twenty rows below the composer, so the tail held only popup rows and the proof refused the send. fm-control exit and relaunch then stopped at "the exit command could not be sent" before the fork's exit-dialog answer could run. A /- or $-prefixed payload now reads the whole bounded window. Verified live on Claude Code 2.1.280 through Herdr 0.9.0 with the exit-dialog and submit-confirm live guards.
… loaded hosts The watcher leg ran the remote relaunch under the production 120-second bound and the unreachable leg checked for its probe after a fixed four seconds; a loaded host outlasted both. Give the relaunch the suite's hang bound and wait for the probe itself.
# Conflicts: # tests/fm-spawn-dispatch-profile.test.sh
# Conflicts: # AGENTS.md # bin/fm-crew-state.sh # docs/configuration.md # tests/fm-crew-state.test.sh # tests/fm-spawn-dispatch-profile.test.sh
…mes onto upstream's new tests Upstream's Pi dollar-footer test drove the fake Herdr with FM_HERDR_* variables, which the fork's Herdr server launch now strips, so it uses the FAKE_HERDR_* names every other fake in the suite uses. Upstream's AI-trailer hook installer requires a secondmate home to be a git worktree, so the fork's child-session secondmate fixture is initialized the same way upstream's own fixtures are.
…ener The orphan-reaping precondition required the listener's parent to be pid 1, but on WSL (and under a systemd user manager) an orphan is adopted by a child subreaper instead. Prove the orphaned state by the listener no longer having this suite's shell among its ancestors.
…eanup Upstream's spawn now installs a read-only state/<id>.git-hooks directory, so the plain rm -rf in the fork's Claude exit-dialog and transcript live guards and in the remote secondmate lifecycle suite left fixture roots behind. Clean up through tests/lib.sh's fm_test_remove_tree, which restores write permission first, and fail when the fixture root survives.
# Conflicts: # bin/fm-composer-lib.sh # bin/fm-supervise-daemon.sh # docs/configuration.md # docs/herdr-backend.md # tests/fm-afk-launch.test.sh # tests/fm-spawn-dispatch-profile.test.sh
…tection Upstream's afk-launch suite pins FM_TEST_HARNESS=claude through its test seam, which outranked the real codex ancestor #13's two end-to-end cases launch under to prove the daemon receives the detected primary harness. Switch the seam off for exactly those two launches.
…omposer #13 made a bare agent-glyph row in a Claude primary's pane read unknown (it is a themed shell prompt), and moved the daemon suite's delivery fixtures to Claude's framed idle composer. Upstream's two new unknown-wake delivery tests still typed into a bare glyph, so they use the same fixture.
Owner
Author
Conflict resolutions in detailBrings upstream kunchenguid/firstmate Conflict resolutions
Semantic follow-ups after the textual merge
Fork fixes verified intact
Keeping up with main and upstreamFork PRs #3, #9, #11, #10 and #13 merged to main during this work and are folded in with merge commits, and a second upstream merge brings in the three upstream commits that landed meanwhile.
Local verification
|
… run of the suite failed on a different test partway through. I did not re-run it. **The change.** First I discarded the previous round's uncommitted edits to `tests/fm-remote-secondmate-lifecycle-e2e.test.sh`. Then, in `cleanup()`, I deleted only the left-behind check: the `if [ -e "$TMP_ROOT" ] ... exit 1; fi` block (4 lines). The `fm_test_remove_tree "$TMP_ROOT"` call and its comment are still there. No other file changed. That check was the only thing failing on CI: every test in the suite passed there, then cleanup exited 1 because a late writer was still filling the fixture root. **The local run.** I ran `bin/fm-test-run.sh tests/fm-remote-secondmate-lifecycle-e2e.test.sh` once. The host load average was about 190, and the run took 1393s, against 278s on CI. Every test passed until this one: `not ok - the dead remote secondmate was not auto-relaunched: check: secondmate ios auto-relaunch failed after remote endpoint dead on its configured host: remote job worker stopped before this job completed` The suite then stopped with exit 1. That test runs before `cleanup()`, so my change can't affect it, and it passed on CI. My guess is that it's a timing failure caused by the overloaded host, but I haven't confirmed that. The next CI run will show whether this check now passes. I left ci-2 (PR must be raised via no-mistakes) alone as instructed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
go with option 1
Context: option 1 was: a worker merges upstream Firstmate into the captain's fork as a PR, resolving the clashing files so both sides' changes survive, through the full automated review, with the captain making the merge call; after it merges, the normal Firstmate update brings everything to the running Firstmate and its second mate. The captain wants every upstream Firstmate update, and never opens PRs against upstream kunchenguid/firstmate; the fork dardant/firstmate is home base. At the time of asking, upstream/main (b805823) was 68 commits ahead of the fork's main (36f9b51), the fork had 6 commits of its own (fixes #1, #2, #4, #5, #8 and an earlier upstream merge), and a merge clashed in bin/fm-test-run.sh, docs/herdr-backend.md, tests/fm-contributions.test.sh, tests/fm-control.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh, tests/fm-secondmate-harness.test.sh and tests/fm-secondmate-liveness.test.sh.
What Changed
8d2ee291) into the fork. This brings in upstream's new features: the opt-in Claude away supervision host (bin/fm-supervision-host.sh,bin/fm-supervision-engine-lib.sh), away supervision for non-Pi primaries, the Devin CLI crewmate/scout adapter, the opt-in fleet activity ledger, Gerrit publish/watch on forge-bound projects, per-home Claude/Pi worker account pins, a configurable ship-branch prefix, and stripping of AI co-author trailers from fleet-launched commits. It also brings upstream's manybin/fixes, including watcher eviction and teardown safety, merge refusal when required checks are unreported, auto-relaunch of dead secondmates, and the Calm queued-input handling in the Pi extension. Upstream's docs and skill rewrites come along too.bin/fm-test-run.sh,docs/herdr-backend.md, and the contributions, control, secondmate-harness, secondmate-liveness and remote-secondmate-lifecycle tests) so the fork's own fixes (fix(bin): keep transcripts for Claude agents launched from a Claude primary #1, fix(bin): keep Firstmate FM_* variables out of a Herdr server a state read starts #2, fix(bin): answer Claude's background-work exit confirmation in fm-control exit and relaunch #4, test: make host-sensitive Firstmate suites pass on WSL, loaded hosts, and hosts without ruby #5, fix(bin): keep CI lint partitions inside a 16 GB runner #8, fix(bin): prove the primary harness owns the pane before away-mode injection #13) survive next to upstream's changes. Adds two fork-sidebinfixes on top:fm_backend_sourcecan load adapters from zsh again, and a Herdr Claude slash command is only treated as sent once it is confirmed past its completion popup, with bounded Ctrl+U clears.🤖 Generated with Claude Code
Risk Assessment
Testing
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ No issues found.
⏭️ **Test** - skipped
Step was skipped.
✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.