Repository navigation
chore(sync): merge upstream main into house - #211
Merged
Merged
Conversation
kunchenguid#6708) * fix: reopen a remote-reply continuity break after repair A later break for the same route and reason was swallowed after the operator resolved the first one, because the status line matched for the life of the log. The continuity ingest now appends again when the cursor has moved or retirement has reset that episode, and an unchanged re-read still appends nothing. status_event_recorded is unchanged. Its other callers are the pending-reply escalation, which already decides its own episode, the parent-channel note append, and the remote document transfer note. * no-mistakes(review): seed continuity episode for already recorded break line * no-mistakes(document): document when a remote-reply continuity break reopens * no-mistakes(ci): The adapter now stores the continuity episode record before it appends the `blocked` line, so a failed store appends nothing. The changes are uncommitted in the worktree, in `bin/fm-procevent-remote-reply.sh` and `tests/fm-remote-reply.test.sh`. **Invariant:** a continuity `blocked` line is on the parent status log only when the episode record for that break is already stored. Only the `continuity-broken` branch of `cmd_ingest` appends that line, so the fix is at that one place. **What changed in `cmd_ingest`:** - It decides first whether the break needs a line, then stores the episode record, then appends. - If storing the record fails, it stops with "cannot record continuity episode" and appends nothing. - If the line is already on the log and no episode record exists, it stores the record and does not append. **One addition you did not ask for:** if the append fails after the record is stored, the adapter puts the earlier record back (or deletes the new one when none existed). Without that, the retry would read as the unchanged repeat and append nothing, which is the issue 6701 failure again. **Deviation from your test instruction:** the test does not use `chmod` on the cursor directory. `write_continuity_episode` runs `chmod 700` on that directory before every write, so a read-only directory is made writable again. The test instead makes `mktemp` fail for the episode's temporary file in that directory, the same way the existing receipt-failure test does. **Tests added to `tests/fm-remote-reply.test.sh`:** - Store failure: the second break exits 1, appends no `blocked` line and opens no decision. One retry after storage recovers exits 3 and appends a single line. - Append failure: with the status log read-only, the third break exits 1 and appends nothing. One retry after the log is writable exits 3 and appends a single line. **Verification:** `bash tests/fm-remote-reply.test.sh` ends with "ALL TESTS PASSED" with the fix. Against the script at commit b750c828 the same test file fails at "a continuity break appended its line before its episode was stored". `shellcheck -S warning` reports only an unused loop variable at line 667 of the test file, which this change does not touch. I ran no other test files. The Greptile Review check log could not be retrieved, so I worked from the finding text alone * fix: record a continuity break's reader position on its status line A later break at another cursor is then a different line, so the existing duplicate check appends it and reopens the decision. An unchanged re-read builds the same line and appends nothing. * fix: reopen a continuity break after an identical restore A retirement that puts the same bytes back used to rebuild the recorded line, so the later break stayed closed. The retirement count on that line makes the later break distinct. * no-mistakes(review): remove continuity match for full-prefix line without retirement count * no-mistakes(document): clarify what a continuity break status line records * fix: remove the reply cursor before recording retirement A stop between those steps must leave the count unchanged, so an unchanged continuity break still builds the same line.
…is retired (kunchenguid#6733) * fix(control): drop busy_gen when an incarnation is retired A deliberate exit removed the busy sidecar and left busy_gen in the task record, so the two records disagreed about whether that incarnation was still observable. * no-mistakes(review): drop GNU-only chmod and unreached sidecar-absent branch * no-mistakes(review): correct lock comment to name the deadlock * no-mistakes(ci): The test `test_exit_drops_meta_busy_gen_with_the_sidecar` in tests/fm-control.test.sh now compares the whole task record (the `state/<id>.meta` file), so the Greptile finding is fixed. Invariant: after `exit` retires an incarnation, the task record must equal the record from before `exit` with only the `busy_gen` line removed. This test is the only place in the change that asserts the record survives the rewrite, so it is the only site to fix. The other `busy_gen` tests assert that the line stays, and they do not go through the rewrite. What changed: before `exit`, the test writes the record without its `busy_gen` line to `expected.meta`. After `exit`, the test runs `diff` between that expected copy and the real record, and fails with the diff output if they differ. This one comparison replaces the two earlier checks (no `busy_gen` line left, and the `window` line present), because it covers both. I did not change bin/fm-control.sh or any other file. How I know it works: - I ran `bash tests/fm-control.test.sh`: exit code 0, 45 lines starting with `ok`, no other lines. - I temporarily changed the rewrite in bin/fm-control.sh to also drop the `harness` line. The test then failed with `not ok - exit should drop only busy_gen from the task record:` and the diff `< harness=codex`. The earlier `window`-only check would have passed that rewrite. I restored bin/fm-control.sh afterwards; `git status` shows only tests/fm-control.test.sh modified. - `bash -n` and `shellcheck` on the test file report no new warnings from the edit. The change is not committed; the working tree holds it
…#6484) * test(secondmate-harness): scope fake ps -codex label for pid 5252 to the liveness probe Closes kunchenguid#6456 * no-mistakes(ci): Updated the collision test to log and assert that PID 5252 was queried before selecting 4242. Full fm-secondmate harness suite passes --------- Co-authored-by: YifuGu <ironerumi@users.noreply.github.com>
* fix(calm): share the standalone Pi Calm working-ship widget slot Firstmate Calm and the user-global standalone Pi Calm both install an animated working-ship widget during agent runs. Each claimed its own Pi widget key, so a session loading both (the main Firstmate home) rendered two boats. Pi replaces widgets under one key, so claiming the shared "calm-working-ship" slot keeps dual-install sessions to a single boat while a Firstmate-only session is unchanged. Pins the shared slot contract in the working-ship module test so the key cannot silently diverge again. * test(calm): pin the shared working-ship widget key in CI, document dual-install The key-parity assertion inside the Pi fixture only runs where the @earendil-works/pi-coding-agent package is installed, so CI never exercised it. Add a source-level twin that needs nothing but the tracked file, and note in docs/calm.md that the boat shares the standalone Pi Calm working-row widget slot. * no-mistakes(review): Add executable dual-install widget replacement coverage * no-mistakes(review): Guard shared widget cleanup with disposal ownership * no-mistakes(document): Document shared Calm working-ship slot behavior * test(calm): read the standalone Calm slot from its own module The dual-install check registered both boats itself under the shared slot, so it could only prove that Pi replaces a widget under one key: it would still pass if the standalone Pi Calm extension installed its boat under a different key, which is the two-boat regression the check exists to prevent. Read the standalone extension's own working-ship module when it is installed - FM_STANDALONE_CALM_SHIP, else ~/.pi/agent/extensions/calm - and drive the check with the key that module exports, so a rename on either side registers two widgets and fails naming both keys. A pinned shared-slot contract still covers a machine without the extension, and the run reports which side it used instead of passing silently over an absent extension. Verified: the touched Pi Calm suite passes and reads the installed standalone extension; with a copy of it whose key is renamed to calm-working-ship-v2 the suite fails naming the drift. * no-mistakes(review): Gate stock-row restoration by shared-widget ownership * no-mistakes(review): Removed redundant widget-key source assertions * no-mistakes(document): Document shared Calm working-ship widget ownership
…id#6649) * fix(herdr): make exact-resume presentation-lock wait instead of a bounded timeout The exact-resume path in bin/fm-spawn.sh used the same 50-attempt-then- give-up lock acquire as the new-task-create path, but the two paths are not equivalent on contention: a create has no prior state to strand and can safely fall back to a flat layout, while a resume is recovering a specific existing identity that a concurrent recovery may legitimately be holding the lock for. Giving up there does not degrade gracefully, it hard-fails the resume outright. The suite's own concurrent cross-home recoveries test already asserts both concurrent recoveries succeed with a genuine reclaim, and the file's header comment already (inaccurately) claimed lock contention falls back to the ordinary flat layout for both paths alike, so the intended contract was always that recoveries serialize and both succeed, not that either one refuses under a short bound. Give spawn_herdr_presentation_order_lock_acquire a wait mode that uses this file's own established fm_lock_acquire_wait idiom (already used for its other fleet-shared locks) instead of the bounded loop, and use it only at the exact-resume call site. The new-task-create call site is unchanged and keeps its bounded-then-flat-fallback behavior, which is already covered by its own passing test. Dead-owner PID-liveness reclaim inside fm_lock_try_acquire still bounds the wait against a holder that crashed mid-hold. Adds a deterministic regression test that holds the shared session lock from an unrelated process for well past the old bound, then asserts the resume succeeds with a genuine reclaim and took close to the full hold duration, so a fix that merely widens the bound rather than genuinely waiting is still caught. The existing concurrent cross-home recovery test exercises this under real timing but does not reliably outlast a fixed bound on its own. Corrects the header comment's claim that create and resume share one bounded-then-flat-fallback behavior on lock contention; they no longer do. * no-mistakes(document): Document Herdr recovery waiting for presentation lock * no-mistakes(document): Update stale hard-refusal claim in verification log * no-mistakes(ci): Fixed the Greptile finding on tests/fm-backend-herdr-presentation-e2e.test.sh:1389 by bounding the resume lock-wait regression's spawn_task call. Added an optional 4th `deadline_seconds` arg to the `spawn_task` helper (defaults to empty, so all ~20 other existing call sites are unaffected and unwrapped by `timeout`). The lock-wait test now passes `LOCK_WAIT_HOLD_SECONDS + 60` (90s) as the deadline, and a dedicated check for exit code 124 emits a clear "hung for over Xs instead of waiting out a Ys lock hold" diagnostic before falling through to the existing pass/fail assertions, which are unchanged. No product code was touched. Verified with `bash -n`, `shellcheck -x` (no warnings), a standalone reproduction of the timeout/no-timeout/success paths, the project's `bin/fm-lint.sh --fast` on the file (clean), and the full `tests/fm-lint.test.sh` suite (all 46 assertions pass) * no-mistakes(ci): Replaced the direct `timeout "$deadline_seconds"` call in `spawn_task()` (tests/fm-backend-herdr-presentation-e2e.test.sh) with the repo's portable bounded-execution helper: sourced `bin/fm-timeout-lib.sh` at the top of the file and changed `deadline_cmd=(timeout "$deadline_seconds")` to `deadline_cmd=(fm_run_timed "$deadline_seconds")`. This removes the GNU/BSD `timeout` dependency that would fail with exit 127 on a stock macOS host without coreutils, while preserving identical semantics (exit 124 on bound-hit, command's own exit otherwise), which the existing `[ "$LOCK_WAIT_STATUS" -eq 124 ]` diagnostic check already relies on. Verified: `bash -n` syntax check, `bin/fm-lint.sh --fast` clean, full `tests/fm-lint.test.sh` suite (46/46 pass), and a standalone repro confirming `fm_run_timed` returns 124 on timeout and 0 on success identically to the prior `timeout` call. No other direct `timeout` calls exist in this file or elsewhere in the PR's diff, so no sibling sites remain * fix(herdr): gate exact-resume lock wait behind --herdr-resume-lock-wait Keep refuse-by-default on presentation-order lock contention for Herdr exact resume. Callers that need concurrent recoveries to serialize must pass --herdr-resume-lock-wait; unbounded blocking on a third-party session lock is never the default. Update docs and the real-Herdr e2e suite so the default path asserts the refusal and the opt-in path asserts the wait. * no-mistakes(test): Fix e2e test's lost exit status after if/fi with no else branch * docs(herdr): stop advertising --herdr-resume-lock-wait on --relaunch The relaunch path reuses the recorded endpoint and never takes the presentation-order lock, so the flag is inert there. Drop it from the --relaunch usage line and state where the flag applies. * no-mistakes(review): Clarify lock-wait docs; simplify bash-3.2-safe spawn_task helper * no-mistakes(ci): Fixed ci-1 (Greptile P2). In tests/fm-backend-herdr-presentation-e2e.test.sh, the failure cleanup `cleanup_all` stopped only `LOCK_CONTENTION_OWNER_PID`. It now also stops `LOCK_REFUSE_HOLDER_PID` and `LOCK_WAIT_HOLDER_PID`, the holders of the two new contention cases, so a `fail` before their explicit `wait` no longer leaves them running. Both new PIDs are initialised empty next to the existing one, and each is cleared right after its successful `wait` so cleanup never touches a finished PID. I changed nothing else. `bash -n` passes. The real Herdr e2e run passed both new cases ("default resumed identity refuses session lock contention" and "--herdr-resume-lock-wait waits out session lock contention instead of refusing"). The full run hit my 550s timeout in a later, unrelated case, after the new cases passed
….2 (kunchenguid#6762) * fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 Stock macOS bash 3.2 holds a HUP or TERM until a running command substitution's child exits, and the watcher read every pane through $(fm_backend_capture ...). A blocked backend read therefore held the watcher's stop for as long as the read lasted, and a stopped watcher left the hung read orphaned. tests/fm-watch-triage.test.sh test_term_stops_a_watcher_blocked_inside_a_poll failed on /bin/bash 3.2 for this reason while passing on bash 5. Pane captures now go through watcher_capture, which runs the read as a waited background process group recorded like a check's, so the stop is honored at once and watcher_cleanup stops a read still in flight along with its per-call output file. * no-mistakes(review): Run drain-ring idle capture in watcher shell, add regression test * no-mistakes(document): Document watcher TERM handling for blocked checks and captures * no-mistakes(test): Silence bash 3.2 setpgid race noise from watcher captures * fix(bin): verify the capture group and scope the stop claim to pane reads watcher_capture now confirms its background read leads its own process group, as run_check_capture already does, so watcher_cleanup never relies on a group that set -m failed to create. The comment and continuity doc now say only fm_backend_capture pane reads go through watcher_capture; agent-state and composer-state reads still run inside command substitutions.
* feat(spawn): add per-home worker tool exclusions Add an optional per-home config/crew-exclude-tools file listing tool names to hide from workers, one per line, with blank lines and # comments allowed. It applies to every ship and scout launch and relaunch in that home, is never inherited by another home, and does not affect secondmate agents. Pi and pi-signed apply it through --exclude-tools, which also covers MCP tool names. Any other runtime, and a raw launch command, refuses the launch when the list is non-empty rather than ignoring it. Malformed entries are refused before provisioning, and before a relaunch stops a running worker. Exclusions that match no tool in the worker's loaded registry are reported as unverified warnings in its status record instead of refusing the worker. Closes kunchenguid#6744 * no-mistakes(review): Preserve UTF-8 exclusion paths and verify Pi lifecycle behavior * no-mistakes(document): Clarify worker tool exclusion documentation * no-mistakes(ci): Fixed ci-1 in bin/fm-exclude-tools-lib.sh: a failed read now returns an error before printing names, so all shared launch and relaunch callers refuse rather than silently dropping exclusions. Added deterministic regression coverage for a file disappearing after readability checks across Pi/pi-signed ship and scout launches. Reproduced the original failure; verified 83 spawn checks, 77 relaunch checks, direct parser/runtime failure cases, full targeted lint, Bash syntax, and git diff --check. Relaunch tests passed with existing fixture-cleanup permission warnings. ci-2 remains unchanged per the user's decision; the outer executor owns the fresh CI run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge pinned upstream
mainat47aff866dbe0612bd43df66d8fa76576e06a2b3einto the fork'shouseline with a true merge commit.Fork
mainwas fast-forwarded non-forcibly from99da16d9to that tip; the pinned tip matched the intake tip.House behavior is retained, and the pipeline fixed a merge interaction between house's Treehouse project-lock handling and upstream's new presentation-lock wait.
Land this PR with a merge commit.
Intent
update our house branch with the latest upstream main commits for the firstmate, quota-axi,backpass,no-mistakes forks we run. teport back on changes and conflicts
Context for reading that ask. This task covers the firstmate fork only; backpass is already current; quota-axi and no-mistakes are synced by sibling tasks. The fork jazz127/firstmate runs
house(default branch, merge commits only) as the line we run from; itsmainmirrors upstream kunchenguid/firstmatemain. Forkhouseis at 852c67d and forkmainat 99da16d. Upstreammainis at 47aff86, with 7 upstream commit(s) not yet in house: 2e8cd9e fix(bin): reopen the remote-reply continuity decision on a later break (#6708); e7c3bdc fix(control): drop busy_gen from the task record when an incarnation is retired (#6733); 237e1cf test: cover PID collisions in harness ancestry detection (#6484); 53b5bc1 fix: share Pi Calm's working-ship widget slot (#1854); 23e71b3 feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (#6649); 0f34fab fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (#6762); 47aff86 feat: add per-home worker tool exclusions (#6750).Upstream commits
2e8cd9e9fix(bin): reopen the remote-reply continuity decision on a later break (fix(bin): reopen the remote-reply continuity decision on a later break kunchenguid/firstmate#6708)e7c3bdc1fix(control): drop busy_gen from the task record when an incarnation is retired (fix(control): drop busy_gen from the task record when an incarnation is retired kunchenguid/firstmate#6733)237e1cf3test: cover PID collisions in harness ancestry detection (test: cover PID collisions in harness ancestry detection kunchenguid/firstmate#6484)53b5bc11fix: share Pi Calm's working-ship widget slot (fix: share Pi Calm's working-ship widget slot kunchenguid/firstmate#1854)23e71b3dfeat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn (feat(bin): add opt-in --herdr-resume-lock-wait to fm-spawn kunchenguid/firstmate#6649)0f34fab4fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 (fix(bin): let TERM stop a watcher blocked in a pane capture on bash 3.2 kunchenguid/firstmate#6762)47aff866feat: add per-home worker tool exclusions (feat: add per-home worker tool exclusions kunchenguid/firstmate#6750)Conflicts and resolutions
Five files had textual conflicts:
bin/fm-control.sh: retain both the house Dock and upstream tool-exclusion library imports.bin/fm-spawn.sh: retain house--seat luna|maindocumentation alongside upstream--herdr-resume-lock-wait.bin/fm-watch.sh: adopt upstream's direct capture calls, result variables, waited process groups, and cleanup; preserve the house two-second tmux capture bound through the existing signal-forwardingfm_exec_timedrunner.The first combination using
fm_run_timedleft an orphan in a synthetic stop regression; the signal-forwarding runner resolved that merge-caused failure.docs/watcher-continuity.md: document both process-group stop handling and the retained capture bound, including the existing limitation for other tmux queries.tests/fm-control.test.sh: retain both the house footer-busy Pi refusal and upstream busy-generation retirement tests and invocations.The pipeline also fixed semantic merge issue R2 in
4f25ec61: presentation-lock acquisition now releases the Treehouse project lock before waiting, reacquires it afterwards, and revalidates protected state before provisioning.This preserves opt-in/default/batch behavior and the house slot-ownership checks while letting the presentation owner retake the project lock.
The documentation commit
daacc683updates the lock-order ownership comment and corrects control/recovery documentation.No house behavior was dropped.
Both the original house base and pinned upstream remain ancestors of the pushed head
daacc683a843d2624ef66a3f3c9eed2b27425f98.Validation and accepted limits
Review, documentation, lint, and push completed through no-mistakes.
Test was explicitly approved with this recorded exception: "upstream merge into house; scripted suites are the accepted evidence, live drive not required".
The accepted Test evidence is synthetic/offline; its live verdict is inconclusive, with 0 of 16 scenarios driven live in that run.
The R2 counterfactual observed one failure in one pre-fix case: the presentation owner could not retake the project lock.
All six post-fix cases passed: opt-in wait, batch forwarding, default refusal, fresh creation, project-space creation, and a changed journal.
The other five modes have no pre-fix failure measurement in that sample.
Reproduction artifacts are
accepted-test-evidence/lock-order-before.log,lock-order-after.log, andlock-order-transcript.txtunder the evidence directory below; archived selectors preserve the executable-interface path.Focused synthetic checks cover busy retirement, PID handling, exclusions, watcher cancellation, local reply-decision reopening, and Calm widget behavior.
The full secondmate, reply, and dispatch subjects exceeded their selected 240/300/360-second bounds; their remaining cases are not claimed as passes.
The fake-SSH source selector returned no-result exit 64, so direct local delta-reader/ingest checks provide narrower offline reply evidence.
The stock worktree-settle fixture suppressed process identity and failed its task-set lock; a disposable fixture copy delegated other queries to
/bin/psand passed without changing the tracked fixture.The first relaunch case in
tests/fm-control-relaunch.test.shstill reports endpoint cwdunknown; later cases did not execute.The same failure was reproduced on an exact clean checkout of untouched house
852c67dca8054c3d76a70491d985cabbaf3a2096and left unchanged as instructed.Baseline evidence:
/tmp/fm-hf-firstmate-upstream-sync-1007b/relaunch-house-baseline.logandrelaunch-house-baseline.json.Herdr lifecycle scenarios are not driven live in accepted validation: lock contention/default refusal, opt-in serialization including batch recovery, exact husk replacement, and presentation layout/focus preservation.
Native harness/TUI, actual tool-registry enforcement, and actual rendered UI checks are also excluded.
An earlier canceled Test attempt's runtime evidence is not accepted for this PR; targeted readback found its recorded Herdr session and tmux socket absent.
The driver made no default-session mutations.
Earlier review warning R1 concerned inherited upstream metadata-retirement failure propagation; its helpers are source-identical to pinned upstream and were left unchanged.
That comparison is source-only; no filesystem-failure injection was performed.
House CI:
House checkspassed on the pushed head at the evidence capture below.evidence-artifact: /tmp/fm-hf-firstmate-upstream-sync-1007b/pr-evidence.json
evidence-command: python3 /tmp/fm-hf-firstmate-upstream-sync-1007b/capture-pr-evidence.py
evidence-captured: 2026-10-07T16:38:25Z
Pipeline
Updates from git push no-mistakes