Conversation
14f6321 to
a936a81
Compare
|
Speaking as Kun's firstmate: First look on HEAD Attestation: MISMATCH — body attestation |
…ktree (#3) * fix(bin): refuse teardown when a task's endpoint close fails (kunchenguid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (kunchenguid#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(orca): validate uuid::path worktree identity against recorded worktree * no-mistakes(document): Document Orca composite identity validation * no-mistakes(review): Allow colons in valid Orca worktree paths * test(orca): allow :: in composite worktree path half The review fix removed the broad colon rejection so valid Unix paths containing ':' or '::' are accepted when they match the recorded worktree. Update the refusal case to a path half that does not match worktree=, which is the malformed shape the guard must still reject. * no-mistakes(test): Updated stale Orca fixture; targeted teardown test passes --------- Co-authored-by: Amin Roudaki <roudaky@gmail.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: mdc2122 <mdc2122@users.noreply.github.com>
The OMP Calm extension lived untracked at .omp/extensions/fm-calm-omp.ts, so it loaded only in sessions whose cwd was the firstmate home. Move it to extensions/fm-calm-omp/ as an omp plugin package and add bin/fm-omp-calm-install.sh, which links it into OMP's user plugin scope so /calm-omp and the working boat load in every omp session. The package stays outside .omp/extensions/ because OMP de-duplicates extension entries by absolute path, not realpath: a project-local copy plus a global link would load twice in home sessions. The installer retires the legacy copy, removing an identical file and renaming a divergent one to .bak.
… regression tests
Vendor working-ship into extensions/fm-calm-omp/lib, copy to ~/.local/share/fm-calm-omp on install, and link via omp plugin link so /calm-omp works in every session without firstmate checkout dependencies. Co-authored-by: Cursor <cursoragent@cursor.com>
…ktree (#3) * fix(bin): refuse teardown when a task's endpoint close fails (kunchenguid#4510) * fix(teardown): refuse a cleanup whose endpoint close failed bin/fm-teardown.sh discarded both the exit status and the stderr of every fm_backend_kill call, so a close that genuinely failed was indistinguishable from one that succeeded. Teardown continued past it, deleted the task's durable records, returned its worktree, and reported the cleanup as completed. The deleted metadata is the only record of which endpoint belongs to the task, so such a close did not merely leave a stray session behind, it stranded one: nothing was left on disk naming it. The adapters could not carry that signal either. Driven against the real code, every backend arm returned 0 for a genuine failure exactly as it did for an already-exited endpoint, so there was nothing for the four call sites to propagate even once they stopped swallowing it. The tmux arm now resolves a close that did not succeed against the window's exact recorded identity, since kill-window fails the same way for a window that is gone and one that is still there. The Orca arm reports a close its missing CLI never attempted. Both stay silent for an endpoint that is already legitimately gone, and the remaining arms are unchanged: their close-command timing cannot be established without the real Zellij, Orca, and cmux binaries, and a gate that refused ordinary cleanup of an already-exited session would be worse than the defect. docs/verification/runtime-backends.md records what each backend can prove. A reported close failure now reaches teardown's existing retain-and-stop refusal before the records naming the endpoint are removed, matching where the Herdr confirmed-gone gates already sit for the same hazard, and the retained records let a rerun finish once the close works. * no-mistakes(review): refuse unreadable tmux close re-read; honor --force override * no-mistakes(review): drop unreachable Orca force arm; prove CLI-absent close * no-mistakes(document): document endpoint-close refusal in its backend and retirement owners * no-mistakes(ci): The two reported failing checks are NOT code defects. Both "CI" (run 34935529184) and "Require no-mistakes" (run 34935529206) returned conclusion=action_required with zero jobs and 0s duration (run_started_at == updated_at), which is this repo's workflow-approval gate holding the run before any job starts. No job executed, so nothing in the diff could have caused them; two unrelated branches (fm/captain-hold-json-nonref, fm/presenter-core-l1) show the identical shape in the same time window. Verified the change locally instead: bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok, and all suites the diff touches pass (fm-teardown-endpoint-safety 25/25 including the five new endpoint-close cases, fm-backend-orca, fm-backend, fm-backend-tmux-smoke, fm-backend-cmux, fm-backend-zellij, fm-backend-herdr). Separately, I found and fixed a genuinely flaky test that the phase rules require me to make deterministic: tests/fm-tmux-agent-liveness.test.sh intermittently failed "an idle shell pane must classify dead" (verdict ambiguous, comms=[bash sleep]). It is selected by --changed for this diff, so it would run against this PR once CI is approved. Root cause, established by instrumenting the pane's process group: the idle window was created by `new-session` with no command, so it inherited tmux's default-shell, i.e. whoever runs the suite. ps on the pane tty showed `-zsh` -> `bash` -> `sleep`, all sharing pgid==tpgid, i.e. the host operator's shell configuration spawning a periodic helper directly into the pane's FOREGROUND process group, which is the one surface the classifier reads. `sleep` classifies as `other`, so fg_other=1 and the verdict became `ambiguous` instead of `dead` whenever that helper overlapped the 10s poll window. Every other window in the suite runs an explicit command via new_window; the idle case was the only one whose process group the host defined. Fix (smallest root-cause, test-only, 1 line + explanatory comment): create the idle window with an explicit bare `/bin/sh` (`-- /bin/sh`), the same shell the neighbouring background case already execs. Its foreground group is now exactly one process (verified: `/bin/sh` alone), so no host configuration can inject into it. This flake is pre-existing and NOT caused by this PR: an interleaved A/B showed base commit da5e658 failing the identical case (2/6 runs) alongside head (3/7 runs), and the diff only extracted the tmux inventory read into a helper with identical semantics while never touching fm_backend_tmux_foreground_comms. After the fix: 8/8 consecutive passes, with lint and the coverage guard still clean. Change left uncommitted in the working tree * feat(calm): add flag-gated Claude Code Calm mode (kunchenguid#4565) * feat(calm): ship the Claude Code Calm and sailboat mod behind the function-hooks flag Add .claude/mods/firstmate-calm, a Claude Code mod (function-hooks plugin) that brings Calm to Claude Code: the sailboat replaces the stock working row through a Raster repainted on the sprite's own tick, and tool, tool-group, mid-turn narration, and canonically classified operational user rows draw at zero height. /calm is registered by the hooks module itself and toggles the same per-home config/calm preference the Pi extension uses, so one choice applies on either harness; rows redraw retroactively on toggle and stay hidden across claude --continue. The mod loads only while Claude Code's default-off CLAUDE_CODE_ENABLE_FUNCTION_HOOKS flag is on. Nothing sets that flag in any settings file, and the plugin carries no command file, skill, agent, or classic hook, so it is a complete no-op while the flag is off. The trusted project auto-loads it through an .agents/skills symlink, the only path Claude Code scans for project plugins. Extract the working-ship geometry, bounce track, cadences, and freeze/resume state into a harness-neutral sprite core inside the mod (Claude Code refuses hooks-module imports from outside the plugin folder) and have the Pi widget paint that core's frames as standard ANSI, byte for byte as before; the Pi suite stays green. Classify operational rows through a port of bin/fm-operational-input.sh's classify command guarded by a corpus parity test against the shell owner. Tests: portable Node checks (plugin shape, sprite parity with Pi's rendering, Raster packing, policy, classifier parity), the mod's own claude plugin test suites behind a default-on wrapper, and an opt-in live TUI guard proving the flag-off no-op, the moving boat, hidden rows, the persisted toggle, and resume on Claude Code 2.1.272. Docs: record the version-scoped Claude Code evidence and the three bounded gaps in docs/calm-mode-feasibility.md, describe the Claude Code contract in docs/calm.md, and make the shared preference, layout, and contributor notes harness-neutral. * no-mistakes(review): Preserve colliding final replies and strengthen parser parity * no-mistakes(review): Preserve final replies and strengthen canonical parity checks * no-mistakes(review): Require exact function-hooks opt-in before Calm activation * no-mistakes(review): Clarify Calm module loading and activation boundaries * no-mistakes(review): Reset Calm presentation state across session starts * no-mistakes(document): Refresh Calm session lifecycle documentation * feat(calm): paint the Claude Code working ship in Claude's own theme colors The captain picked the "Claude native" palette for the Claude Code mod's Raster: every water cell takes the spinner blue of the active theme family (#93a5ff dark, (#d77757), one water color and one boat color. The family follows the `theme` setting's prefix, read at load through $.config.list and re-read on a config.set of that row, with `auto` and custom themes falling back to the dark set. The Pi extension keeps its standard ANSI blue and yellow, byte for byte. Rename the shared sprite's color classes from hue names to `water` and `boat`, since each harness now maps them to its own colors; geometry, motion, cadence, and the activation gate are untouched. Tests cover both palettes' packing and the family rule under Node, and the plugin kit drives every theme value, a theme change mid-session, the Calm-off pass-through, and inertness of the menu read while the flag is off. The docs describe the Claude Code colors and record the guard passing on 2.1.273. * no-mistakes(review): Use light palette for unresolved Claude themes * no-mistakes(document): Refresh Claude Calm verification evidence * fix(orca): validate uuid::path worktree identity against recorded worktree * no-mistakes(document): Document Orca composite identity validation * no-mistakes(review): Allow colons in valid Orca worktree paths * test(orca): allow :: in composite worktree path half The review fix removed the broad colon rejection so valid Unix paths containing ':' or '::' are accepted when they match the recorded worktree. Update the refusal case to a path half that does not match worktree=, which is the malformed shape the guard must still reject. * no-mistakes(test): Updated stale Orca fixture; targeted teardown test passes --------- Co-authored-by: Amin Roudaki <roudaky@gmail.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: mdc2122 <mdc2122@users.noreply.github.com>
omp 18.1.22 changed from a Bun-compiled single binary (comm=omp) to a bun script (comm=bun, argv 'bun .../.bun/bin/omp'), which broke harness detection: the session lock, agent-process classifier, tmux backend liveness, and fm-harness verdict all missed it. Add fm_omp_args_are_omp() in fm-session-lock-lib.sh as the single owner of the bun-launcher rule (argv[0] is bun, next arg is a path ending in /omp) and consult it from the ancestry matcher, the agent-process classifier, the tmux foreground-args liveness check, and the fm-harness verdict. Anchored exactly like the binary shape so ompd/comp and bun processes carrying other harness names in their arguments never match.
Upstream 616049a (kunchenguid#4498) redesigned both geometry and palette; 0f242b9 (kunchenguid#4554) and 3a82299 reverted only the palette, leaving the Unicode geometry. Restore the original Jul 30 sprite from 621299a: ASCII <| / |> sail, \__/ hull, ~-cycle water — one yellow boat over blue water.
Apple Terminal's default palette renders ANSI 34 as a dark violet; its cyan (36) is the hue this water was always meant to read as.
* fix(fm-send): gate the post-interrupt composer clear on restored-prompt proof muse restores the cancelled prompt into its composer after Escape, but only when the composer was empty at cancel time - fresh typed input survives the interrupt untouched. fm_send_clear_after_interrupt sent C-u unconditionally, clobbering a message the captain was typing when a wake arrived mid-composition. The clear now fires only when the composer's extracted content is a suffix of the cancelled run's recorded started prompt (the provable restored text, read from the bound muse session log). A mismatch, an unreadable composer, or an unprovable restored prompt skips the clear with a warning instead. Supporting changes: - fm_backend_composer_content dispatcher plus tmux/herdr/cmux adapters (zellij already had one) expose normalized composer content. - fm_busy_muse_last_run_prompt extracts the last run's started prompt; fm_busy_muse_matching_logs now scans the bounded log prefix for the metadata record because muse 1.3.0 prepends a retained_frame wrapper. - _fm_composer_select_cursorless exempts a separator directly adjacent below a bare composer: muse 1.3.0 closes its composer with a pure rule that the lone-separator veto misread as a dangling pi edge. * no-mistakes(review): warn on unreadable composer in post-interrupt clear gate * no-mistakes(test): accept muse 1.3.0 glyph in live e2e matcher * no-mistakes(document): Clarify proof-gated Muse composer clearing * fix(fm-control): proof-gate the post-interrupt composer clear too The same clobber fm-send's clear had is reachable through the control plane: send_interrupt_keys sent C-u unconditionally after Escape for muse, so an interrupt or busy exit could eat input the captain was typing. The clear now runs through fm_busy_muse_restored_prompt_verdict (bin/fm-busy-lib.sh), the single owner of the restored-prompt proof shared with fm-send's --key Escape path. The control plane's contract differs from fm-send's: a composer that cannot be proven free of the restored prompt dies loudly rather than skipping, because the next lifecycle line would concatenate onto whatever remains; only provably-fresh input skips the clear, with a warning, since there is then nothing restored to remove. fm-send's clear is refactored onto the same helper so both planes share one proof implementation. * no-mistakes(review): Require stable composer before restored-prompt clearing * no-mistakes(review): Reject mixed readable and failed composer samples * no-mistakes(review): Normalize composer whitespace before restored prompt comparison * no-mistakes(test): Fix fm-control stable composer test fixture * no-mistakes(test): Fix control composer stability test fixtures * no-mistakes(document): Clarify conditional Muse composer clearing * fix(tests): bind muse send cases to a session log for the proof-gated clear The post-interrupt composer clear is now gated on fm_busy_muse_restored_prompt_verdict, which needs a muse session binding and a readable pane. Extend the send-case fake tmux with capture-pane and cursor_y, add muse_session_fixture to bind each muse case to a session log carrying the restored prompt, and wire it into the alias and failed-clear tests. Also quote proof=agent-alive in fm-control.sh (SC2100). * no-mistakes(review): Normalize multiline restored prompt boundaries consistently * no-mistakes(document): Document proof-gated interrupt clear and restore-wait knobs * no-mistakes(test): pin terminal color env in live muse e2e launch * no-mistakes(document): fix stale interrupt-clear wording in fm-send comment * fix(bin): guard rm -rf against empty INSTALL_DIR (SC2115) Pre-existing fork-main lint debt flagged by the PR merge-checkout lint job. The :? expansion aborts before rm -rf /* if the var is ever empty. * no-mistakes(review): Captain: removed destructive installer cleanup path * no-mistakes(test): Fix Muse multiline composer clearing * no-mistakes(document): Classify new launcher incident verification document * fix(tests): expect C-c for the muse post-interrupt composer clear The pipeline live-verified that muse 1.3 Ctrl+U clears only the current line, so the clear key moved to Ctrl+C (commit 5bfda54), but these assertions still expected C-u and failed the serial-4 job. * no-mistakes(document): Updated Orca identity and interrupt documentation --------- Co-authored-by: mdc2122 <mdc2122@users.noreply.github.com>
* fix(bin): report verified PR state for passed runs (kunchenguid#4624) * fix(bin): derive passed PR state from PR record A completed no-mistakes run with outcome=passed does not prove the associated pull request merged or closed. A parked gate can be approved on other evidence, so the old crew-state label could report an open PR as merged and make teardown look safe when unlanded work still exists. For passed runs, derive the crew-state detail from the run or task PR identity, accept a matching merge-poll retirement receipt as local merged evidence, and otherwise perform a bounded forge read. If the identity is absent or unreadable, report the run as passed with unknown PR state instead of inventing a merged claim. Fixes kunchenguid#4607 * no-mistakes(review): Add bounded GitLab merge-request state reads * no-mistakes(review): Preserve network-free inactive crew-state scans * no-mistakes(document): Document PR record readers in shared library * fix: restore published contribution follow-up (Fixes kunchenguid#4469) (kunchenguid#4627) * fix: restore published contribution follow-up (Fixes kunchenguid#4469) * fix(review): Fix contribution freshness and merge actor routing * fix(review): Restore issue triage and scope contribution follow-up * fix(test): test: assert one wake per contribution signal * fix(document): Document contribution follow-up * fix: restore truthful terminal delivery evidence * fix(review): Disclose unsupported contributions and deduplicate watcher wakes * fix(review): Preserve unmeasured unsupported contributions across Bearings * fix(review): Deduplicate shared contribution wakes and isolate diagnostics * fix(ci): Captain, fixed the CI failure by updating the PR-security fake GitHub interface to support the contribution observer’s API reads. Verified with shellcheck, git diff --check, the full contribution suite, and a focused merged-poll retirement reproduction. The full PR-security script was not allowed to complete locally after its expanded observer path made it substantially slower * fix(bin): make remote report transfers explicit and fail-open (kunchenguid#4658) * fix(bin): make a remote-reply document gap self-clearing and re-attemptable A remote mate's undelivered document raised a keyed `blocked` decision that nothing could ever resolve, and any `data/*.md` substring in any mirrored line was an unconditional fetch instruction. A mate announcing a report it had not written yet therefore manufactured a permanent, factually false blocker, and its own explanation of the false alarm manufactured more. The reader has no permanence vocabulary: a report still being written refuses exactly like a path that will never exist. So an undelivered document is now a durable, re-attemptable obligation under `state/remote-replies/<id>.pending-docs`, re-attempted on the next delta and on the channel's own quiet poll, and retired with a matching `resolved` line naming the local copy once it arrives. The cursor still advances and no delta stalls on one bad pointer. Only a structured `report=data/....md` pointer now offers a document, so a path merely mentioned in prose - including one under another home's mirror tree, which is provably not that mate's to serve - is never fetched. Offers are deduplicated across the whole delta, the escalation names each missing document once and carries the reader's own reason instead of discarding it, and a strictly increasing notice ordinal keeps a later escalation from being swallowed as duplicate bytes. A mirrored line still lands once whichever pointer form it was first written under. * no-mistakes(review): Require structured pointer token boundaries * no-mistakes(review): Unify boundary-safe pointer extraction and rewriting * fix(bin): identify a mirrored line independently of its delivery state Two defects in the boundary-safe pointer work. The at-most-once check compared only the all-remote and all-local renderings of a line, so it could not recognize a mixed one. A line offering two documents where only the first was deliverable mirrored as local-plus-remote; once the second arrived, a cursor-loss whole-log recapture rendered the same line all-local, matched neither alternate, and mirrored a second time. A line's identity is now the canonical form every boundary-valid pointer would take once delivered, derived by the same parser that does extraction and rewriting, so it no longer depends on which documents happened to be deliverable at the time. The pointer map was passed to awk through the process environment. A delta may carry up to the configured 1 MiB bound, and an expanded map of delivered pointers can exceed the platform's exec argument limit, so awk would fail to start; because no caller checked, the empty result would have been appended as blank lines while the cursor advanced past dropped status content. The map now travels in a file, and every call site checks the exit status and stops the ingest rather than committing a delta it could not render. Both passes now run once per stream instead of twice per line. * no-mistakes(review): Abort ingest when document pointer extraction fails * no-mistakes(review): Exclude structured cross-home pointers from document transfer * fix(bin): fail open on an undeliverable remote document instead of tracking it Narrow the remote-reply document fix to the scope the diagnosis actually requires, as decided after measuring a simpler alternative. A document the reader cannot deliver now fails open. The mate's line is mirrored with its own pointer, the cursor advances, and one unkeyed note carries the reader's reason. A note never enters the open-decision fold, so it cannot stand open the way the original keyed block did - which removes the never-clearing false blocker by construction rather than by resolving it. That makes the durable self-clearing obligation unnecessary, so it goes: the per-mate pending-documents record, its notice ordinal and resolved announcements, and the poll-side retry. Canonical line identity goes too, and with it a way to silently drop a genuine status line; mirroring is back to at-most-once on exact bytes. The cross-home exclusion goes as well: under fail-open a cross-home report= either fails harmlessly or is a nested remote report this mate genuinely holds, which is now relayed again. Kept: fetching only on a structured report= pointer, the boundary-correct parser, the file-based rewrite map, and checked extraction and rewrite exit status. The parser now scans behind a sentinel byte so a rejected candidate can no longer give the text right after it a false leading boundary. The reported incident is covered end to end: a report path announced in prose before it exists raises no decision, and the report still arrives through the ledger publisher's structured offer once written. * no-mistakes(review): Preserve source-line identity across remote reply replays * no-mistakes(document): Document remote reply transfer and replay semantics * no-mistakes(lint): Fix staging truncation lint checks * fix(calm): preserve substantive mid-turn responses (kunchenguid#4655) * Preserve substantive Calm mid-turn text * no-mistakes(review): Distinguish newline-preserved replies from short narration * no-mistakes(document): Document Calm mid-turn preservation boundaries * no-mistakes(ci): Fixed the flaky contribution watcher test by increasing its bounded checkpoint from 5 to 15 seconds, allowing diagnostics to surface under slower CI load. Verified with `bash tests/fm-contributions.test.sh` and `git diff --check` * feat(bin): merge a still-open green PR from its armed poll under yolo=on When a task's merge poll finds the PR still open and the task meta records yolo=on, the watcher now invokes bin/fm-pr-merge.sh itself instead of only reporting the still-open PR and waiting for a firstmate turn. The merge path's own live verification remains the only green check: a red, unmergeable, or unverifiable PR is refused there and the poll stays armed. A confirmed landing is recognized by the merge-notified marker the merge path commits with its durable outcome, so a queued or unconfirmed acceptance keeps polling and the landed wake is emitted exactly once. The task's recorded yolo posture now resolves as standing merge authority ahead of the away-posture record, which is consulted only for a task without it. * no-mistakes(review): bound the yolo merge attempt and identity-check its refreshed poll snapshot * no-mistakes(test): Bound watcher tests by terminating descendant process groups --------- Co-authored-by: Joseph Kim <jokim1@gmail.com> Co-authored-by: Mickaël Rémond <mremond@process-one.net> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: mdc2122 <mdc2122@users.noreply.github.com>
…oviders pi's --list-models hides every provider with no configured auth, including extension-registered providers whose models --model still resolves. The stale enabledModels pin 'muse-code/muse-spark-1.3' in ~/.pi/agent/settings.json (synced from omp's config, where muse-code is a built-in provider) prints 'No models match pattern' on every pi startup but does not fail the launch; the no-mistakes fix agent runs on agent_config.pi.model=zai-glm53-cliproxy/zai-glm53-max and the observed deaths were 30m agent_timeout SIGTERMs, not model failures. Correct the adapter reference and qualify dispatch-auth.md's authoritative-negative claim so an empty listing for an unauthenticated provider is treated as missing evidence rather than a block.
a936a81 to
33abf4b
Compare
|
Closing per triage feedback (polluted tip); will re-raise the narrow docs fix from a clean branch if still needed. |
Intent
Fix the no-mistakes fix agent's dead model seat. Every run that needs a fix agent dies: pi exited: exit status 143: Warning: No models match pattern "muse-code/muse-spark-1.3" — observed twice on yolo-automerge (run 01M2NT51H43R2GEKDWCHMKJP3T) and twice on remote-omp-verify-07 (runs 01M2PNAX9M7AE27XEM2TF8ZJ12, 01M2PVMZG0PMQ9772YNNMD8QBF). The pin muse-code/muse-spark-1.3 exists in omp's model list but pi rejects it, and the fix agent runs under pi.
What Changed
Risk Assessment
🚨 High: The changed interrupt handling can still destroy captain-authored input on a reachable normal sequence, directly undermining the fix's safety invariant and affecting both control paths.
Testing
Focused live Pi validation passed for nonfatal stale-pin startup and omission output. The authenticated model path central to the fix-agent intent could not be exercised without provider credentials/configuration.
Evidence: Pi launch with stale enabledModels pin
Source: Pi launch with stale enabledModels pin
RC=0 Warning: No models match pattern "muse-code/muse-spark-1.3" Captain, aye — on deck and ready.Evidence: Pi model listing omission
Source: Pi model listing omission
RC=0 No models matching "muse-code/muse-spark-1.3"Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
.agents/skills/bearings/SKILL.md- branch carries 18 commit(s) that exist on your local main branch but were never pushed to origin/main; these may be unintended bundled work (proposed PR changes 80 file(s)):Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.
✅ No issues found.
bin/fm-busy-lib.sh:744- The restore proof is not actually unambiguous: it classifies any stable composer text that is a suffix of the cancelled prompt asrestored(case "$prompt" in *"$content")). A real sequence is: the cancelled prompt isplease run tests, Muse restores nothing because the composer was non-empty or the prompt was already cleared, the captain then types fresh inputtests, andfm-sendorfm-controlreads that stable text; it matches the suffix and sendsC-c, deleting the captain's new input. This violates the change's stated no-clobber invariant on both callers atbin/fm-send.sh:318andbin/fm-control.sh:395. Use an unambiguous restore signal or require exact restored-content equality; the current suffix heuristic cannot distinguish restored text from valid fresh suffix input.git show --stat --oneline a936a81506a28e5e83275ab5a44dec6e661a9e1agit diff --check b430bf50d9aea5d1a810cab1bd293670bbca64be..a936a81506a28e5e83275ab5a44dec6e661a9e1aLive validation:⚠️ inconclusive - 2 of 3 scenarios driven live against the product
PI_CODING_AGENT_DIR=<isolated dir> pi --offline --no-session --no-tools -p 'ping'with staleenabledModelspinPI_CODING_AGENT_DIR=<isolated dir> pi --offline --list-models 'muse-code/muse-spark-1.3'bash tests/fm-harness-adapter-references.test.sh✅ **Document** - passed
✅ No issues found.
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
✅ No issues found.