diff --git a/.agents/skills/bearings/SKILL.md b/.agents/skills/bearings/SKILL.md index 6a5c6b81d6a..a9256d73b28 100644 --- a/.agents/skills/bearings/SKILL.md +++ b/.agents/skills/bearings/SKILL.md @@ -89,11 +89,13 @@ Board answers are acted on later under the normal authority rules; this skill's ## Lavish board mode `/bearings lavish` adds one deliverable beside the unchanged chat digest: the interactive fleet board, a myfirstmate-styled Lavish page where the captain answers Captain's Call items directly instead of replying in chat. -`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, Lavish session establishment, the any-origin answer binding, and arm-if-absent registration - so the per-invocation work is composing the payload and running its `build`. +`bin/fm-bearings-board.sh` owns every board mechanic - the stable board path, fm-bearings-board.v1 payload validation, template injection, live Lavish session verification and ended-session reopening, the any-origin answer binding, and listener registration - so the per-invocation work is composing the payload and running its `build`. Compose the payload from the same snapshot with the same ranking judgment as the chat digest, plus these board rules: - A Captain's Call decision key is the captain-held TASK ID from `decisions_open` (legacy `-decision-` rows are already task ids); a merge card's key is `merge.`; the Charted Next dispatch picker's key is `dispatch.charted`. +- Before carding a hold, check that its SUBJECT has not already landed, and omit it when it has. `build` drops a card whose task or PR appears in the payload's own landed rows, and one whose task is no longer an open captain call. When a hold waits on one specific PR, put that PR in the card's `pr_url`. When it concerns a published version, put the artifact and numeric three-part version in the card's structured `subject`; landed rows for releases carry the same identity, and a matching or newer version drops the card. Identity matching is structured only, so verify any subject without one of these identities against current reality before carding it. +- Never author a `reconcile` option on any card. `build` gives every decision card the standard reconcile choice itself, and the payload validator reserves that value across all card types; recommendations must name an authored option. - Compose exactly one decision card per captain-held task id. When one task carries multiple questions, consolidate all of them and their options into that card; never emit duplicate cards with the same task-id key. - Decision cards carry agent-authored copy: a short noun-phrase title, one-line `about` and `decide` context rows, and option labels with hints, with the recommended option marked. - Card `type` (decision, merge, credential) is your composing judgment from the row's content; no backlog field types a card for you. @@ -103,14 +105,20 @@ Compose the payload from the same snapshot with the same ranking judgment as the - Every Captain's Call item and every Underway, Recently Landed, and Charted Next row carries an explicit `repo` field. Fill it from the snapshot and task records wherever known; use null or an empty string only as the deliberate genuinely-no-repo marker, in which case the template may show the internal id. Ids otherwise stay in the payload only as the routing channel, and composed reasons name blockers in plain words. Run `build` once after composing the payload. -Its serve-first sequence publishes the board, establishes or resumes its Lavish session with `lavish-axi`, and only then binds and arms the polling source; use the session URL it prints in the chat digest. -Never bind or arm the board before that session exists. -Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and the watcher's ordinary reconcile restarts it, so no conversational turn ever blocks on the board. +Its serve-first sequence publishes the board, establishes and verifies its Lavish session with `lavish-axi`, reopens an ended session when necessary, and only then binds the answer source and proves a live polling listener; use the session URL it prints in the chat digest. +Never bind or arm the board before its session is listed open. +Never run `lavish-axi poll` for the board yourself: the armed source's supervised runner owns the blocking poll, and both the build and the watcher's ordinary reconcile repair a missing listener, so no conversational turn ever blocks on the board. ### Handling a board wake A board answer arrives as an ordinary `procevent lavish ` check wake. Identify it by comparing the wake source id with `bin/fm-procevent-lavish.sh source-id "$(bin/fm-bearings-board.sh path)"`, regardless of which answer kinds the result contains; then load `process-event-sources` and follow its contract for the result read, adapter classification, and the handled acknowledgement. Decision answers need no routing from you: the runner feeds the board's binding into `bin/fm-captain-hold.sh`'s one keyed-answer intake, which closes or releases each answered captain-held task at answer time; reconcile any `skipped:` key yourself with a direct `answer`, and when the captain's answer is "later", record it as a deferral with `bin/fm-captain-hold.sh hold --reason "" --until ` instead of a closure. +A current structured Reconcile selection closes nothing: the versioned board context carries its exact selected option separately from any typed note, and the adapter routes that selection only into a durable re-check request while preserving the note as provenance. +The rollout-compatible old context still feeds ordinary non-reconcile answers, but its bare or separator-annotated reconcile values and every structurally uncertain choice feed neither intake and remain announced for deliberate handling. +Verify the call's latest state, then retire the request through `bin/fm-captain-hold.sh reconcile close --evidence-file ` when it turns out to be moot, or `reconcile note --note-file ` when it is genuinely still open. +Both outcomes refuse without that pending board-created request, and `bin/fm-captain-hold.sh reconcile list` names every request still outstanding. +A remote-secondmate card whose task is absent from the main backlog remains on the board unchanged, but its reconcile request is refused in the main home until the separately tracked owner-aware routing follow-up can query and mutate the authoritative secondmate home; handle the announced capture without claiming that a request or reconciliation succeeded. +`captain-hold-lifecycle` owns why a reconcile may never be recorded as the captain's answer. Route the non-decision keys yourself: - `merge.` is the captain's explicit merge order; follow the merge ruling below. diff --git a/.agents/skills/bearings/assets/board-template.html b/.agents/skills/bearings/assets/board-template.html index 3ba3356dc9e..614bef7426b 100644 --- a/.agents/skills/bearings/assets/board-template.html +++ b/.agents/skills/bearings/assets/board-template.html @@ -548,22 +548,26 @@ var fd = new FormData(form); var value = fd.get("answer"); var note = (fd.get("note") || "").trim(); - /* picked option, optionally annotated; a bare note is itself the answer */ - var answer = value ? (note ? value + " - " + note : value) : note; - if (!answer) return; - if (utf8ByteLength(answer) > 512) { + var displayAnswer = value ? (note ? value + " - " + note : value) : note; + if (!displayAnswer) return; + if (utf8ByteLength(displayAnswer) > 512) { answerLimit.textContent = "Answer is too long to queue (512 bytes maximum)."; answerLimit.classList.add("is-visible"); return; } if (window.lavish && window.lavish.queuePrompt) { - /* close carries the composer-declared close mode: "release" frees a - captain-gated work item instead of completing a question task */ - var ctxData = { question: item.key, answer: answer }; + /* The versioned context keeps the selected option separate from its + note, while close carries the composer-declared completion mode. */ + var ctxData = { + schema: "fm-bearings-answer.v1", + question: item.key, + selection: value || "", + note: note + }; if (item.close) ctxData.close = item.close; window.lavish.queuePrompt( - "Captain's Call answer - " + item.title + ": " + answer, - { tag: "choice", text: item.title + " -> " + answer, element: form, + "Captain's Call answer - " + item.title + ": " + displayAnswer, + { tag: "choice", text: item.title + " -> " + displayAnswer, element: form, data: ctxData } ); } diff --git a/.agents/skills/captain-hold-lifecycle/SKILL.md b/.agents/skills/captain-hold-lifecycle/SKILL.md index 424e16daf3b..bfdfb0148c9 100644 --- a/.agents/skills/captain-hold-lifecycle/SKILL.md +++ b/.agents/skills/captain-hold-lifecycle/SKILL.md @@ -24,7 +24,8 @@ After inventorying the whole report and review surface, run `bin/fm-captain-hold A completed investigation and an ended visual review use this same owner and completion command; a visual tool, including Lavish, never owns a parallel completion policy. Run the command in the originating work's authoritative `FM_HOME`; secondmate-owned work registers in that secondmate home's backlog, and a question already held anywhere is never re-registered as a second row. Do not close a captain-held task merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down. -Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call and only `answer` closes it. +Holding the work item the question gates is safe for exactly that reason: cleanup keeps such a row open with the finished work's deliverable recorded and returns it to the queue, so it still reads as the captain's own call. +Only `answer` with the captain's words or an evidence-backed `reconcile close` may close it. Never close anything the captain owns without recording what he actually said: `bin/fm-captain-hold.sh answer` writes his exact words into the task and closes it in the same act, with `--release` when the answer frees a captain-gated work item to proceed instead of completing a question. When the answer changes what a task must build, follow `AGENTS.md` section 7's Validate contract to preserve the captain's words in the brief and steer the worker. @@ -32,6 +33,15 @@ When the captain says "later", that is an answer too: re-hold with `bin/fm-capta "A keyed answer closes its matching captain-held task" is one capability with one owner, `bin/fm-captain-hold.sh answers`, and every channel that carries a captain answer feeds it the same task id and answer; a channel never maps keys to tasks, records a decision, or closes anything itself. Chat already feeds it through `bin/fm-send.sh --resolve-key`, and a captured-answer source feeds it once bound with `bin/fm-captain-hold.sh bind `; bind before arming the source, and key each structured question by the held task's id. An unbound source and a key that names no captain-held task both simply feed nothing: the answer is still captured and firstmate is still woken, and closing falls back to the direct command above. +One answer value is reserved and closes nothing: `reconcile` means "go re-check reality", never "the captain answered", so the shared intake refuses it from every channel and creates nothing. +A bound captured source uses a separate seam: its adapter omits reconcile from keyed answers and emits the selected task id through `reconciles`, the generic runner feeds that into `reconcile-requests`, and the intake verifies the source binding and the local captain-held task before filing the durable board request. +A remote-secondmate card whose task is absent from the main backlog therefore remains announced but cannot create a main-home request; owner-aware request and mutation routing to the authoritative secondmate home is a separate follow-up. +That board-created request is yours to work off in the turn that receives it: `bin/fm-captain-hold.sh reconcile close --evidence-file ` records the EVIDENCE and closes a moot call, while `reconcile note --note-file ` annotates a genuinely active call and leaves it held. +Both outcomes refuse unless that task still has the pending request created by the captain's board selection, so neither is a standalone way to mutate a captain call. +A normal captain answer also retires any pending request because the call is settled, including close, release, and idempotent replay paths. +A retirement failure makes the command fail without reversing the already-durable answer, close, or note, and `reconcile list` keeps the surviving request visible for retry. +`reconcile list` names every request still outstanding. +Never use `answer` for an evidence-only moot call: `answer` records what the captain said, while `reconcile close` records verified evidence. A captain-held task closed outside this owner leaves no durable answer, so the completion gate keeps failing until `answer` records the decision the captain actually gave. Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create held tasks. Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose. @@ -49,8 +59,8 @@ The absence of a routed work item is not a divergence and the guard never requir 3. Hold that task - or create one captain-held task for the review's open questions - with a concise reason carrying the question and options. 4. Run `complete` with the full captain-held inventory for that review pass. 5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat. -6. Close each call only through `answer` (or a channel that feeds `answers`), through `--until` when the captain defers it, or confirm a channel already closed it. -7. Confirm Bearings reflects the outcome: answered calls leave Captain's Call, released work resumes, and deferred calls sit in Charted Next with their date. +6. Close each call only through `answer` (or a channel that feeds `answers`), close a board-requested moot call through evidence-backed `reconcile close`, record a still-active reconciliation through `reconcile note`, use `--until` when the captain defers it, or confirm a channel already closed it. +7. Confirm Bearings reflects the outcome: answered or reconciled-moot calls leave Captain's Call, released work resumes, active reconciliations remain held, and deferred calls sit in Charted Next with their date. `bin/fm-captain-hold.sh --help` owns command syntax, close modes, legacy-identity compatibility, completion attestation, retry behavior, and close ordering. `docs/captain-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy. diff --git a/.agents/skills/firstmate-coding-guidelines/SKILL.md b/.agents/skills/firstmate-coding-guidelines/SKILL.md index a9e21543077..3f9fe85937e 100644 --- a/.agents/skills/firstmate-coding-guidelines/SKILL.md +++ b/.agents/skills/firstmate-coding-guidelines/SKILL.md @@ -98,9 +98,10 @@ Every such check needs two tests, because they fail for different reasons: - A portable regression in `tests/` that pins the logic with real processes and no harness, so CI enforces the classifier everywhere it runs tmux. Drive the signals apart deliberately and assert the verdict survives losing one; assert the divergence itself so the case cannot go quietly vacuous. Confirm which signal a given construction actually blinds on each supported platform rather than assuming, because the same trick can break different sources on macOS and Linux. -- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`), env-gated and self-skipping, that exercises every INSTALLED harness for real and fails naming the harness and version. +- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`) that exercises every INSTALLED harness for real and fails naming the harness and version. Report an absent harness explicitly rather than passing silently over it, and refuse a pass that checked nothing. - This guard is opt-in and on-demand because standard CI has neither harness binaries nor credentials; run it after every harness upgrade and before trusting refreshed per-harness evidence. + Open it with `fm_live_gate` from `tests/lib.sh`, which is the single owner of that decision: a guard that spends no model tokens runs by default wherever its tools are installed, a guard that submits prompts stays opt-in, and its own variable or `FM_LIVE` forces it on (an absent tool then fails rather than skips) or off. + The portable serial CI lane has no credentials and installs the public Pi package, so token-free guards exercise the available Pi surfaces while unavailable tools capability-skip; run a prompt-submitting guard after every harness upgrade and before trusting refreshed per-harness evidence. Record the dated per-harness result in `docs/verification/runtime-backends.md`, and point at the live guard as the command that refreshes it, rather than leaving a version-scoped observation to rot into a false claim. @@ -111,6 +112,12 @@ Move or delete evidence only after the current owner and regression pointer are After all documentation, review-fix, and lint-fix commits, review the complete branch diff again against those criteria rather than reviewing only the latest commit. Run `bin/fm-doc-audience-check.sh`; it enforces classification, README setup routing, local link targets, and owner pointers without keyword-linting legitimate evidence prose. +## No-mistakes test configuration + +Never configure a deterministic suite-walk `commands.test` in any repository's no-mistakes config, whether it selects the full suite, changed tests, a family, or a fixed script list. +Targeted validation belongs to the no-mistakes evidence path, while CI owns broad deterministic regression coverage. +Firstmate PR #3644 demonstrated the cost: pinning a 75-162-script walk took 32.7 minutes per validation, while removing it restored the 3.6-minute targeted-validation posture. + ## Repo style rules - Put one full sentence per line in tracked Markdown. diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index c2abb959072..0c804b3f513 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -3,7 +3,7 @@ name: harness-adapters description: >- Agent-only reference for firstmate harness operations. Use before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. - Contains verified facts for claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, cursor, gemini, and muse. + Contains verified facts for claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, and omp. user-invocable: false metadata: internal: true @@ -41,6 +41,7 @@ Muse and Gemini are verified only for crewmate and scout work, never a secondmat `../../../bin/fm-harness.sh` prints firstmate's own harness from verified environment markers, then process ancestry. Only `FM_PI_HARNESS=pi-signed` at the launch boundary together with `PI_CODING_AGENT=true` selects Pi-signed; shared unmarked launcher ancestry remains Pi. +omp publishes no marker of its own; `FM_OMP_HARNESS=omp` is Firstmate's launch marker and the anchored process name `omp` is its ancestry evidence, as `references/harness/omp.md` records. `../../../bin/fm-spawn.sh` owns worker marker establishment, while the README launch command owns the signed-primary boundary. `../../../bin/fm-harness.sh crew` resolves `config/crew-harness`, where absent or `default` means firstmate's own harness. `../../../bin/fm-harness.sh secondmate` resolves `config/secondmate-harness` -> `config/crew-harness` -> firstmate's own harness. @@ -91,7 +92,9 @@ A new tool remains undispatchable until the `verify` plan, its harness entry, ev "kimi": "references/harness/kimi.md", "cursor": "references/harness/cursor.md", "gemini": "references/harness/gemini.md", - "muse": "references/harness/muse.md" + "muse": "references/harness/muse.md", + "rovo": "references/harness/rovo.md", + "omp": "references/harness/omp.md" } } ``` diff --git a/.agents/skills/harness-adapters/references/harness/muse.md b/.agents/skills/harness-adapters/references/harness/muse.md index b3642390cb7..a0a9df3e004 100644 --- a/.agents/skills/harness-adapters/references/harness/muse.md +++ b/.agents/skills/harness-adapters/references/harness/muse.md @@ -30,7 +30,7 @@ It resolves non-secret `XDG_CONFIG_HOME` and `XDG_DATA_HOME` absolutely before p With neither worker-reachable credential, spawn refuses. Unauthenticated Muse otherwise waits forever at `Sign in at this page: https://auth.meta.com/oauth/device/?code=XXXX-XXXX` and `Waiting for approval…`, which resembles a wedge. -Escalate the refusal as a needed credential. +Before escalating the refusal as a needed credential, check the [worker launch environment contract](../../../../../docs/configuration.md#worker-launch-environment-configlaunch-env-allowlist) for a withheld environment grant. ## Foreign personal context diff --git a/.agents/skills/harness-adapters/references/harness/omp.md b/.agents/skills/harness-adapters/references/harness/omp.md new file mode 100644 index 00000000000..ee78d1b1bba --- /dev/null +++ b/.agents/skills/harness-adapters/references/harness/omp.md @@ -0,0 +1,56 @@ +# omp (Oh My Pi) + +Verified for crew, scout, secondmate, and primary work on Herdr on 2026-09-05 with omp 18.1.11, building on the 2026-09-02 adapter investigation against 18.1.2. +omp is a Pi fork, so `references/harness/pi.md` is the nearest relative; every difference from Pi is stated here. +Cross-harness provider and credential identity is owned by `references/common/model-and-effort.md`. + +## Operating facts + +| Fact | Value | +|---|---| +| Binary | `omp`, a single Bun-compiled executable resolved from `PATH` by `../../../bin/fm-spawn.sh`; a missing binary refuses the spawn. | +| Launch | Foreign markers cleared (`CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, `FM_PI_HARNESS`, `GEMINI_CLI`, Cursor's), `FM_OMP_HARNESS=omp OMP_SKIP_SETUP=1`, then `omp --config <.omp/fm-worker-overlay.yml> --auto-approve --cwd [--model] [--thinking] -e state/.omp-ext.ts `; a secondmate passes no `-e` and relies on auto-discovery. | +| Busy state | `../../../bin/fm-busy-lib.sh` source `omp-ext`: the per-task extension marks busy at `agent_start` and idle at `agent_end` only when `willContinue` is not true; `ctx.isIdle()` is deliberately not consulted because it reads false at a natural TUI `agent_end` (`session_stop` is awaited before settle). | +| Exit command | `/quit` (`/exit` and `/q` are aliases). | +| Interrupt | Single Escape; the composer is left empty, no clear key. | +| Skill invocation | No separate verified form beyond normal command behavior; use natural language when the exact command is uncertain. | +| Model flag | `--model /` (fuzzy patterns are accepted by omp but bypass Firstmate's pre-launch check). | +| Effort flag | `--thinking `, a superset of the shared vocabulary, so every level including `max` maps straight across. | +| Model discovery | `omp models [--json]` lists built-in and auto-discovered providers only; extension-registered providers such as `claude-bridge` never appear, so those models pass through the spawn unvalidated with a stderr notice. `omp usage` shows provider windows; `quota-axi` covers the `claude` provider when the bridge is in use. | +| Marker | None of omp's own (verified: `PI_CODING_AGENT` absent from the binary, no `PI_CODING_AGENT_DIR` or `OMP_PROFILE` in the default profile). `FM_OMP_HARNESS=omp` is Firstmate's launch marker; ancestry matches the exact process name `omp`. | +| Composer | Pinned to `composer.shape: borderless` by the overlay, a bare `❯` (U+276F) row the shared classifier already reads; busy text is `Working…` (U+2026), the only spelling the omp busy regex accepts (the three-dot form its headless `-p` mode writes never reaches a supervised pane), with the status row's braille spinner plus elapsed cell as the second signal. | +| Autonomy | `--auto-approve` owns approval (omp forces `tools.approvalMode: yolo` for the session under it); the overlay pins `plan.defaultOnStartup: false`, `prewalk.enabled: false`, `retry.usageReservePolicy: auto`. | +| Trust | No project-trust gate at all; a fresh profile shows a provider-login wizard instead, suppressed by `OMP_SKIP_SETUP=1`. | +| Resume | `-c/--continue` and `-r/--resume` exist but carry no verified pane-resume contract; use deterministic relaunch. | + +Keep the instructions as one positional argument; a second positional never surfaced as a submitted message. +The openai-codex models reach an extension-registered tool through omp's `xd://` virtual-file bridge: the model reads `xd://fm_watch_arm_omp` for the description and writes `xd://fm_watch_arm_omp` to invoke it, so a transcript or rpc stream shows a `write` to that path rather than a direct `fm_watch_arm_omp` call; both are the same invocation (verified 18.1.11). +omp cold start is roughly twenty seconds to the first agent turn, paid once per worker. + +## Detection + +`../../../bin/fm-harness.sh` tests `FM_OMP_HARNESS=omp` before `CLAUDECODE`, like Cursor's markers, and its ancestry walk matches the anchored process name `omp` above the interpreter fallback. +The omp template in `../../../bin/fm-spawn.sh` clears every foreign marker at its own launch boundary, and `FM_OMP_HARNESS=omp` counts only under a real `omp` ancestor, so the marker inherited by any other launch is inert: an omp secondmate's workers keep their own identity and an inherited `CLAUDECODE` cannot outrank a worker that omp launched. +`../../../bin/fm-session-lock-lib.sh` matches the same anchored name for session-lock ownership, and `../../../bin/backends/tmux.sh` classifies it `agent` for liveness. +The optional claude-bridge extension runs a nested executable literally named `claude` as a sibling of tool execution, never an ancestor of it, so omp's own tool calls detect as omp; that subtree is never walked by a Firstmate script. + +## Worker posture overlay + +The captain's own `~/.omp/agent/config.yml` is never written; the tracked `.omp/fm-worker-overlay.yml` is passed with `--config` for the one session and pins only the settings whose captain-level values would park an unattended worker on a prompt, change its pinned model, or make its composer unreadable. +`../../../bin/fm-spawn.sh`'s header owns the exact list and the reason for each pin. + +## Extension loading + +omp auto-discovers `/.omp/extensions/*.ts` (top level only, cwd only, no ancestor walk, no trust dialog) and the active profile's `agent/extensions/`; `.pi/extensions/` is not a discovery root. +A file that is both auto-discovered and named with `-e` loads twice, so the per-task worker extension lives in `state/` and a secondmate launch names no `-e` at all. +There is no `agent_settled` event; `agent_end` plus `willContinue` replaces it. + +## Primary integration + +The omp primary follows the Pi extension-owned watcher model through `../../../docs/supervision-protocols/omp.md`: `.omp/extensions/fm-primary-omp-watch.ts` arms `bin/fm-watch-arm.sh --restart` through the `fm_watch_arm_omp` tool and owns every successor, and `.omp/extensions/fm-primary-turnend-guard.ts` answers omp's blocking `session_stop` hook by forcing one continuation when `../../../bin/fm-turnend-guard.sh` returns 2, bounded per turn by omp's `stop_hook_active` flag. +The same file ports the `tool_call` seatbelts and delivers the session-start digest through `before_agent_start` on the Run tier; omp's `session_start` carries no reason, so the source is derived (first start `startup` or `resume` from the launch line, later in-process starts `clear`, `session_compact` as `compact`). +omp has no asynchronous Stop-hook equivalent, so the Claude auto-arm model does not apply; `fm_supervision_model` classifies omp as `extension`, and `fm_omp_extension_owns_supervision` in `../../../bin/fm-wake-lib.sh` is the ownership proof that tolerates the extension's own watcher hand-off. +The Pi supervision branch is out of scope for omp; every actionable wake is delivered to main. +Launch a primary with plain `omp` inside the home (`FM_OMP_HARNESS=omp omp` when starting from a Claude pane); `../../../bin/fm-session-start.sh` prints `OMP_WATCH_EXTENSION: not loaded` when the running session has not loaded both tracked extensions. +`FM_OMP_LIVE_E2E=1 ../../../tests/fm-omp-primary-live-e2e.test.sh` is the opt-in live guard; `../../../tests/fm-omp-harness.test.sh` is the portable regression. +A secondmate registered with `remote=1` in `data/secondmates.md`, spawned through the ordinary `../../../bin/fm-spawn.sh --secondmate` path, is refused on omp until a remote host verifies it, as is `../../../bin/fm-remote-secondmate-control.sh launch`; there is no `--remote` flag. diff --git a/.agents/skills/harness-adapters/references/harness/opencode.md b/.agents/skills/harness-adapters/references/harness/opencode.md index 0d0eb6912fe..8b37a8d35ad 100644 --- a/.agents/skills/harness-adapters/references/harness/opencode.md +++ b/.agents/skills/harness-adapters/references/harness/opencode.md @@ -37,6 +37,7 @@ The primary integration was verified on 2026-07-08 with OpenCode 1.17.6. Throwing from `session.idle` does not block `opencode run`, so the primary adapter treats the event as passive and uses `client.session.promptAsync` to force one follow-up turn when `../../../bin/fm-turnend-guard.sh` returns 2. The follow-up was verified in the interactive TUI. `opencode run` can exit before displaying a queued follow-up, so the adapter steps aside in headless mode. +On native Windows, the operational-input adapter runs its Bash helper through `bash`; macOS and Linux invoke it directly. The companion `.opencode/plugins/fm-primary-watch-arm.js` owns normal TUI watcher supervision, wakes it with `client.session.promptAsync`, and coordinates with the guard before a blind-turn follow-up. The PreToolUse-equivalent watcher-arm seatbelt blocks by throwing from `tool.execute.before`. diff --git a/.agents/skills/harness-adapters/references/harness/pi.md b/.agents/skills/harness-adapters/references/harness/pi.md index ebbca27ddc6..0455efe6681 100644 --- a/.agents/skills/harness-adapters/references/harness/pi.md +++ b/.agents/skills/harness-adapters/references/harness/pi.md @@ -44,6 +44,7 @@ Pi sets `PI_CODING_AGENT=true` for its children as its harness-detection marker. The primary turn-end behavior was verified on 2026-07-09 with Pi 0.80.5. `.pi/extensions/fm-primary-turnend-guard.ts` listens for logical-run `agent_settled`, not per-tool-loop `turn_end`, and uses `pi.sendUserMessage(..., { deliverAs: "followUp" })` to force one guarded follow-up when `../../../bin/fm-turnend-guard.sh` returns 2. Without `deliverAs: "followUp"`, Pi rejects the send while the agent is still processing. +On native Windows, the extension runs its session-start, both PreToolUse, turn-end, and operational-input Bash helpers through `bash`; macOS and Linux invoke those helpers directly. The primary watcher protocol also requires `.pi/extensions/fm-primary-pi-watch.ts`. The Pi engine auto-discovers both tracked project-local extensions once the project is trusted. diff --git a/.agents/skills/harness-adapters/references/harness/rovo.md b/.agents/skills/harness-adapters/references/harness/rovo.md new file mode 100644 index 00000000000..7cb313d0c48 --- /dev/null +++ b/.agents/skills/harness-adapters/references/harness/rovo.md @@ -0,0 +1,76 @@ +# Rovo CLI + +Verified 2026-09-02 on Rovo CLI 202609.1.2 for crewmate/scout work only. +Not verified, and not naturally verifiable, as a secondmate or primary: rovo has no turn-end hook and no primary supervision protocol, the same gap that scopes muse to crewmate/scout. + +## Operating facts + +| Fact | Value | +|---|---| +| Binary | `resolve_rovo_binary` in `../../../bin/fm-spawn.sh` resolves `PATH`, then falls back to `$HOME/.local/bin/rovo`; spawning refuses if neither is executable. | +| Launch | Bare `rovo run --yolo` (no positional brief), the kimi launch-then-send shape: a readiness gate on the `Welcome to Rovo!` banner, then a typed absolute brief pointer, then a delivery-confirmation gate. A positional brief is dead-on-arrival (see "Launch and readiness" below). | +| Models | `--model `, discovered from the in-session `/models` command or ACP `session/new`; the observed live list (GPT-5.6 Terra/Sol/Luna, GPT-5.5, GPT-5.4, several Claude Sonnet/Opus/Haiku ids, Gemini 3 ids) is per-account and must never be hardcoded. | +| Busy state | Rendered-tail fallback, isolated to rovo like Grok's - the animated `Rovo is thinking...` line, matched by `fm_busy_rovo_tail_busy` in `../../../bin/fm-busy-lib.sh` - because rovo's `eventHooks` fire at tool granularity only (`on_tool_start`/`on_tool_end`), never at turn-end, so no semantic writer exists to arm. | +| Exit command | `/exit` (also `/quit`, and a single idle Ctrl-C); prints `Run rovo --restore to resume your conversation`. | +| Interrupt | Single Escape is the cancel key and prints `Agent cancelled`; `../../../bin/fm-control-lib.sh` records its acknowledgement source as `none` (see "Interrupt: confirmed under real tmux" below), the same conservative choice as claude/codex/grok/kimi/cursor. | +| Skill invocation | `/`, the Claude/Grok form, but see "Skill-loading interop gap" below - a rovo worker cannot invoke a firstmate skill until that gap is resolved. | +| Autonomy | `--disable-permission-checks` (alias `--yolo`) runs every file CRUD operation and bash command without confirmation, though its own printed caveat keeps permission checks on tools accessing Atlassian data and user-provided MCP servers, which crew/scout tasks never touch. | +| File access | rovo confines every file-tool operation to its launch worktree by default, so the standard instructions/steering/status/report loop - whose files live in the firstmate home outside the worktree - fails until granted. `../../../bin/fm-spawn.sh`'s `rovo_config_override_flag` grants `toolPermissions.allowedExternalPaths` at launch, folded into the single `--config-override` (see Effort), for exactly this task's brief directory, steering inbox, and status file. The grant lifts the file tools only; rovo's bash tool stays worktree-confined regardless, so the crewmate status line's `echo ... >> status` lands only because the worker falls back to its own file tool for the append. See `../../../../docs/verification/rovo.md`. | +| Trust dialog | None observed on a clean launch in a fresh worktree; `--yolo` clears crew/scout's confirmation prompts, but it is not the only launch grant the standard flow needs - see File access for the required `allowedExternalPaths` grant. | +| Environment marker | `ATLASSIAN_AGENT_TYPE=rovo` (most specific) and `ROVODEV_CLI=1`, both set on rovo's tool subprocesses alongside `AGENT=rovodev_cli`, none of which rovo scrubs from an inherited `CLAUDECODE`/`CURSOR_AGENT`/etc - so `../../../bin/fm-harness.sh` tests rovo's markers before the `CLAUDECODE` line (the same ordering hazard cursor already documents, issue #3517) and `../../../bin/fm-spawn.sh` clears foreign markers at the launch boundary too. | +| Process name | `comm=rovo` on the tool subprocess and the `rovo run` process itself, because the installed wrapper execs the generation's `rovo` shim so argv[0] stays `rovo` even though the on-disk binary is `atlassian_cli_rovodev`. | +| Composer | The existing bordered `box` shape family (`╭─╮ │ │ ╰─╯`) `../../../bin/fm-composer-lib.sh` already reads, with an empty composer showing de-emphasized suggestion chips and a `? for shortcuts.` hint, and a busy footer reading `Enter to queue, Ctrl+Enter to steer`. | +| Effort | `agent.efficiencyLevel`, accepted `low\|medium\|high\|max` (default `medium`, no CLI `--effort` flag), set live through rovo's single `--config-override` flag - folded into the SAME JSON object as the mandatory `allowedExternalPaths` grant, never emitted as a standalone override, because `--config-override` is single-value (see `../../../../docs/verification/rovo.md`) - with an `xhigh` request recorded in task metadata but omitted from that object per `../../../references/common/model-and-effort.md`'s record-and-omit contract because rovo has no `xhigh`. | + +## Detection + +`../../../bin/fm-harness.sh` checks `ATLASSIAN_AGENT_TYPE=rovo` and `ROVODEV_CLI=1` before the `CLAUDECODE` line, then falls back to ancestry (`rovo)` case, beside `kimi)`). +Both layers matter for the same reason cursor's do: marker ordering covers a rovo session a human started by hand under an inherited foreign marker, while `../../../bin/fm-spawn.sh`'s launch-boundary `env -u` clearing covers every firstmate-launched worker regardless of ordering. + +## Launch and readiness + +The launch template clears `CLAUDECODE`, `PI_CODING_AGENT`, `GROK_AGENT`, and `FM_PI_HARNESS` inline (rovo's own foreign-marker exposure), and the shared outer wrap clears `CURSOR_AGENT`/`CURSOR_INVOKED_AS` like every other non-cursor harness. +rovo launches BARE (`rovo run --yolo`, plus any `--model`/`--config-override` flags) and takes its brief only after the TUI comes up - the same launch-then-send shape as kimi, wired through the same shared readers (`fm_backend_capture`, `fm_backend_composer_state`, `fm_backend_send_text_submit`): + +1. **Readiness gate** (`rovo_wait_for_ready` in `../../../bin/fm-spawn.sh`): poll for the fresh-launch `Welcome to Rovo!` ASCII banner, falling back to composer-empty. The banner is the primary signal because the composer-empty fallback is weaker for rovo than for kimi - rovo's idle composer renders an inline placeholder chip whose luminance sits above the ghost-strip threshold (see "Composer ghost text" below), so it can read non-empty. +2. **Typed pointer**: `Read the brief at and follow it exactly.`, submitted through `fm_backend_send_text_submit` (the exact wording and mechanism kimi uses). +3. **Delivery gate** (`rovo_wait_for_delivery`): composer empty AND either the echoed pointer text (`Read the brief at`) has scrolled into view or rovo's `Context:` footer percentage has advanced off zero. rovo's real footer is `Context: N.N% NN.NK/NNNK` (e.g. `Context: ▎ 3.3% 30.1K/922K`); the delivery regex tolerates the bar glyph and arbitrary spacing but anchors to the digits before the `%`, so the always-nonzero denominator (`.../922K`) can never masquerade as usage. + +A positional brief is dead-on-arrival: `rovo run --yolo ""` loads, never enters a working state, and drops back to an idle shell within about 10-15 seconds - confirmed independently four times over a raw PTY and once under real tmux 3.6a with the exact `fm-spawn.sh` send-keys shape. `--startup-receipt` cannot rescue that shape either: it requires "prompt-free interactive mode" (`Invalid value: --startup-receipt requires prompt-free interactive mode in a terminal`), so it cannot gate a launch that will have a message typed into it. The launch-then-send shape, by contrast, is confirmed live end to end (bare launch -> `Welcome to Rovo!` -> typed pointer -> `Rovo is thinking` for a real bash tool call -> clean `/exit`); see `../../../../docs/verification/rovo.md`. +rovo leaves no worktree-resident artifact and no firstmate-owned sidecar at all, and has no readiness receipt or session-id to record. + +## Composer ghost text: a known, unfixed gap + +rovo's empty composer renders an inline placeholder chip (e.g. `Summarize my open tasks`) directly inside the bordered content row, not merely as a separate suggestion list below it. +Measured live, that placeholder's foreground is `38;2;162;163;165` (luminance ~163), while real typed text in the same box is `38;2;206;207;210` (luminance ~207) - a real gap, but one that sits entirely above `../../../bin/fm-composer-lib.sh`'s default `FM_COMPOSER_GHOST_LUMA_MAX` of 128, so `fm_composer_strip_ghost` does not strip it and a fresh rovo composer can misclassify as `pending` instead of `empty`. +Raising the shared default to catch it is not safe: muse's own real, must-not-be-stripped prompt glyph measures luminance ~149.9, below rovo's ghost luminance, so no single global threshold can keep muse's real glyph while dropping rovo's ghost chip. +This is deliberately left unfixed rather than patched with a threshold change that would risk muse's already-verified behavior; a real fix needs a harness-scoped signal the shared composer classifier does not currently carry. +The practical consequence is bounded to composer-emptiness consumers - steering into an idle rovo pane may see a non-empty verdict and retry through the normal doorbell ladder rather than deliver on the first try. +It does not block the launch-then-send gates: readiness leads with the `Welcome to Rovo!` banner (not composer-empty), and while the delivery gate does require composer-empty as one conjunct, it runs while rovo is actively processing the just-delivered brief - the placeholder chip renders only at idle rest, not mid-turn - so the composer reads genuinely empty during the delivery window. + +## Interrupt: confirmed under real tmux + +The original verification scout (`fm-rovo-smoke-s1`, PTY smoke) observed a single Escape print `Agent cancelled` during a running tool call. +A follow-up live check under real tmux 3.6a - an isolated `tmux -L ` session/window, not the shared fleet session - reproduced the scout's exact finding: a single Escape sent during a genuine mid-flight bash tool call printed `Agent cancelled` in the captured pane. +The launch-then-send live guard (`../../../../tests/fm-rovo-signals-live-e2e.test.sh`) now reproduces it over a raw PTY too: an earlier single fixed-timer Escape landed unreliably (the interrupt instant is timing-sensitive over a bare PTY), so the guard sends Escape across the live tool-call window until the cancel renders - a deterministic way to reproduce a timing-sensitive interrupt, and confirmed to print `Agent cancelled` every run. +Escape is the interrupt key and is what `fm_control_interrupt_key` returns. +`fm_control_interrupt_ack_source` still records `none` for rovo - the same conservative choice already made for claude/codex/grok/kimi/cursor, a control-plane fact independent of whether the render happens to appear - so the control plane sends the key and lets its own postcondition, not a parsed string, decide whether the agent actually stopped. +The interrupt key and its rendered evidence are now fully corroborated rather than in tension with the code. + +## OAuth token lifetime + +The access token lasts about one hour, but `rovo` refreshes it silently and non-interactively from a stored refresh token (about four weeks' lifetime) with no browser prompt and no visible interruption - this is standing captain-corrected guidance, not this task's own discovery, and this task's own live checks corroborated it empirically: `rovo auth status` showed `Access token expired ... but a refresh token is present`, then a plain `rovo run` completed successfully and a follow-up `rovo auth status` showed a freshly valid token with no interactive step in between. +Treat the ~1h access-token lifetime as an ordinary operational fact, not a non-negotiable-safety blocker: a rovo worker does not need to be scoped short to survive it. +`rovo auth login` (interactive browser OAuth) is needed only after roughly four weeks of disuse or if the refresh token itself is invalidated. + +## Skill-loading interop gap + +rovo's skill loader rejects every firstmate skill: `Invalid skill definition in .../SKILL.md: 'metadata -> internal': Input should be a valid string`, because firstmate's `metadata.internal` is a boolean and rovo's schema wants a string. +This blocks `/no-mistakes` and every other firstmate skill invocation inside a rovo worker until firstmate's `SKILL.md` frontmatter is made rovo-compatible (a separate, deferred follow-up - it touches every skill file and the installer contract, per `../../firstmate-coding-guidelines/SKILL.md`). +A `no-mistakes`-mode rovo ship crewmate is blocked by this gap; a rovo scout, which invokes no skill, is unaffected. + +## ACP as a future upgrade + +`rovo acp` (Agent Client Protocol) and `rovo serve --non-interactive` expose a fully structured, machine-readable turn lifecycle: `session/prompt` returns a real `{"stopReason":"end_turn"}`, and `session/cancel` is a protocol-native interrupt. +This is a cleaner done-signal than any current adapter has, but consuming it means firstmate runs a JSON-RPC client and owns the session lifecycle itself - a new backend-shaped surface, not a drop-in TUI adapter - so it is out of scope here. +It remains a deliberate future upgrade for a rovo-as-structured-backend follow-up, not a near-term path; do not build it as part of this TUI-path adapter. diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index e8550505cd6..4b2629113a6 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -118,13 +118,15 @@ Supported by tests: - the handled acknowledgement is generation-keyed to the exact source and sequence, private, path-safe, durable, and idempotent, and is the only thing that stops re-announcement; - one identity-matched owner per canonical source, across homes that share one underlying source store; - registration and ownership transitions share one per-source boundary, release is generation-bound, and uncertain process identity preserves the source for retry; -- ownership moves only once a whole generation is gone, so a crashed runner leader whose owned process group is still running never reads as stale: that surviving group is stopped before any replacement starts, and the claim is kept for retry when it cannot be; +- leaderless PID/PGID-reuse ambiguity preserves the claim without signalling or replacement, as owned by the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); +- runner lifetime, owner-lease, and launch-pacing guarantees follow the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent); - stored argv is executed directly, so an argument containing spaces or shell metacharacters is never re-split or interpreted; - oversized output is bounded rather than published whole or silently dropped. The `when` adapter's guarantees are part of the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent). **Not true, and never to be claimed:** at-least-once, no-loss, or lossless delivery, and no generic exactly-once effect either - the handled acknowledgement only stops re-announcement, it says nothing about whether a paired external effect performed before the acknowledgement call actually completed, so a crash between that effect and the call can still repeat the effect on the next replay. +Also never claim that a source cannot refresh its owning home's lease: that rule is confused-agent-grade and a deliberately marker-stripping source is out of scope, per the operating contract in [`docs/configuration.md`](../../../docs/configuration.md#process-to-event-sources-stateprocevent). The currently published `lavish-axi poll` destructively clears feedback before returning it. A result lost after that clearing and before the runner reads the process output is unrecoverable, and no firstmate wrapper can close that source-side window. diff --git a/.agents/skills/quota-array-dispatch/SKILL.md b/.agents/skills/quota-array-dispatch/SKILL.md index 157696c05e1..c2b9f05ece5 100644 --- a/.agents/skills/quota-array-dispatch/SKILL.md +++ b/.agents/skills/quota-array-dispatch/SKILL.md @@ -28,6 +28,7 @@ The helper maps each harness to its primary provider family and applies the prov An `exhausted_now` runway vetoes the candidate. The helper selects a candidate only when its applicable quota has a known `effectivePercentRemaining` greater than zero. This is an optional narrow helper with a known limitation: it maps each harness to one primary provider family only, so a candidate whose established provider differs from that primary family is checked against the wrong quota row. +omp has no primary family, so the helper keys an `omp:` candidate on its model prefix, mapping only `openai-codex/` and `claude-bridge/` and refusing every other prefix; the helper's header owns that mapping. Authoritative multi-provider routing - including provider discovery from the harness catalog and quota matching by that explicit provider - stays owned by this skill's intake procedure above and AGENTS.md section 4, not by the helper. Use it only when the brief already fixed the candidate order and every candidate's provider is the harness's primary family. It does not replace the reasoning-class, runway-feasibility, or authentication gates above. diff --git a/.agents/skills/secondmate-provisioning/SKILL.md b/.agents/skills/secondmate-provisioning/SKILL.md index 432fb362c98..46c856eb204 100644 --- a/.agents/skills/secondmate-provisioning/SKILL.md +++ b/.agents/skills/secondmate-provisioning/SKILL.md @@ -106,9 +106,11 @@ A remote home is a standalone clone on another machine, so that host imports the Neither path moves the host's Firstmate copy, and the host-local launch never re-targets that copy after the parent has already synced the home. `/updatefirstmate` is the one path that still follows that copy: it first updates the remote code root from its own origin, then syncs the home to that refreshed code-root commit. SSH exit 255 preserves the route and reports unknown completion; it never triggers local respawn or failover. -The same placement-specific launch and deferred bootstrap sweep also propagate the primary's declared inherited local material: `config/crew-dispatch.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`, and the one shared captain-preference file `data/captain-shared.md`. +The same placement-specific launch and deferred bootstrap sweep also propagate the primary's inherited local material declared by [`fm_config_inherit_items`](../../../bin/fm-config-inherit-lib.sh), whose owner also defines which items are session-scoped. Because these paths are gitignored, that propagation is a separate, primary-authoritative copy independent of the tracked-files fast-forward: it re-converges every live home whether or not its tracked files advanced, and it touches only the declared items. -Propagation failures warn without blocking secondmate launch or session-start continuation, and the destination keeps whatever safely validated state the helper left behind. +Propagation failures warn without blocking a local secondmate launch or session-start continuation; a remote prelaunch transfer failure refuses that launch. +The destination keeps whatever safely validated state the helper left behind. +For inherited config files, local propagation and the remote sender preserve the destination item on source inspection errors and mirror only proven absence; [`fm-config-inherit-lib.sh`](../../../bin/fm-config-inherit-lib.sh) owns this boundary. Inheritance copies the literal `config/crew-harness` file, so a secondmate's own crewmates use the primary's crewmate harness only when it names a concrete adapter such as `codex`; an unset or `default` value has nothing concrete to inherit, and the secondmate's own crewmates fall back to the secondmate's own or detected harness instead. Inherited `config/backend` becomes that secondmate home's local runtime-backend default for future spawns only; it never retargets, rewrites, migrates, stops, or restarts an already-live worker endpoint. A present primary value always converges byte-exact into validated secondmate homes, and primary absence removes the destination so those homes keep runtime auto-detection. @@ -127,7 +129,7 @@ Keep every `data/learnings.md` fully local by captain decision; route fleet-gene No AGENTS.md reread nudge is needed at spawn or respawn because the agent reads instructions fresh on launch; only the bootstrap sweep's running-home instruction-surface advance needs that AGENTS.md re-read. Bootstrap reports successful AGENTS.md re-read sends as `BOOTSTRAP_INFO:` and only emits `NUDGE_SECONDMATES:` when that send fails and needs retry. A separate, literal-content config reread is required whenever inherited `config/*` material changes under an already-running secondmate. -For a local home, after each successful allowlisted config write, both the locked bootstrap convergence path and mid-session `bin/fm-config-push.sh` use the shared propagation report to build one per-home generation-specific private instruction file from the validated destination post-write bytes for only the allowlisted config items that actually changed for that home (`config/crew-dispatch.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`), in deterministic allowlist order. +For a local home, after each successful allowlisted config write, both the locked bootstrap convergence path and mid-session `bin/fm-config-push.sh` use the shared propagation report to build one per-home generation-specific private instruction file from the validated destination post-write bytes for only the declared config items that actually changed for that home, in declaration order. Each changed path is printed with clear begin/end delimiters and the destination file's full exact new bytes unparsed, or the explicit token `ABSENT` when propagation removed the destination copy. The instruction uses only minimal framing that these are defaults/rules and do not remove judgment; it never includes SHA values, selected profiles, parsed summaries, or any other generated interpretation. `data/captain-shared.md` is not a config file and is never inlined into this instruction file or message. @@ -141,7 +143,8 @@ Successfully delivered generations are retained only within a bounded per-home s A remote home receives the same allowlisted bytes through `fm-remote-inherit.sh` and gets one marked re-read instruction after a changed transfer. The parent records that nudge before delivery, retains it after a failed send, and retries the exact same route during locked bootstrap convergence. It does not receive a pointer to a primary-local generation path that cannot exist on that host. -These config values remain defaults and rules only; they must not harden `fm-spawn` to reject a deliberate runtime choice that differs from the configured defaults. +Inherited harness and runtime-backend defaults must not harden `fm-spawn` to reject a deliberate runtime choice that differs from those defaults. +The [worker launch environment contract](../../../docs/configuration.md#worker-launch-environment-configlaunch-env-allowlist) separately governs explicit environment grants. For already-live secondmates, use `bin/fm-config-push.sh` to push a mid-session inherited local-material change without running the tracked-file fast-forward. It uses the same live-home discovery and propagation helper as bootstrap, reports each item as `pushed`, `unchanged`, `skipped`, or `error`, and follows the config-reread contract above for changed or pending generations. `bin/fm-home-seed.sh` refuses to copy a missing or placeholder charter. @@ -252,6 +255,7 @@ It refuses retirement while that cleanup is uncertain or unavailable, preserving Raw deletion is unsupported because a blocking process-event child can outlive its home. With `--force`, teardown is the explicit discard path. +The worktree-slot ownership contract in `bin/fm-teardown.sh` still applies: `--force` never authorizes returning a descendant pool slot that another task may own. It kills child windows, discards child work and state inside the secondmate home, removes the route, releases the lease, and removes the retired secondmate home. If forced teardown contends with a fresh task publication in any affected home, one command refuses without publishing or removing task state; treat that refusal as terminal and inspect the other operation before retrying. Relaunch and non-forced teardown remain outside that serialization. diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index 64d809c798d..c5209051a44 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -3,6 +3,7 @@ name: stuck-crewmate-recovery description: >- Agent-only playbook for stuck or missing ordinary Firstmate direct reports. Use when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer. + Also use on the inverse case: a live crewmate reporting the no-mistakes pipeline dead, unreachable, or timed out. Reconciles recorded work before escalating from targeted inspection through safe relaunch or failure. user-invocable: false metadata: @@ -39,6 +40,25 @@ Preserve its uncommitted changes and commits, keep the same task identity, and r Do not use a fresh generic spawn while the recorded worktree is unaccounted for, because allocating another worktree can split one task across two copies. If the worktree or ownership cannot be reconciled safely, leave all state intact and report the task failed or blocked with the conflicting evidence. +## A live crewmate claiming the pipeline is dead + +This is the inverse of the dead-endpoint case above: the worker is alive and the pipeline it declares dead usually is too. +A drive call blocks until the next gate or outcome, far longer than a harness lets one command run, and the daemon accepts a response immediately and runs the round in the background. +So a crewmate's timed-out, killed, or errored drive call leaves it waiting on a read it never got, and the "the daemon is gone" conclusion it draws from that is a guess, not evidence. + +Read the two authoritative sources yourself before believing the claim: + +1. `no-mistakes daemon status` for the socket. +2. `no-mistakes axi status --run ` for the run, or `bin/fm-crew-state.sh `, which already folds this contradiction in and reports a non-socket daemon-or-timeout `blocked:` line over a running or fixing run with fresh activity as superseded because the run is alive. + +A refused connection or missing socket from `daemon status` is positive daemon-down evidence and must be escalated even if the persisted run record still says running or fixing; that record can be stale after the daemon exits. +Otherwise, if the run is still running or fixing with recent activity, the claim is wrong: steer the crewmate to reattach with `no-mistakes axi run` from its own worktree, which is safe and idempotent while the run still matches its `HEAD`, and tell it a timeout is not daemon death. +Nothing reaches the captain in that case. + +Never restart, stop, or update the shared daemon on a crewmate's claim. +It is one instance serving every lane and home, so a restart kills other lanes' in-flight runs. +Only positive socket refusal or absence is a daemon-down finding; escalate that finding, or a failed run record that names a daemon error, to the captain. + ## Live-endpoint escalation Escalate in order: diff --git a/.gitattributes b/.gitattributes index dfdb8b771ce..7bdc6ed6b5a 100644 --- a/.gitattributes +++ b/.gitattributes @@ -1 +1,2 @@ +# Bash parses shell scripts with LF line endings on every supported platform. *.sh text eol=lf diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 692d5a838cc..b7811d17829 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -9,6 +9,10 @@ on: permissions: contents: read +concurrency: + group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }} + cancel-in-progress: ${{ github.event_name == 'pull_request' }} + jobs: lint: name: Lint @@ -69,11 +73,18 @@ jobs: set -eu npm install -g tasks-axi@0.2.5 tasks-axi --version + - name: Install the Pi package for the Pi extension tests + run: | + set -eu + npm install -g @earendil-works/pi-coding-agent typescript@5.9.3 + npm ls -g --depth 0 @earendil-works/pi-coding-agent + tsc --version - name: Run portable parallel shard 1 run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" bin/fm-test-run.sh --lane portable-parallel-1 \ + --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-parallel-1.json" - name: Upload shard 1 timing artifact if: always() @@ -165,6 +176,16 @@ jobs: set -eu npm install -g tasks-axi@0.2.5 tasks-axi --version + # The Pi extension tests read the installed Pi package's own types and + # runtime, so without it they gate-skip and pass silently. It is a public + # npm package and needs no credential, so CI can hold the real thing. + - name: Install the Pi package for the Pi extension tests + run: | + set -eu + npm install -g @earendil-works/pi-coding-agent typescript@5.9.3 + npm ls -g --depth 0 @earendil-works/pi-coding-agent + tsc --version + command -v cc >/dev/null || { echo "::error::a C compiler is required for portable process fixtures"; exit 1; } - name: Run portable serial shard ${{ matrix.shard }} env: # job-total rather than a literal, so shrinking or growing the matrix @@ -175,7 +196,10 @@ jobs: run: | set -eu mkdir -p "$RUNNER_TEMP/fm-test" + # The Pi package is installed above and CI provides npm and tsc, so + # any missing typecheck prerequisite is a broken lane, not a valid skip. bin/fm-test-run.sh --lane "$FM_SERIAL_LANE" \ + --fail-on-gate-skip 'Pi extension typecheck prerequisite not found' \ --json "$RUNNER_TEMP/fm-test/fm-test-timing-portable-serial-${FM_SERIAL_SHARD}.json" - name: Upload portable serial shard ${{ matrix.shard }} timing artifact if: always() @@ -357,6 +381,60 @@ jobs: shell: bash run: bin/fm-test-run.sh tests/fm-update-windows.test.sh + reconciliation-windows: + name: Windows reconciliation (${{ matrix.subject }}) + runs-on: windows-latest + timeout-minutes: 10 + strategy: + fail-fast: false + matrix: + include: + - subject: core + script: "" + case: "" + - subject: copilot-launch + script: tests/fm-spawn-dispatch-profile.test.sh + case: test_copilot_threads_model_effort_and_hooks + - subject: legacy-rollback + script: tests/fm-teardown.test.sh + case: test_legacy_record_rolls_the_stamp_back_when_the_marker_write_fails + steps: + - uses: actions/checkout@v6 + - uses: actions/setup-node@v7 + with: + node-version: 24 + - name: Install tasks-axi for the rollback fixture + if: matrix.subject == 'legacy-rollback' + shell: bash + run: | + set -eu + npm install -g tasks-axi@0.2.5 + tasks-axi --version + - name: Run focused native Windows contracts + if: matrix.subject == 'core' + shell: bash + env: + FM_LIVE: "0" + run: | + set -eu + node --version + bin/fm-test-run.sh --jobs 1 --per-script-timeout-secs 180 \ + tests/fm-reconcile-validation.test.sh \ + tests/fm-backend-herdr-treehouse.test.sh \ + tests/fm-pi-windows-shell-invocation.test.sh + - name: Run named Windows reconciliation case + if: matrix.subject != 'core' + shell: bash + env: + FM_LIVE: "0" + FM_TEST_ONLY: ${{ matrix.case }} + FM_TEST_SCRIPT: ${{ matrix.script }} + run: | + set -eu + command -v jq >/dev/null || { echo "::error::jq is required for the selected fixture"; exit 1; } + git --version + bin/fm-test-run.sh --jobs 1 --per-script-timeout-secs 300 "$FM_TEST_SCRIPT" + harness-package-compatibility: name: Harness package compatibility runs-on: ubuntu-latest diff --git a/.no-mistakes.yaml b/.no-mistakes.yaml index e259441a597..3f3aad29dd5 100644 --- a/.no-mistakes.yaml +++ b/.no-mistakes.yaml @@ -22,37 +22,16 @@ document: evidence destination, and unique safety facts, then review the complete branch diff again after every documentation or lint fix. -# Pin lint to the same owner CI runs instead of leaving it to no-mistakes' +# Pin lint to the same owner CI invokes instead of leaving it to no-mistakes' # default handling, which does not invoke the repository's canonical lint gate. -# `bin/fm-lint.sh` owns the complete lint definition, including GitHub workflow -# lint via pinned actionlint in `bin/fm-lint-workflows.sh`, and -# `.github/workflows/ci.yml` invokes it directly, with parity asserted by -# `tests/fm-lint.test.sh` and `tests/fm-lint-workflows.test.sh`. +# `bin/fm-lint.sh` owns the context-sensitive ShellCheck modes and GitHub +# workflow lint via pinned actionlint in `bin/fm-lint-workflows.sh`. +# Invocation wiring is asserted by `tests/fm-lint.test.sh` and +# `tests/fm-lint-workflows.test.sh`. # -# Pin the test baseline to the repository's own runner rather than leaving each -# gate agent to chain `bash tests/a.test.sh && bash tests/b.test.sh` by hand. -# `bin/fm-test-run.sh --changed` selects only the families the branch's changed -# files map to, runs concurrency-admitted scripts with bounded concurrency, keeps -# every unproven stateful script serial, and applies its own generous per-script -# bound - so a verification round is neither a guessed short timeout nor a -# serial chain. `bin/fm-test-run.sh` owns all of that (see its header). -# -# real-herdr-gated is excluded for the same reason the portable CI lanes exclude -# it: those scripts drive a live Herdr lab, and the dedicated required Herdr lane -# in .github/workflows/ci.yml owns that coverage. A gate baseline must not start -# real Herdr sessions on whatever machine it happens to run on. -# -# This is still NOT a complete tests/*.test.sh walk, and must not become one. -# Local no-mistakes Test is intent-targeted validation of whether the change -# meets its brief; .github/workflows/ci.yml owns broad regression (behavior -# suite, platform, security, Herdr, tmux, and lifecycle coverage). A full-suite -# override here would duplicate CI and defeat the targeted Test contract. The -# configured command is a baseline only: because firstmate always supplies -# --intent, the Test step still runs its intent-targeted evidence agent on top -# of it. +# Keep commands.test absent; the firstmate-coding-guidelines skill owns this policy. commands: lint: 'bin/fm-lint.sh' - test: 'bin/fm-test-run.sh --changed --exclude-family real-herdr-gated' # Publish each run's test evidence to the orphan no-mistakes/evidence branch linked from the PR. # The evidence is not committed to the feature or default branch. diff --git a/.omp/extensions/fm-primary-omp-watch.ts b/.omp/extensions/fm-primary-omp-watch.ts new file mode 100644 index 00000000000..93749f09dcd --- /dev/null +++ b/.omp/extensions/fm-primary-omp-watch.ts @@ -0,0 +1,1073 @@ +// Firstmate primary watcher bridge for omp (Oh My Pi). +// +// A port of .pi/extensions/fm-primary-pi-watch.ts for the omp fork. The arm, +// successor, retry, and replacement-handoff logic is the Pi contract verbatim; +// the omp-specific differences are stated once here: +// - omp auto-discovers this file from /.omp/extensions with no trust +// gate, so an omp primary or secondmate started inside its home loads it +// without -e (naming it both ways loads it twice - verified, omp 18.1.11). +// - pi.sendUserMessage returns synchronously (no promise) in omp, so "Pi +// accepted the follow-up" collapses to "the call returned"; consumption is +// still tracked at before_agent_start / message_start exactly as on Pi. +// - omp reports no session_shutdown reason, so EVERY shutdown with a pending +// actionable close persists the replacement handoff and the next owning +// session_start, in this process or a later one, replays it. Replaying a +// wake main has already drained is harmless (the queue is durable and the +// drain is idempotent); losing one across /new is not. +// - The Pi supervision branch is out of scope for omp: every actionable wake +// is delivered to main, so no branch offer is made and no calm presentation +// hooks exist. +// - The arming tool is fm_watch_arm_omp and its human fallback +// /fm-watch-arm-omp; the loaded-build marker is state/.omp-watch-extension-loaded. +// +// Session-generation ownership (stated once here): +// omp emits session_shutdown for ordinary same-process replacements (/new, +// /resume, /fork) as well as terminal quit. This extension binds one generation +// per session activation. Only the active live generation may start, stop, +// rearm, or clear the arm child. An owning replacement session_start (or fresh +// factory bind) arms its new generation without a model turn. A replacement +// handoff carries actionable closes that were still pending delivery; its +// durable state lives at state/extensions/omp-primary-watch/session-replacement-actionable.json. +// Stale callbacks from a prior generation are no-ops against the active replacement. +// +// Delivery versus consumption (stated once here): +// A main follow-up is delivered once omp accepts it (sendUserMessage returns). +// The successor pipeline never waits for the model to read it: a follow-up +// queued while main is streaming joins the running run without ever raising +// before_agent_start, so waiting on that event stalls every later close. +// Consumption is tracked only so a replacement can replay a follow-up omp had +// not consumed. An idle main consumes at before_agent_start; a streaming main +// consumes at the user message_start carrying the exact wake text; either +// event finishes the pending record, and a still-unconsumed record rides the +// replacement handoff. +import { spawn, spawnSync, type ChildProcess } from "node:child_process"; +import { createHash } from "node:crypto"; +import { mkdirSync, readFileSync, renameSync, unlinkSync, writeFileSync } from "node:fs"; +import { dirname, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +// typebox resolves inside omp's extension loader (verified, omp 18.1.11); the +// injected TypeBox compatibility shim keeps it available for tool parameters. +import { Type } from "typebox"; +// The operational-input encoder is shared with the omp extensions; its owner +// resolves bin/fm-operational-input.sh relative to its own location, which is +// the same repository root this file lives in. +import { encodeFirstmateOperationalInput } from "../../.pi/extensions/lib/fm-operational-input.ts"; + +// The omp extension API surface this file uses. omp is a Pi fork and ships no +// separately installable type package, so the contract is declared locally +// rather than imported from the Pi package name. +type ExtensionAPI = { + on?: (event: string, handler: (event: any, ctx: any) => unknown) => void; + sendUserMessage: (content: string, options?: { deliverAs?: string }) => unknown; + registerCommand?: (name: string, command: { description: string; handler: (args: string, ctx: any) => Promise | void }) => void; + registerTool?: (tool: Record) => void; +}; + +type ArmResult = { + ok: boolean; + message: string; +}; + +type LockOwnership = "owned" | "missing" | "other"; + +type CloseClassification = { + kind: "actionable" | "failure"; + message: string; +}; + +type PendingActionableClose = { + version: 1; + token: string; + message: string; + predecessorArmPid: string; + delivered?: true; +}; + +type ReplacementActionableHandoff = { + version: 2; + pending: PendingActionableClose[]; +}; + +type UnconsumedWake = { + content: string; + pending: PendingActionableClose; +}; + +type SessionGeneration = { + id: number; + stopping: boolean; + replacement: boolean; + child: ChildProcess | null; + retryTimer: ReturnType | null; + cleanupTimer: ReturnType | null; + retryFailures: number; + restoring: boolean; + seq: number; + pendingActionables: PendingActionableClose[]; + cleanupFailure: string; + // Main follow-ups omp has accepted but not yet consumed, by pending token. + // Never cleared at shutdown: a delivery continuation that runs after the + // replacement began reads it to tell a main-queued wake (replayed) from a + // branch-handled one (finished). + unconsumedWakes: Map; + // A verified successor's failure close that arrived while the pipeline was + // still delivering the wake it was started for; its bounded retry runs once + // that delivery settles instead of being skipped by the single-flight guard. + deferredClose: { message: string; predecessorArmPid: string } | null; +}; + +const extensionFile = fileURLToPath(import.meta.url); +const extensionDir = dirname(extensionFile); +const root = resolve(extensionDir, "../.."); +const fmHome = process.env.FM_HOME || process.env.FM_ROOT_OVERRIDE || root; +const fmRoot = process.env.FM_ROOT_OVERRIDE || root; +const state = process.env.FM_STATE_OVERRIDE || `${fmHome}/state`; +const config = process.env.FM_CONFIG_OVERRIDE || `${fmHome}/config`; +const armScript = `${fmRoot}/bin/fm-watch-arm.sh`; +const marker = `${state}/.omp-watch-extension-loaded`; +const handoffDir = `${state}/extensions/omp-primary-watch`; +const actionableHandoff = `${handoffDir}/session-replacement-actionable.json`; +const extensionVersion = `sha256:${createHash("sha256").update(readFileSync(extensionFile)).digest("hex")}`; +const retryBaseMs = positiveInteger("FM_WATCH_REARM_RETRY_BASE_MS", 250); +const retryMaxMs = positiveInteger("FM_WATCH_REARM_RETRY_MAX_MS", 4000); +const retryLimit = positiveInteger("FM_WATCH_REARM_RETRY_LIMIT", 5); +// 35s on Windows so the budget stays above arm's MSYS confirm default (30s in +// bin/fm-watch-arm.sh): a slow but successful Git Bash cold start must not be +// SIGTERMed mid-confirmation. Conditioned on win32 so other platforms keep 12s. +const armReadyTimeoutMs = positiveInteger( + "FM_OMP_ARM_READY_TIMEOUT_MS", + process.platform === "win32" ? 35000 : 12000, +); +const armRetireTimeoutMs = positiveInteger("FM_WATCH_ARM_RETIRE_TIMEOUT_MS", 1000); +const repairOnlyHint = "call fm_watch_arm_omp again only after a later notification says the cycle is missing, failed, or unhealthy"; +const shuttingDownMessage = "watcher: not armed - omp session is shutting down"; + +let nextGenerationId = 0; +let nextHandoffId = 0; +let activeGeneration: SessionGeneration | null = null; +let replacementHandoff: PendingActionableClose[] | null = null; +type ReplacementActionableReceiver = (pending: PendingActionableClose) => void; +type ActionableDeliveryClaim = { + owner: SessionGeneration; + settlement: Promise<"delivered" | "failed">; +}; +type ReplacementCoordinator = { + receiver: ReplacementActionableReceiver | null; + pending: PendingActionableClose[]; + nextTokenId: number; + deliveries: Map; +}; +type ReplacementCoordinatorGlobal = typeof globalThis & { + __firstmateOmpWatchReplacements?: Map; +}; +const replacementCoordinatorGlobal = globalThis as ReplacementCoordinatorGlobal; +const replacementCoordinators = replacementCoordinatorGlobal.__firstmateOmpWatchReplacements ??= new Map(); +function replacementCoordinatorFor(handoff: string): ReplacementCoordinator { + const existing = replacementCoordinators.get(handoff); + if (existing) return existing; + const created: ReplacementCoordinator = { + receiver: null, + pending: [], + nextTokenId: 0, + deliveries: new Map(), + }; + replacementCoordinators.set(handoff, created); + return created; +} +const replacementCoordinator = replacementCoordinatorFor(actionableHandoff); +const armReadiness = new WeakMap>(); +const armClose = new WeakMap>(); +// Children the extension itself asked to exit; their close is not a failure +// of the successor and never earns a deferred retry. +const armRetired = new WeakSet(); +const armRecovery = new WeakMap(); +const armPendingActionable = new WeakMap(); + +function positiveInteger(name: string, fallback: number): number { + const value = Number(process.env[name]); + if (!Number.isFinite(value) || value <= 0) return fallback; + return Math.floor(value); +} + +function parentPid(pid: string): string { + const result = spawnSync("ps", ["-o", "ppid=", "-p", pid], { encoding: "utf8" }); + if (result.status !== 0) return ""; + return result.stdout.trim(); +} + +function pidAlive(pid: string): boolean { + try { + process.kill(Number(pid), 0); + return true; + } catch { + return false; + } +} + +function lockOwnership(): LockOwnership { + let lockPid = ""; + try { + lockPid = readFileSync(`${state}/.lock`, "utf8").trim(); + } catch { + return "missing"; + } + if (!/^[0-9]+$/.test(lockPid) || lockPid === "1") return "other"; + let pid = String(process.pid); + for (let i = 0; i < 8; i += 1) { + if (pid === lockPid) return "owned"; + pid = parentPid(pid); + if (!pid || pid === "1") break; + } + return pidAlive(lockPid) ? "other" : "missing"; +} + +function markLoaded(): void { + if (lockOwnership() === "other") return; + mkdirSync(state, { recursive: true }); + writeFileSync(marker, `${extensionVersion}\n${process.pid}\n`); +} + +function actionableLine(output: string): string { + const lines = output.split(/\r?\n/); + return lines.find((line) => /^(signal:|stale:|check:|heartbeat($|:))/.test(line)) || ""; +} + +function completedActionableLine(output: string): string { + const newline = output.lastIndexOf("\n"); + return newline < 0 ? "" : actionableLine(output.slice(0, newline + 1)); +} + +// The text omp carries in a user message_start: sendUserMessage wraps a string +// as one text part, so the joined text parts equal the sent content. +function userMessageText(content: unknown): string { + if (typeof content === "string") return content; + if (!Array.isArray(content)) return ""; + const parts: string[] = []; + for (const part of content) { + if ( + typeof part === "object" && part !== null && + (part as { type?: unknown }).type === "text" && + typeof (part as { text?: unknown }).text === "string" + ) { + parts.push((part as { text: string }).text); + } + } + return parts.join("\n"); +} + +function nodeErrorCode(error: unknown): string { + return typeof error === "object" && error !== null && "code" in error + ? String((error as { code?: unknown }).code ?? "") + : ""; +} + +function createPendingActionable(message: string, predecessorArmPid: string): PendingActionableClose { + return { + version: 1, + token: `${process.pid}-${Date.now()}-${++replacementCoordinator.nextTokenId}`, + message, + predecessorArmPid, + }; +} + +function validatePendingActionable(value: unknown): PendingActionableClose { + if ( + typeof value !== "object" || value === null || + (value as { version?: unknown }).version !== 1 || + typeof (value as { token?: unknown }).token !== "string" || + !/^[0-9]+-[0-9]+-[0-9]+$/.test((value as { token: string }).token) || + typeof (value as { message?: unknown }).message !== "string" || + !actionableLine((value as { message: string }).message) || + typeof (value as { predecessorArmPid?: unknown }).predecessorArmPid !== "string" || + !/^[0-9]*$/.test((value as { predecessorArmPid: string }).predecessorArmPid) || + ((value as { delivered?: unknown }).delivered !== undefined && + (value as { delivered?: unknown }).delivered !== true) + ) { + throw new Error(`invalid omp replacement actionable handoff at ${actionableHandoff}`); + } + return value as PendingActionableClose; +} + +function validateReplacementHandoff(value: unknown): PendingActionableClose[] { + if ( + typeof value !== "object" || value === null || + (value as { version?: unknown }).version !== 2 || + !Array.isArray((value as { pending?: unknown }).pending) || + (value as { pending: unknown[] }).pending.length === 0 + ) { + throw new Error(`invalid omp replacement actionable handoff at ${actionableHandoff}`); + } + const pending = (value as { pending: unknown[] }).pending.map(validatePendingActionable); + if (new Set(pending.map((item) => item.token)).size !== pending.length) { + throw new Error(`invalid omp replacement actionable handoff at ${actionableHandoff}`); + } + return pending; +} + +function writeReplacementHandoff(pending: PendingActionableClose[]): void { + replacementHandoff = [...pending]; + mkdirSync(handoffDir, { recursive: true }); + const temporary = `${actionableHandoff}.tmp-${process.pid}-${++nextHandoffId}`; + const handoff: ReplacementActionableHandoff = { version: 2, pending }; + try { + writeFileSync(temporary, `${JSON.stringify(handoff)}\n`, { mode: 0o600 }); + renameSync(temporary, actionableHandoff); + } catch (error) { + try { + unlinkSync(temporary); + } catch { + // Preserve the original handoff publication error. + } + throw error; + } +} + +function persistReplacementHandoff(pending: PendingActionableClose[]): void { + if (pending.length === 0) return; + writeReplacementHandoff(pending); +} + +function loadReplacementHandoff(): PendingActionableClose[] { + try { + const pending = validateReplacementHandoff(JSON.parse(readFileSync(actionableHandoff, "utf8"))); + replacementHandoff = pending; + return [...pending]; + } catch (error) { + if (nodeErrorCode(error) === "ENOENT") { + replacementHandoff = null; + return []; + } + throw error; + } +} + +function mergeReplacementHandoff(pending: PendingActionableClose): void { + let stored: PendingActionableClose[] = []; + try { + stored = validateReplacementHandoff(JSON.parse(readFileSync(actionableHandoff, "utf8"))); + } catch (error) { + if (nodeErrorCode(error) !== "ENOENT") throw error; + } + if (!stored.some((item) => item.token === pending.token)) stored.push(pending); + writeReplacementHandoff(stored); +} + +function clearReplacementHandoff(pending: PendingActionableClose): void { + try { + const stored = validateReplacementHandoff(JSON.parse(readFileSync(actionableHandoff, "utf8"))); + const remaining = stored.filter((item) => item.token !== pending.token); + if (remaining.length === stored.length) return; + if (remaining.length > 0) { + writeReplacementHandoff(remaining); + } else { + replacementHandoff = null; + unlinkSync(actionableHandoff); + } + } catch (error) { + if (nodeErrorCode(error) !== "ENOENT") throw error; + } +} + +function classifyClose(stdout: string, stderr: string, code: number | null, signal: NodeJS.Signals | null): CloseClassification { + const combined = `${stdout}\n${stderr}`.trim(); + const reason = actionableLine(combined); + if (reason) return { kind: "actionable", message: reason }; + const healthy = combined.split(/\r?\n/).find((line) => /^watcher: healthy\b/.test(line)); + if (healthy) { + return { + kind: "failure", + message: `watcher: FAILED - omp extension arm child found an external healthy watcher instead of owning wake delivery\n${healthy}`, + }; + } + const failed = combined.split(/\r?\n/).find((line) => /^watcher: FAILED/.test(line)); + if (failed) return { kind: "failure", message: failed }; + if (signal) { + return { + kind: "failure", + message: `watcher: FAILED - omp extension arm child ended from ${signal}${combined ? `\n${combined}` : ""}`, + }; + } + if (code && code !== 0) { + return { + kind: "failure", + message: `watcher: FAILED - fm-watch-arm.sh exited ${code}${combined ? `\n${combined}` : ""}`, + }; + } + return { + kind: "failure", + message: "watcher: FAILED - omp extension arm cycle ended without an actionable reason", + }; +} + +function createGeneration(): SessionGeneration { + return { + id: ++nextGenerationId, + stopping: false, + replacement: false, + child: null, + retryTimer: null, + cleanupTimer: null, + retryFailures: 0, + restoring: false, + seq: 0, + pendingActionables: [], + cleanupFailure: "", + unconsumedWakes: new Map(), + deferredClose: null, + }; +} + +function activateGeneration(generation: SessionGeneration): void { + activeGeneration = generation; +} + +function generationIsLive(generation: SessionGeneration): boolean { + return activeGeneration === generation && !generation.stopping; +} + +function stopGeneration(generation: SessionGeneration): ChildProcess | null { + generation.stopping = true; + if (generation.retryTimer) clearTimeout(generation.retryTimer); + if (generation.cleanupTimer) clearTimeout(generation.cleanupTimer); + generation.retryTimer = null; + generation.cleanupTimer = null; + const child = generation.child; + if (child) child.kill("SIGTERM"); + generation.child = null; + return child; +} + +async function waitForGenerationChildClose(armChild: ChildProcess | null): Promise { + if (!armChild) return; + const closed = armClose.get(armChild); + if (!closed) return; + await new Promise((resolveWait) => { + const timer = setTimeout(resolveWait, armRetireTimeoutMs); + void closed.then(() => { + clearTimeout(timer); + resolveWait(); + }); + }); +} + +async function stopSessionGeneration(generation: SessionGeneration, replacement: boolean): Promise { + generation.replacement = replacement; + let persistedTokens = ""; + try { + if (replacement && generation.pendingActionables.length > 0) { + persistReplacementHandoff(generation.pendingActionables); + persistedTokens = generation.pendingActionables.map((pending) => pending.token).join("\n"); + } + } catch (error) { + const detail = error instanceof Error ? error.message : String(error); + for (const pending of generation.pendingActionables) { + if (replacementCoordinator.pending.some((item) => item.token === pending.token)) continue; + replacementCoordinator.pending.push({ + ...pending, + message: `${pending.message}\n\nwatcher: FAILED - omp extension could not persist a replacement-session actionable wake\n${detail}`, + }); + } + throw error; + } finally { + const child = stopGeneration(generation); + await waitForGenerationChildClose(child); + } + const currentTokens = generation.pendingActionables.map((pending) => pending.token).join("\n"); + if (replacement && currentTokens && currentTokens !== persistedTokens) { + persistReplacementHandoff(generation.pendingActionables); + } +} + +const cleanupOnProcessExit = () => { + if (activeGeneration) stopGeneration(activeGeneration); +}; +process.once("exit", cleanupOnProcessExit); + +export default function (pi: ExtensionAPI) { + let generation = createGeneration(); + activateGeneration(generation); + + async function sendWake( + owner: SessionGeneration, + message: string, + pending?: PendingActionableClose, + ): Promise { + if (!generationIsLive(owner)) return false; + const content = encodeFirstmateOperationalInput( + "watcher", + `FIRSTMATE WATCHER WAKE: ${message}\n\nRun bin/fm-wake-drain.sh first and handle the queued wake. Watcher continuity is extension-owned.`, + ); + if (pending) owner.unconsumedWakes.set(pending.token, { content, pending }); + try { + await pi.sendUserMessage(content, { deliverAs: "followUp" }); + } catch (error) { + if (pending) owner.unconsumedWakes.delete(pending.token); + throw error; + } + // Accepted by omp (sendUserMessage returns synchronously there; awaiting a + // non-promise resolves at once). A generation replaced while omp was + // accepting it may have lost the follow-up with the old session, so report + // it undelivered and let the replacement replay the still-pending record. + return generationIsLive(owner); + } + + // omp consumed a main follow-up: an idle main at before_agent_start, a + // streaming main at the user message_start that joins the running run. + function consumeWake(owner: SessionGeneration, text: string): void { + for (const [token, wake] of owner.unconsumedWakes) { + if (wake.content !== text) continue; + owner.unconsumedWakes.delete(token); + wake.pending.delivered = true; + try { + finishPendingActionable(owner, wake.pending); + } catch (error) { + surfaceCleanupFailure(owner, error); + schedulePendingCleanup(owner); + } + return; + } + } + + function confirmHandlingDelivery(recovery: { generation: string; watcherPid: string }): { + ok: boolean; + detail: string; + } { + try { + const result = spawnSync( + "bash", + [armScript, "--handling-delivered", recovery.generation, "--watcher-pid", recovery.watcherPid], + { + cwd: fmRoot, + encoding: "utf8", + env: { ...process.env, FM_HOME: fmHome, FM_STATE_OVERRIDE: state, FM_ROOT_OVERRIDE: fmRoot }, + }, + ); + if (result.status === 0) return { ok: true, detail: "" }; + const stderr = (result.stderr || "").trim(); + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation was rejected (status=${result.status ?? "none"} generation=${recovery.generation} watcherPid=${recovery.watcherPid})${stderr ? `\n${stderr}` : ""}`, + }; + } catch (error) { + const message = error instanceof Error ? error.message : String(error); + return { + ok: false, + detail: `watcher: FAILED - handling delivery confirmation could not be executed (generation=${recovery.generation} watcherPid=${recovery.watcherPid})\n${message}`, + }; + } + } + + function confirmHandlingDeliveryWithRetry( + owner: SessionGeneration, + recovery: { generation: string; watcherPid: string }, + ): { ok: boolean; detail: string } { + const snapshot = (): { generation: string; watcherPid: string } => { + const current = owner.child ? armRecovery.get(owner.child) : undefined; + return current ?? recovery; + }; + const first = confirmHandlingDelivery(snapshot()); + if (first.ok) return first; + return confirmHandlingDelivery(snapshot()); + } + + async function deliverActionableWake( + owner: SessionGeneration, + message: string, + pending: PendingActionableClose, + recovery?: { generation: string; watcherPid: string }, + ): Promise { + if (!generationIsLive(owner)) return false; + if (recovery) { + const confirmed = confirmHandlingDeliveryWithRetry(owner, recovery); + if (!confirmed.ok) { + const watcherPid = recovery.watcherPid; + if (!pidAlive(watcherPid)) { + await retireArm(owner.child); + } + return await sendWake(owner, `${message}\n\n${confirmed.detail}`, pending); + } + } + // No supervision branch on omp: every actionable wake goes to main. + return await sendWake(owner, message, pending); + } + + function surfaceFailure(owner: SessionGeneration, message: string): void { + void sendWake(owner, message).catch(() => { + // omp owns delivery errors; continuity restoration never waits on prompting. + }); + } + + function enqueuePendingActionable( + owner: SessionGeneration, + pending: PendingActionableClose, + ): void { + if (owner.pendingActionables.some((item) => item.token === pending.token)) return; + owner.pendingActionables.push(pending); + if (owner.stopping && owner.replacement) { + let replacementPending = pending; + try { + mergeReplacementHandoff(pending); + } catch (error) { + const detail = error instanceof Error ? error.message : String(error); + replacementPending = { + ...pending, + message: `${pending.message}\n\nwatcher: FAILED - omp extension could not persist a late replacement-session actionable wake\n${detail}`, + }; + } + if (replacementCoordinator.receiver) { + replacementCoordinator.receiver(replacementPending); + } else if (replacementPending !== pending) { + replacementCoordinator.pending.push(replacementPending); + } + } + } + + function finishPendingActionable(owner: SessionGeneration, pending: PendingActionableClose): void { + clearReplacementHandoff(pending); + const index = owner.pendingActionables.findIndex((item) => item.token === pending.token); + if (index >= 0) owner.pendingActionables.splice(index, 1); + owner.cleanupFailure = ""; + } + + function surfaceCleanupFailure( + owner: SessionGeneration, + error: unknown, + ): void { + const detail = error instanceof Error ? error.message : String(error); + if (owner.cleanupFailure === detail) return; + owner.cleanupFailure = detail; + surfaceFailure(owner, `watcher: FAILED - omp extension could not clear a delivered replacement-session actionable wake\n${detail}`); + } + + function schedulePendingCleanup(owner: SessionGeneration): void { + if (!generationIsLive(owner) || owner.cleanupTimer) return; + const timer = setTimeout(() => { + if (owner.cleanupTimer === timer) owner.cleanupTimer = null; + void processPendingActionables(owner); + }, retryDelay(1)); + timer.unref(); + owner.cleanupTimer = timer; + } + + async function processPendingActionables(owner: SessionGeneration): Promise { + if (!generationIsLive(owner) || owner.restoring || owner.pendingActionables.length === 0) return; + owner.restoring = true; + const attemptedCleanup = new Set(); + try { + while (generationIsLive(owner) && owner.pendingActionables.length > 0) { + for (const delivered of owner.pendingActionables.filter((item) => item.delivered && !attemptedCleanup.has(item.token))) { + attemptedCleanup.add(delivered.token); + try { + finishPendingActionable(owner, delivered); + } catch (error) { + surfaceCleanupFailure(owner, error); + } + } + // A record omp has accepted but not consumed is neither redelivered + // nor finished here: consumption finishes it, replacement replays it. + const pending = owner.pendingActionables.find( + (item) => !item.delivered && !owner.unconsumedWakes.has(item.token), + ); + if (!pending) break; + const existingClaim = replacementCoordinator.deliveries.get(pending.token); + if (existingClaim && existingClaim.owner !== owner) { + const settlement = await existingClaim.settlement; + if (!generationIsLive(owner)) return; + if (settlement === "delivered") { + pending.delivered = true; + continue; + } + if (replacementCoordinator.deliveries.get(pending.token) === existingClaim) { + replacementCoordinator.deliveries.delete(pending.token); + } + } + let settleClaim: (settlement: "delivered" | "failed") => void = () => {}; + const settlement = new Promise<"delivered" | "failed">((resolveSettlement) => { + settleClaim = resolveSettlement; + }); + const deliveryClaim = { owner, settlement }; + replacementCoordinator.deliveries.set(pending.token, deliveryClaim); + const releaseClaim = (): void => { + if (replacementCoordinator.deliveries.get(pending.token) === deliveryClaim) { + replacementCoordinator.deliveries.delete(pending.token); + } + }; + try { + // A new restoration supersedes whatever became of the previous + // successor; only a failure during this delivery is retried after it. + owner.deferredClose = null; + const restoration = await restoreAfterActionableClose(owner, pending.predecessorArmPid); + if (!generationIsLive(owner)) { + settleClaim("failed"); + releaseClaim(); + return; + } + const message = restoration.failure ? `${pending.message}\n\n${restoration.failure}` : pending.message; + const delivered = await deliverActionableWake(owner, message, pending, restoration.recovery); + if (!delivered) { + settleClaim("failed"); + releaseClaim(); + return; + } + const awaitingConsumption = owner.unconsumedWakes.has(pending.token); + if (awaitingConsumption && !generationIsLive(owner)) { + // omp accepted the follow-up, then the session was replaced before + // this continuation ran: the shutdown persisted the still-pending + // record, so a replacement waiting on this claim must replay it. + settleClaim("failed"); + releaseClaim(); + return; + } + settleClaim("delivered"); + if (!awaitingConsumption) { + // omp consumed it before this ran. + pending.delivered = true; + try { + finishPendingActionable(owner, pending); + } catch (error) { + surfaceCleanupFailure(owner, error); + } + } + releaseClaim(); + } catch (error) { + settleClaim("failed"); + releaseClaim(); + throw error; + } + } + } catch (error) { + const detail = error instanceof Error ? error.message : String(error); + surfaceFailure(owner, `watcher: FAILED - omp extension could not deliver an actionable wake\n${detail}`); + } finally { + if (generationIsLive(owner)) { + owner.restoring = false; + if (owner.pendingActionables.some((pending) => pending.delivered)) schedulePendingCleanup(owner); + // No bare arm is launched here. A generation without a child at this + // point has either delivered a typed restoration failure after its + // bounded retries, which hands repair to main through fm_watch_arm_omp + // (one more silent launch past the bound could hold a hung child that + // the repair call would then report as "unchanged"), or lost a + // verified successor during the delivery, which takes the ordinary + // bounded, lock-checked retry it would have taken had the pipeline + // been idle. + const deferred = owner.deferredClose; + owner.deferredClose = null; + if (deferred && !owner.child && !owner.retryTimer) { + scheduleRetry(owner, deferred.message, deferred.predecessorArmPid); + } + } + } + } + + const receiveReplacementActionable: ReplacementActionableReceiver = (pending) => { + if (!generationIsLive(generation)) return; + enqueuePendingActionable(generation, pending); + void processPendingActionables(generation); + }; + + function retryDelay(attempt: number): number { + return Math.min(retryMaxMs, retryBaseMs * 2 ** Math.max(0, attempt - 1)); + } + + function waitForRetry(attempt: number): Promise { + return new Promise((resolveRetry) => { + const timer = setTimeout(resolveRetry, retryDelay(attempt)); + timer.unref(); + }); + } + + function waitForReadiness(armChild: ChildProcess): Promise { + const readiness = armReadiness.get(armChild); + if (!readiness) return Promise.resolve(false); + return new Promise((resolveReady) => { + const timer = setTimeout(() => resolveReady(false), armReadyTimeoutMs); + timer.unref(); + void readiness.then((ready) => { + clearTimeout(timer); + resolveReady(ready); + }); + }); + } + + async function retireArm(armChild: ChildProcess | null): Promise { + if (!armChild) return true; + armRetired.add(armChild); + armChild.kill("SIGTERM"); + const closed = armClose.get(armChild); + if (!closed) return false; + return new Promise((resolveRetired) => { + const timer = setTimeout(() => resolveRetired(false), armRetireTimeoutMs); + timer.unref(); + void closed.then(() => { + clearTimeout(timer); + resolveRetired(true); + }); + }); + } + + async function restoreAfterActionableClose(owner: SessionGeneration, predecessorArmPid: string): Promise<{ + failure: string; + recovery?: { generation: string; watcherPid: string }; + }> { + let failure = ""; + for (let attempt = 0; attempt <= retryLimit; attempt += 1) { + if (!generationIsLive(owner)) return { failure: "" }; + const replacement = startArm(owner, predecessorArmPid); + const successorChild = owner.child; + if (replacement.ok && successorChild && await waitForReadiness(successorChild)) { + return { failure: "", recovery: armRecovery.get(successorChild) }; + } + if (replacement.ok) { + failure = "watcher: FAILED - omp extension could not verify a ready successor watcher"; + if (!(await retireArm(successorChild))) { + return { + failure: `${failure}\nwatcher: FAILED - omp extension could not restore watcher continuity because the unready successor arm did not exit within ${armRetireTimeoutMs}ms`, + }; + } + } else { + failure = /(?:read-only|no live session)/.test(replacement.message) + ? `watcher: FAILED - omp extension cannot restore continuity because this session no longer owns the lock\n${replacement.message}` + : `watcher: FAILED - omp extension could not start the successor watcher cycle\n${replacement.message}`; + if (/(?:read-only|no live session)/.test(replacement.message)) break; + } + if (attempt === retryLimit) break; + await waitForRetry(attempt + 1); + } + return { failure: `${failure}\nwatcher: FAILED - omp extension could not restore watcher continuity after ${retryLimit} retries` }; + } + + function scheduleRetry(owner: SessionGeneration, message: string, predecessorArmPid: string): void { + if (!generationIsLive(owner) || owner.child || owner.retryTimer) return; + const ownership = lockOwnership(); + if (ownership !== "owned") { + surfaceFailure(owner, `watcher: FAILED - omp extension cannot restore continuity because this session no longer owns the lock\n${message}`); + return; + } + owner.retryFailures += 1; + if (owner.retryFailures > retryLimit) { + surfaceFailure(owner, `watcher: FAILED - omp extension could not restore watcher continuity after ${retryLimit} retries\n${message}`); + return; + } + const timer = setTimeout(() => { + if (owner.retryTimer === timer) owner.retryTimer = null; + if (!generationIsLive(owner)) return; + const result = startArm(owner, predecessorArmPid); + if (!result.ok) { + surfaceFailure(owner, `watcher: FAILED - omp extension could not launch a continuity retry\n${result.message}`); + } + }, retryDelay(owner.retryFailures)); + timer.unref(); + owner.retryTimer = timer; + } + + function startArm(owner: SessionGeneration, predecessorArmPid = ""): ArmResult { + if (!generationIsLive(owner)) return { ok: false, message: shuttingDownMessage }; + const ownership = lockOwnership(); + if (ownership === "other") return { ok: false, message: "watcher: read-only - session lock is held by another firstmate session" }; + if (ownership === "missing") { + return { + ok: false, + message: "watcher: not armed - no live session holds the lock; run bin/fm-session-start.sh to reclaim it, then call fm_watch_arm_omp to re-arm", + }; + } + markLoaded(); + if (owner.child) { + return { + ok: true, + message: `watcher: unchanged - omp extension already owns an arm child; no manual re-arm needed; ${repairOnlyHint}`, + }; + } + if (owner.retryTimer) { + return { + ok: true, + message: `watcher: unchanged - omp extension already owns a scheduled continuity retry; no manual re-arm needed; ${repairOnlyHint}`, + }; + } + const id = ++owner.seq; + const env = { + ...process.env, + FM_HOME: fmHome, + FM_ROOT_OVERRIDE: fmRoot, + FM_CONFIG_OVERRIDE: config, + FM_WATCH_ARM_SCRIPT: armScript, + FM_WATCH_PREDECESSOR_ARM_PID: predecessorArmPid, + }; + const armChild = spawn("bash", ["-lc", "config_dir=\"${FM_CONFIG_OVERRIDE:-$FM_HOME/config}\"; [ -f \"$config_dir/x-mode.env\" ] && . \"$config_dir/x-mode.env\"; exec \"$FM_WATCH_ARM_SCRIPT\" --restart"], { + cwd: fmRoot, + env, + stdio: ["ignore", "pipe", "pipe"], + }); + owner.child = armChild; + let stdout = ""; + let stderr = ""; + let settled = false; + let readinessSettled = false; + let verified = false; + let resolveReadiness: (ready: boolean) => void = () => {}; + let resolveClosed: () => void = () => {}; + const readiness = new Promise((resolveReady) => { + resolveReadiness = resolveReady; + }); + armReadiness.set(armChild, readiness); + const closed = new Promise((resolveClosedChild) => { + resolveClosed = resolveClosedChild; + }); + armClose.set(armChild, closed); + const settleReadiness = (ready: boolean): void => { + if (readinessSettled) return; + readinessSettled = true; + verified = ready; + resolveReadiness(ready); + }; + const observeEstablishedArm = (): void => { + const combined = `${stdout}\n${stderr}`; + const recovery = combined.match(/^watcher: started pid=([0-9]+).* recovery-generation=([A-Za-z0-9._-]+)$/m); + if (recovery) armRecovery.set(armChild, { watcherPid: recovery[1], generation: recovery[2] }); + if (/^watcher: (?:started|attached)\b/m.test(combined)) { + settleReadiness(true); + } + const reason = completedActionableLine(stdout) || completedActionableLine(stderr); + if (reason && !armPendingActionable.has(armChild)) { + const pending = createPendingActionable(reason, String(armChild.pid ?? "")); + armPendingActionable.set(armChild, pending); + enqueuePendingActionable(owner, pending); + } + }; + const releaseChild = (): void => { + if (owner.child === armChild) owner.child = null; + }; + armChild.stdout.on("data", (chunk: Buffer) => { + stdout += chunk.toString(); + observeEstablishedArm(); + }); + armChild.stderr.on("data", (chunk: Buffer) => { + stderr += chunk.toString(); + observeEstablishedArm(); + }); + armChild.on("close", (code: number | null, signal: NodeJS.Signals | null) => { + if (settled) return; + settled = true; + resolveClosed(); + settleReadiness(false); + releaseChild(); + const classification = classifyClose(stdout, stderr, code, signal); + const predecessor = String(armChild.pid ?? ""); + if (classification.kind === "actionable") { + const pending = armPendingActionable.get(armChild) ?? createPendingActionable(classification.message, predecessor); + enqueuePendingActionable(owner, pending); + if (!generationIsLive(owner)) return; + owner.retryFailures = 0; + void processPendingActionables(owner); + return; + } + if (!generationIsLive(owner)) return; + if (owner.restoring) { + // The pipeline is still delivering the wake this successor was + // started for. A verified successor that failed on its own keeps its + // bounded retry for the end of that delivery; an unready child closing + // here was retired by the restoration itself. + if (verified && !armRetired.has(armChild)) { + owner.deferredClose = { message: classification.message, predecessorArmPid: predecessor }; + } + return; + } + scheduleRetry(owner, classification.message, predecessor); + }); + armChild.on("error", (error: Error) => { + if (settled) return; + settled = true; + resolveClosed(); + settleReadiness(false); + releaseChild(); + if (!generationIsLive(owner)) return; + if (owner.restoring) return; + scheduleRetry(owner, `watcher: FAILED - omp extension arm child ${id} failed: ${error.message}`, String(armChild.pid ?? "")); + }); + return { + ok: true, + message: `watcher: started omp extension arm child ${id}; future ordinary re-arms are automatic; ${repairOnlyHint}`, + }; + } + + function activateOwnedWatch(owner: SessionGeneration): ArmResult { + if (!generationIsLive(owner)) return { ok: false, message: shuttingDownMessage }; + if (lockOwnership() !== "owned") return startArm(owner); + replacementCoordinator.receiver = receiveReplacementActionable; + let pending: PendingActionableClose[] = []; + let loadFailure = ""; + try { + pending = loadReplacementHandoff(); + } catch (error) { + const detail = error instanceof Error ? error.message : String(error); + loadFailure = `watcher: FAILED - omp extension could not load a replacement-session actionable wake\n${detail}`; + } + const inProcessPending = replacementCoordinator.pending.splice(0); + for (const actionable of [...pending, ...inProcessPending]) { + enqueuePendingActionable(owner, actionable); + } + if (owner.pendingActionables.length > 0) { + if (loadFailure) surfaceFailure(owner, loadFailure); + const armResult = startArm(owner, owner.pendingActionables[0].predecessorArmPid); + if (!armResult.ok) { + surfaceFailure(owner, `watcher: FAILED - omp extension could not arm before replacement wake delivery\n${armResult.message}`); + } + void processPendingActionables(owner); + return armResult; + } + const result = startArm(owner); + if (loadFailure) surfaceFailure(owner, `${loadFailure}\n${result.message}`); + return result; + } + + pi.on?.("before_agent_start", (event) => { + consumeWake(generation, String((event as { prompt?: unknown })?.prompt ?? "")); + }); + pi.on?.("message_start", (event) => { + const message = (event as { message?: { role?: unknown; content?: unknown } })?.message; + if (!message || message.role !== "user") return; + consumeWake(generation, userMessageText(message.content)); + }); + + pi.on?.("session_start", async () => { + if (generation.stopping) generation = createGeneration(); + activateGeneration(generation); + markLoaded(); + if (lockOwnership() !== "owned") return; + activateOwnedWatch(generation); + }); + pi.on?.("session_shutdown", async () => { + // omp carries no shutdown reason (verified: `reason` is undefined), so the + // replacement handoff is always persisted when anything is pending; a + // terminal quit then merely replays an already-drained wake next start. + if (replacementCoordinator.receiver === receiveReplacementActionable) replacementCoordinator.receiver = null; + await stopSessionGeneration(generation, true); + }); + + pi.registerCommand?.("fm-watch-arm-omp", { + description: "Arm firstmate watcher supervision through the omp extension instead of foreground bash.", + handler: async (_args, ctx) => { + const result = activateOwnedWatch(generation); + ctx?.ui?.notify?.(result.message, result.ok ? "info" : "warning"); + }, + }); + + pi.registerTool?.({ + name: "fm_watch_arm_omp", + label: "Arm firstmate watcher", + description: "Start the first required omp watcher cycle, or repair one only after a notification says the cycle is missing, failed, or unhealthy. Do not call after ordinary work or ordinary notifications; the omp extension re-arms automatically. Never run bin/fm-watch-arm.sh through bash.", + promptSnippet: "Start the first required omp watcher cycle or repair a cycle reported missing, failed, or unhealthy; ordinary re-arming is automatic.", + promptGuidelines: [ + "Call fm_watch_arm_omp only for the first required cycle or after a notification says the cycle is missing, failed, or unhealthy. Do not call it after ordinary work, turn completion, or ordinary signal, stale, check, or heartbeat handling because the omp extension owns re-arming. Never run bin/fm-watch-arm.sh through bash.", + ], + parameters: Type.Object({}), + execute: async () => { + const result = activateOwnedWatch(generation); + return { + content: [{ type: "text", text: result.message }], + details: result, + }; + }, + }); + + markLoaded(); +} diff --git a/.omp/extensions/fm-primary-turnend-guard.ts b/.omp/extensions/fm-primary-turnend-guard.ts new file mode 100644 index 00000000000..f8c7af8fb64 --- /dev/null +++ b/.omp/extensions/fm-primary-turnend-guard.ts @@ -0,0 +1,620 @@ +// Firstmate turn-end guard, pre-tool seatbelts, and native session-start +// delivery for the omp (Oh My Pi) primary. +// +// A port of .pi/extensions/fm-primary-turnend-guard.ts with the turn-end +// mechanism replaced. Pi could only ASK for a follow-up after agent_settled; +// omp's session_stop hook is awaited before the session settles and can COMPEL +// a continuation, so "no turn ends blind" (docs/turnend-guard.md) is +// structurally enforced here rather than requested. Verified on omp 18.1.11: +// a { continue: true, additionalContext } return started a fresh agent loop, +// and the continuation's own session_stop carried stop_hook_active=true, which +// bin/fm-turnend-guard.sh reads exactly as it reads Claude's payload, bounding +// the guard to one forced continuation per turn (omp's own cap of 8 +// consecutive continuations is the second backstop). session_stop does not +// fire for an interrupted turn or for task/subagent sessions, so a +// supervisor-initiated interrupt is deliberately unguarded (bin/fm-control.sh +// owns that postcondition). +// +// Session-start delivery: omp's session_start payload carries no reason field +// (verified: keys are `type` only), so the source is derived here, following +// the Cursor precedent in docs/sessionstart-nudge.md. The first session_start +// of the process is `startup` (or `resume` when the launch line named +// --continue/-c or --resume/-r); a later session_start in the same process is +// an in-process replacement (/new, /resume, /fork) and maps to `clear`, whose +// wrapper contract re-emits the digest only when this lock owner already +// completed a full startup; session_compact maps to `compact`. +// before_agent_start returning { message } was verified to reach model context +// on omp 18.1.11 (the model quoted an injected marker back), so omp qualifies +// for the Run tier. +import { spawn, spawnSync, type ChildProcess } from "node:child_process"; +import { createHash } from "node:crypto"; +import { existsSync, readFileSync, writeFileSync } from "node:fs"; +import { dirname, resolve } from "node:path"; +import { fileURLToPath } from "node:url"; +// Shared with the Pi extensions; the owner resolves bin/fm-operational-input.sh +// relative to its own location, which is this same repository root. +import { + classifyFirstmateCurrentOperationalText, + encodeFirstmateOperationalInput, +} from "../../.pi/extensions/lib/fm-operational-input.ts"; + +// The omp extension API surface this file uses, declared locally: omp ships no +// separately installable type package and is a Pi fork whose event names match +// where they are used here. +type ExtensionAPI = { + on?: (event: string, handler: (event: any, ctx: any) => unknown) => void; + sendMessage?: (message: unknown) => void; +}; + +type LockOwnership = "owned" | "missing" | "other"; + +const extensionFile = fileURLToPath(import.meta.url); +const extensionDir = dirname(extensionFile); +const root = resolve(extensionDir, "../.."); +const fmHome = process.env.FM_HOME || process.env.FM_ROOT_OVERRIDE || root; +const state = process.env.FM_STATE_OVERRIDE || `${fmHome}/state`; +const marker = `${state}/.omp-turnend-extension-loaded`; +const extensionVersion = `sha256:${createHash("sha256").update(readFileSync(extensionFile)).digest("hex")}`; + +function parentPid(pid: string): string { + const result = spawnSync("ps", ["-o", "ppid=", "-p", pid], { encoding: "utf8" }); + if (result.status !== 0) return ""; + return result.stdout.trim(); +} + +function pidAlive(pid: string): boolean { + try { + process.kill(Number(pid), 0); + return true; + } catch { + return false; + } +} + +function lockOwnership(): LockOwnership { + let lockPid = ""; + try { + lockPid = readFileSync(`${state}/.lock`, "utf8").trim(); + } catch { + return "missing"; + } + if (!/^[0-9]+$/.test(lockPid) || lockPid === "1") return "other"; + let pid = String(process.pid); + for (let i = 0; i < 8; i += 1) { + if (pid === lockPid) return "owned"; + pid = parentPid(pid); + if (!pid || pid === "1") break; + } + return pidAlive(lockPid) ? "other" : "missing"; +} + +function markLoaded(): void { + if (!existsSync(state) || lockOwnership() === "other") return; + writeFileSync(marker, `${extensionVersion}\n${process.pid}\n`); +} + +const sessionstartDeliveryBytes = 512 * 1024; + +type SessionStartContext = { + sessionManager?: { + getSessionId?: () => unknown; + }; +}; + +// The launch line is the only resume evidence omp offers an extension: its +// session_start payload has no reason and no header timestamp is guaranteed. +function launchResumeSource(): "resume" | undefined { + const args = process.argv.slice(2); + for (const arg of args) { + if ( + arg === "-c" || arg === "--continue" || + arg === "-r" || arg === "--resume" || arg.startsWith("--resume=") + ) return "resume"; + } + return undefined; +} +const sessionstartTruncatedMarker = + "\n\nOMP SESSION-START DELIVERY TRUNCATED - the digest exceeded 512 KiB. " + + "Treat omitted context as unread and inspect the named files directly before acting on it."; +const sessionstartManualFallback = + "Run `bin/fm-session-start.sh` now, exactly once, before executing any other instructions."; +const sessionstartIneligibleExit = 3; +const sessionstartRetireTimeoutMs = 1000; + +// One active generation owns native startup from child launch through context +// claim. Replacement activates first, serially retires every predecessor, and +// lets only the matching session id claim one persistent provider prerequisite. +type SessionstartSource = "startup" | "clear" | "resume" | "fork" | "compact"; +type SessionstartResult = + | { kind: "ready"; raw: string } + | { kind: "empty" | "failed" | "ineligible" | "cancelled" }; +type SessionstartMessage = { + customType: "firstmate-sessionstart-nudge"; + content: string; + display: false; + details: { kind: "session-start" }; +}; +type SessionstartGeneration = { + id: number; + sessionId: string; + source: SessionstartSource; + stopping: boolean; + delivered: boolean; + child: ChildProcess | null; + processGroupId: number | null; + childClosed: boolean; + childClose: Promise | null; + stopPromise: Promise | null; + result: Promise; +}; + +let nextSessionstartGenerationId = 0; +let activeSessionstartGeneration: SessionstartGeneration | null = null; + +function sessionIdFromContext(ctx: SessionStartContext): string { + try { + return String(ctx?.sessionManager?.getSessionId?.() ?? ""); + } catch { + return ""; + } +} + +function sessionstartGenerationIsLive(generation: SessionstartGeneration): boolean { + return activeSessionstartGeneration === generation && !generation.stopping; +} + +function signalSessionstartChild(child: ChildProcess, signal: NodeJS.Signals): void { + const pid = child.pid; + if (!pid) return; + if (process.platform === "win32") { + const args = ["/pid", String(pid), "/t"]; + if (signal === "SIGKILL") args.push("/f"); + spawnSync("taskkill", args, { stdio: "ignore" }); + return; + } + try { + process.kill(-pid, signal); + } catch { + try { + child.kill(signal); + } catch { + } + } +} + +function sessionstartProcessGroupAlive(processGroupId: number): boolean { + try { + process.kill(-processGroupId, 0); + return true; + } catch { + return false; + } +} + +function waitForSessionstartProcessGroupExit( + processGroupId: number, + timeoutMs: number, +): Promise { + return new Promise((resolveWait) => { + const startedAt = Date.now(); + const poll = (): void => { + if (!sessionstartProcessGroupAlive(processGroupId) || Date.now() - startedAt >= timeoutMs) { + resolveWait(); + return; + } + setTimeout(poll, 10); + }; + poll(); + }); +} + +function waitForSessionstartClose(generation: SessionstartGeneration, timeoutMs: number): Promise { + if (generation.childClosed || !generation.childClose) return Promise.resolve(); + return new Promise((resolveWait) => { + const timer = setTimeout(resolveWait, timeoutMs); + void generation.childClose?.then(() => { + clearTimeout(timer); + resolveWait(); + }); + }); +} + +function stopSessionstartGeneration(generation: SessionstartGeneration): Promise { + if (generation.stopPromise) return generation.stopPromise; + generation.stopping = true; + generation.stopPromise = (async () => { + const child = generation.child; + if (process.platform === "win32") { + if (!child || generation.childClosed) { + await generation.result; + return; + } + signalSessionstartChild(child, "SIGTERM"); + await waitForSessionstartClose(generation, sessionstartRetireTimeoutMs); + if (!generation.childClosed) { + signalSessionstartChild(child, "SIGKILL"); + await waitForSessionstartClose(generation, sessionstartRetireTimeoutMs); + } + return; + } + const processGroupId = generation.processGroupId; + if (!child || !processGroupId) { + await generation.result; + return; + } + try { + process.kill(-processGroupId, "SIGTERM"); + } catch { + } + await waitForSessionstartProcessGroupExit(processGroupId, sessionstartRetireTimeoutMs); + if (sessionstartProcessGroupAlive(processGroupId)) { + try { + process.kill(-processGroupId, "SIGKILL"); + } catch { + } + await waitForSessionstartProcessGroupExit(processGroupId, sessionstartRetireTimeoutMs); + } + })(); + return generation.stopPromise; +} + +function runSessionstartHook(generation: SessionstartGeneration): Promise { + return new Promise((resolveResult) => { + let settled = false; + let closeChild: () => void = () => {}; + const settle = (result: SessionstartResult): void => { + if (settled) return; + settled = true; + resolveResult(result); + }; + const supervised = process.platform !== "win32"; + const runner = `${root}/bin/fm-sessionstart-run.sh`; + // The internal --pi-prerequisite mode is shared: it is the wrapper's + // "silent exit 3 on an intentional stand-down" contract, not a Pi-only path. + let child: ChildProcess; + try { + child = spawn( + supervised ? "node" : runner, + supervised + ? [ + `${root}/.pi/extensions/lib/fm-sessionstart-supervisor.mjs`, + runner, + "--source", + generation.source, + "--pi-prerequisite", + ] + : ["--source", generation.source, "--pi-prerequisite"], + { + detached: supervised, + stdio: supervised + ? ["ignore", "pipe", "ignore", "ipc"] + : ["ignore", "pipe", "ignore"], + }, + ); + } catch { + settle(generation.stopping ? { kind: "cancelled" } : { kind: "failed" }); + return; + } + generation.child = child; + generation.processGroupId = child.pid ?? null; + generation.childClose = new Promise((resolveClose) => { + closeChild = resolveClose; + }); + const chunks: Buffer[] = []; + let observedBytes = 0; + let retainedBytes = 0; + let truncated = false; + let pendingCompletion: { code: number | null; bytes: number } | null = null; + const unrefSupervisor = (): void => { + if (!supervised) return; + child.unref(); + child.channel?.unref?.(); + const stdout = child.stdout as (NodeJS.ReadableStream & { unref?: () => void }) | null; + stdout?.unref?.(); + }; + const markClosed = (): void => { + if (generation.childClosed) return; + generation.childClosed = true; + if (generation.child === child) generation.child = null; + generation.processGroupId = null; + closeChild(); + }; + const complete = (code: number | null): void => { + unrefSupervisor(); + if (generation.stopping) { + settle({ kind: "cancelled" }); + return; + } + if (code === sessionstartIneligibleExit) { + settle({ kind: "ineligible" }); + return; + } + if (code !== 0) { + settle({ kind: "failed" }); + return; + } + const raw = Buffer.concat(chunks).toString("utf8").trim(); + if (!raw) { + settle({ kind: "empty" }); + return; + } + settle({ + kind: "ready", + raw: truncated ? `${raw}${sessionstartTruncatedMarker}` : raw, + }); + }; + const completePending = (): void => { + if (!pendingCompletion || observedBytes < pendingCompletion.bytes) return; + complete(pendingCompletion.code); + pendingCompletion = null; + }; + child.stdout?.on("data", (chunk: Buffer) => { + observedBytes += chunk.length; + if (retainedBytes >= sessionstartDeliveryBytes) { + truncated = true; + completePending(); + return; + } + const remaining = sessionstartDeliveryBytes - retainedBytes; + const retained = chunk.length <= remaining ? chunk : chunk.subarray(0, remaining); + chunks.push(retained); + retainedBytes += retained.length; + if (retained.length !== chunk.length) truncated = true; + completePending(); + }); + if (supervised) { + child.on("message", (message: unknown) => { + const result = message as { type?: unknown; code?: unknown; bytes?: unknown }; + if (result.type !== "result" || + (typeof result.code !== "number" && result.code !== null) || + typeof result.bytes !== "number") return; + pendingCompletion = { code: result.code, bytes: result.bytes }; + completePending(); + }); + } + child.on("error", () => { + markClosed(); + settle(generation.stopping ? { kind: "cancelled" } : { kind: "failed" }); + }); + child.on("close", (code) => { + markClosed(); + if (supervised) { + settle(generation.stopping ? { kind: "cancelled" } : { kind: "failed" }); + return; + } + complete(code); + }); + }); +} + +function createSessionstartGeneration( + source: SessionstartSource, + sessionId: string, +): SessionstartGeneration { + const previous = activeSessionstartGeneration; + const generation: SessionstartGeneration = { + id: ++nextSessionstartGenerationId, + sessionId, + source, + stopping: false, + delivered: false, + child: null, + processGroupId: null, + childClosed: false, + childClose: null, + stopPromise: null, + result: Promise.resolve({ kind: "cancelled" }), + }; + activeSessionstartGeneration = generation; + generation.result = (async (): Promise => { + if (previous) await stopSessionstartGeneration(previous); + if (!sessionstartGenerationIsLive(generation)) return { kind: "cancelled" }; + return runSessionstartHook(generation); + })(); + return generation; +} + +function sessionstartMessage( + generation: SessionstartGeneration, + result: SessionstartResult, +): SessionstartMessage | undefined { + let raw = result.kind === "ready" ? result.raw : ""; + if (!raw && result.kind === "failed") { + raw = sessionstartManualFallback; + } else if (!raw && ["startup", "clear", "compact"].includes(generation.source) && + result.kind === "empty") { + raw = sessionstartManualFallback; + } + if (!raw) return undefined; + try { + // The wrapper already returns an encoded nudge on a context-preserving + // open, so only an unencoded digest or fallback needs the marker added. + const content = classifyFirstmateCurrentOperationalText(raw) + ? raw + : encodeFirstmateOperationalInput("session-start", raw); + return { + customType: "firstmate-sessionstart-nudge", + content, + display: false, + details: { kind: "session-start" }, + }; + } catch { + return undefined; + } +} + +async function claimSessionstartMessage( + generation: SessionstartGeneration, + ctx?: SessionStartContext, +): Promise { + const result = await generation.result; + if (!sessionstartGenerationIsLive(generation) || generation.delivered) return undefined; + const currentSessionId = ctx ? sessionIdFromContext(ctx) : ""; + if (generation.sessionId && currentSessionId && generation.sessionId !== currentSessionId) { + return undefined; + } + generation.delivered = true; + return sessionstartMessage(generation, result); +} + +// The shared guard reads stop_hook_active exactly as it does from Claude's +// payload: a true value allows the stop, which is what bounds omp to one +// forced continuation per turn. +function runGuard(stopHookActive: boolean): Promise<{ code: number; stderr: string }> { + return new Promise((resolveResult) => { + const child = spawn(`${root}/bin/fm-turnend-guard.sh`, { + stdio: ["pipe", "ignore", "pipe"], + }); + let stderr = ""; + child.stderr.on("data", (chunk) => { + stderr += chunk.toString(); + }); + child.on("error", () => resolveResult({ code: 0, stderr: "" })); + child.on("close", (code) => resolveResult({ code: code ?? 0, stderr })); + child.stdin.end(JSON.stringify({ stop_hook_active: stopHookActive })); + }); +} + +// PreToolUse seatbelts (bin/fm-arm-pretool-check.sh, docs/arm-pretool-check.md; +// bin/fm-cd-pretool-check.sh, docs/cd-guard.md). Both piggyback on this same +// extension file so no extra -e flag is needed: omp auto-discovers this file +// for the turn-end guard, and pi.on("tool_call", ...) can block (verified on +// omp 18.1.2: returning {block: true, reason} refused the bash command and +// surfaced the reason verbatim to the model). Each owner script owns its own +// decision and is inert outside the real primary checkout. +function runChecker(script: string, command: string): Promise<{ code: number; stderr: string }> { + return new Promise((resolveResult) => { + const child = spawn(`${root}/bin/${script}`, ["--command", command], { + stdio: ["ignore", "ignore", "pipe"], + }); + let stderr = ""; + child.stderr.on("data", (chunk) => { + stderr += chunk.toString(); + }); + child.on("error", () => resolveResult({ code: 0, stderr: "" })); + child.on("close", (code) => resolveResult({ code: code ?? 0, stderr })); + }); +} + +function runPretoolCheck(command: string): Promise<{ code: number; stderr: string }> { + return runChecker("fm-arm-pretool-check.sh", command); +} + +function runCdCheck(command: string): Promise<{ code: number; stderr: string }> { + return runChecker("fm-cd-pretool-check.sh", command); +} + +export default function (pi: ExtensionAPI) { + let sessionstartGeneration: SessionstartGeneration | null = null; + let sessionstartExitListenerRegistered = false; + let sessionStarts = 0; + const cleanupSessionstartOnProcessExit = (): void => { + const generation = sessionstartGeneration; + if (!generation) return; + if (process.platform === "win32") { + if (generation.child) signalSessionstartChild(generation.child, "SIGKILL"); + return; + } + const processGroupId = generation.processGroupId; + if (!processGroupId) { + if (generation.child) signalSessionstartChild(generation.child, "SIGKILL"); + return; + } + try { + process.kill(-processGroupId, "SIGKILL"); + } catch { + } + }; + const registerSessionstartExitListener = (): void => { + if (sessionstartExitListenerRegistered) return; + process.once("exit", cleanupSessionstartOnProcessExit); + sessionstartExitListenerRegistered = true; + }; + const removeSessionstartExitListener = (): void => { + if (!sessionstartExitListenerRegistered) return; + process.removeListener("exit", cleanupSessionstartOnProcessExit); + sessionstartExitListenerRegistered = false; + }; + registerSessionstartExitListener(); + + pi.on?.("session_start", (_event, ctx) => { + sessionStarts += 1; + const source: SessionstartSource = sessionStarts === 1 + ? (launchResumeSource() ?? "startup") + : "clear"; + markLoaded(); + registerSessionstartExitListener(); + sessionstartGeneration = createSessionstartGeneration(source, sessionIdFromContext(ctx)); + }); + + pi.on?.("before_agent_start", async (_event, ctx) => { + const generation = sessionstartGeneration; + if (!generation) return undefined; + const message = await claimSessionstartMessage(generation, ctx); + return message ? { message } : undefined; + }); + + // omp's compaction equivalent, delivered the way Pi's is: manual compaction + // is idle and auto-compaction may retry without another before_agent_start, + // so the message is sent directly while sharing generation ownership. + pi.on?.("session_compact", async (_event, ctx) => { + registerSessionstartExitListener(); + const generation = createSessionstartGeneration("compact", sessionIdFromContext(ctx)); + sessionstartGeneration = generation; + const message = await claimSessionstartMessage(generation, ctx); + if (!message || !sessionstartGenerationIsLive(generation)) return; + try { + pi.sendMessage?.(message); + } catch { + generation.delivered = false; + } + }); + + pi.on?.("session_shutdown", async () => { + const generation = sessionstartGeneration; + try { + if (generation) await stopSessionstartGeneration(generation); + } finally { + if (sessionstartGeneration === generation) sessionstartGeneration = null; + removeSessionstartExitListener(); + } + }); + + pi.on?.("tool_call", async (event) => { + if (!event || event.type !== "tool_call" || event.toolName !== "bash") return {}; + const command = String((event.input as { command?: unknown })?.command ?? ""); + if (!command) return {}; + const cdResult = await runCdCheck(command); + if (cdResult.code === 2) { + return { block: true, reason: cdResult.stderr.trim() || "denied by the cd-guard PreToolUse seatbelt" }; + } + const result = await runPretoolCheck(command); + if (result.code !== 2) return {}; + return { block: true, reason: result.stderr.trim() || "denied by the watcher-arm PreToolUse seatbelt" }; + }); + + // The blocking turn boundary. Returning undefined lets the session settle; + // returning { continue: true, additionalContext } compels one more agent + // loop with the guard text attached (verified on omp 18.1.2 and 18.1.11). + pi.on?.("session_stop", async (event) => { + const stopHookActive = Boolean(event && (event as { stop_hook_active?: unknown }).stop_hook_active === true); + const result = await runGuard(stopHookActive); + if (result.code !== 2) return undefined; + let content: string; + try { + content = encodeFirstmateOperationalInput( + "turn-end-guard", + "TURN WOULD END BLIND - supervision is off. " + + "The watcher cycle is missing, failed, or unhealthy. Follow the harness recovery instruction below before ending the turn.\n\n" + + result.stderr, + ); + } catch { + content = "TURN WOULD END BLIND - supervision is off. " + + "The watcher cycle is missing, failed, or unhealthy. Follow the harness recovery instruction below before ending the turn.\n\n" + + result.stderr; + } + return { continue: true, additionalContext: content }; + }); + + markLoaded(); +} diff --git a/.omp/fm-worker-overlay.yml b/.omp/fm-worker-overlay.yml new file mode 100644 index 00000000000..ba914e37542 --- /dev/null +++ b/.omp/fm-worker-overlay.yml @@ -0,0 +1,27 @@ +# Firstmate worker posture for omp (Oh My Pi), passed as `--config` on every +# Firstmate-launched omp session (crewmate, scout, and secondmate alike) by +# bin/fm-spawn.sh, whose header owns why the launch carries it. It is the omp +# analogue of the Pi adapter's `--tui-mode regular` pin: a per-launch overlay, +# never a write to the captain's own ~/.omp/agent/config.yml, which stays +# exactly as the captain set it (model roles, theme, providers, compaction). +# Each key below pins one setting whose captain-facing value would park an +# unattended worker on an interactive prompt, change its pinned model under it, +# or make its composer unreadable to bin/fm-composer-lib.sh. Verified against +# omp 18.1.11's own settings schema (`omp config list`). +composer: + # Renders a bare `❯` (U+276F) row, a glyph the shared composer classifier + # already reads; the shipped default `band` and six other shapes are not in + # its catalogue, and the shape is otherwise a captain-level setting. + shape: borderless +plan: + # `plan.defaultOnStartup: true` opens every session read-only; a worker that + # cannot edit files sits on its brief forever. + defaultOnStartup: false +prewalk: + # Prewalk swaps the active model for the `smol` role after the first edit, so + # a worker pinned with --model would silently change model mid-task. + enabled: false +retry: + # `confirm` pops a dialog when a usage window runs low; `auto` lets omp fall + # through the captain's own fallback chain without a keystroke. + usageReservePolicy: auto diff --git a/.opencode/plugins/lib/fm-operational-input.js b/.opencode/plugins/lib/fm-operational-input.js index ab70015ed3d..64c65fd60ee 100644 --- a/.opencode/plugins/lib/fm-operational-input.js +++ b/.opencode/plugins/lib/fm-operational-input.js @@ -13,11 +13,10 @@ export function encodeFirstmateOperationalInput(root, kind, content) { const script = existsSync(requested) ? requested : `${adapterRoot}/bin/fm-operational-input.sh`; - const child = spawn(process.platform === "win32" ? "bash" : script, [ - ...(process.platform === "win32" ? [script] : []), - "encode", - kind, - ], { + const invocation = process.platform === "win32" + ? { command: "bash", args: [script, "encode", kind] } + : { command: script, args: ["encode", kind] }; + const child = spawn(invocation.command, invocation.args, { stdio: ["pipe", "pipe", "pipe"], }); let stdout = ""; diff --git a/.pi/extensions/fm-branch-supervision.ts b/.pi/extensions/fm-branch-supervision.ts index a56bbd20e6c..7df176403f9 100644 --- a/.pi/extensions/fm-branch-supervision.ts +++ b/.pi/extensions/fm-branch-supervision.ts @@ -78,6 +78,7 @@ import { getAgentDir, keyHint, ModelRuntime, + type ModelRegistry, SessionManager, ToolExecutionComponent, type AgentSession, @@ -641,6 +642,11 @@ export default function (pi: ExtensionAPI) { // extension plus its model_select event, because createBranch runs at wake // time with no context of its own. It is what "follow main" applies. let mainModel: { provider: string; id: string } | null = null; + // Main's own model registry, captured from the contexts Pi hands this + // extension the same way mainModel is. It is the ONLY read path to + // providers an extension registered at runtime (pi-devin-auth's "devin"), + // which the branch's isolated ModelRuntime cannot see on its own. + let mainModelRegistry: ModelRegistry | null = null; // Main's own current effort needs no such tracking: Pi answers it directly // on demand, including at wake time. It throws only when the extension @@ -654,8 +660,9 @@ export default function (pi: ExtensionAPI) { } } - function rememberMainModel(ctx?: { model?: { provider: string; id: string } }): void { + function rememberMainModel(ctx?: { model?: { provider: string; id: string }; modelRegistry?: ModelRegistry }): void { if (ctx?.model) mainModel = { provider: ctx.model.provider, id: ctx.model.id }; + if (ctx?.modelRegistry) mainModelRegistry = ctx.modelRegistry; } function deliverBranchHealthNote(text: string): void { @@ -705,10 +712,55 @@ export default function (pi: ExtensionAPI) { // and same user as main, so stored credentials keep their own semantics // (OAuth stays OAuth, an API key stays an API key) and nothing is ever // installed, converted, derived, or overwritten here. + // A provider that exists only because an extension registered it into + // main's runtime (pi-devin-auth's "devin", whose streamSimple is the custom + // gRPC path no static catalog can express) is invisible to an isolated + // branch runtime until its registration is copied across. The config object + // carries that streamSimple and oauth wiring by reference, so copying it + // reuses the provider's own registration rather than reimplementing its + // wire protocol; the copy is never persisted and stays scoped to this one + // runtime. One registration that fails to compose must not blind the rest, + // so each copy is isolated. A just-registered provider's auth check has not + // run yet, so the copied providers are refreshed here and every caller's + // hasConfiguredAuth verdict is real rather than the provisional entry + // registration leaves behind. + async function copyExtensionProviders(modelRuntime: ModelRuntime): Promise { + if (!mainModelRegistry) return; + let providerIds: readonly string[]; + try { + providerIds = mainModelRegistry.getRegisteredProviderIds(); + } catch { + return; + } + const copied: string[] = []; + for (const providerId of providerIds) { + try { + const config = mainModelRegistry.getRegisteredProviderConfig(providerId); + if (config) { + modelRuntime.registerProvider(providerId, config); + copied.push(providerId); + } + } catch { + // A registration that fails to compose in the isolated runtime leaves + // that provider unavailable, exactly as if it were never copied. + } + } + if (copied.length === 0) return; + try { + await modelRuntime.refresh({ providers: copied, allowNetwork: false }); + } catch { + // A failed availability refresh is answered by hasConfiguredAuth. + } + } + async function resolveBranchModel(provider: string, modelId: string): Promise { const label = `${provider}/${modelId}`; const modelRuntime = await ModelRuntime.create(); - const model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + let model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + if (!model) { + await copyExtensionProviders(modelRuntime); + model = modelRuntime.getModel(provider, modelId) as BranchModel | undefined; + } if (!model) return { ok: false, reason: `${label} is unavailable to the isolated branch runtime` }; if (!modelRuntime.hasConfiguredAuth(provider)) { return { ok: false, reason: `${label} has no configured credentials in the isolated branch runtime` }; @@ -1656,6 +1708,7 @@ ${context.command} let available: string[]; try { const modelRuntime = await ModelRuntime.create(); + await copyExtensionProviders(modelRuntime); available = ctx.modelRegistry .getAvailable() .filter((model) => modelRuntime.getModel(model.provider, model.id) && modelRuntime.hasConfiguredAuth(model.provider)) diff --git a/.pi/extensions/fm-primary-pi-watch.ts b/.pi/extensions/fm-primary-pi-watch.ts index 03fc7cbedd2..9f41893bd0b 100644 --- a/.pi/extensions/fm-primary-pi-watch.ts +++ b/.pi/extensions/fm-primary-pi-watch.ts @@ -604,7 +604,26 @@ export default function (pi: ExtensionAPI) { // also let a check-kind trigger itself slip past main's delivery. const isCheckTrigger = /^check:/.test(message); const scope = scopeForUnreadWake(state, heartbeat); - const eligible = !isCheckTrigger && scope.eligible; + // A signal close containing a needs-decision status file, or a stale close + // for a captain-held task, gets the identical main-only treatment as a + // check-kind trigger. The cross-reference deliberately includes every + // unread decision row: until that row is read, a later signal or stale + // trigger for the same task stays on main. Other tasks and heartbeat + // handling remain independent. + const triggerKeys = /^signal:/.test(message) + ? message + .slice("signal:".length) + .split(/\s+/) + .filter(Boolean) + .map((path) => path.split("/").pop() ?? path) + : /^stale:/.test(message) + ? [message.slice("stale:".length).trim().split(/\s+/, 1)[0]].filter(Boolean) + : []; + const taskIdentity = (key: string): string => + scope.taskByWakeKey[key] ?? scope.taskByWakeKey[key.replace(/^fm-/, "")] ?? key; + const needsDecisionTasks = new Set(scope.needsDecisionKeys.map(taskIdentity)); + const isNeedsDecisionTrigger = triggerKeys.some((key) => needsDecisionTasks.has(taskIdentity(key))); + const eligible = !isCheckTrigger && !isNeedsDecisionTrigger && scope.eligible; const offer = createBranchDispatchOffer(message, scope.projects, heartbeat, eligible); pi.events?.emit?.(FM_BRANCH_DISPATCH_EVENT, offer); return offer.accepted ? offer.settlement : null; diff --git a/.pi/extensions/fm-primary-turnend-guard.ts b/.pi/extensions/fm-primary-turnend-guard.ts index 635f1f9e834..7e192182b17 100644 --- a/.pi/extensions/fm-primary-turnend-guard.ts +++ b/.pi/extensions/fm-primary-turnend-guard.ts @@ -7,6 +7,7 @@ import type { ExtensionAPI } from "@earendil-works/pi-coding-agent"; import { classifyFirstmateCurrentOperationalText, encodeFirstmateOperationalInput, + firstmateShellInvocation, } from "./lib/fm-operational-input.ts"; import { isPidInCurrentAncestry, @@ -238,25 +239,26 @@ function runSessionstartHook(generation: SessionstartGeneration): Promise { return new Promise((resolveResult) => { - const script = `${root}/bin/fm-turnend-guard.sh`; - const child = spawn( - process.platform === "win32" ? "bash" : script, - process.platform === "win32" ? [script] : [], - { stdio: ["pipe", "ignore", "pipe"] }, - ); + const invocation = firstmateShellInvocation(`${root}/bin/fm-turnend-guard.sh`, []); + let child: ChildProcess; + try { + child = spawn(invocation.command, invocation.args, { + stdio: ["pipe", "ignore", "pipe"], + }); + } catch { + resolveResult({ code: 0, stderr: "" }); + return; + } let stderr = ""; - child.stderr.on("data", (chunk) => { + child.stderr?.on("data", (chunk) => { stderr += chunk.toString(); }); child.on("error", () => resolveResult({ code: 0, stderr: "" })); child.on("close", (code) => resolveResult({ code: code ?? 0, stderr })); - child.stdin.end('{"stop_hook_active":false}'); + child.stdin?.on("error", () => {}); + child.stdin?.end('{"stop_hook_active":false}'); }); } @@ -457,14 +464,21 @@ function runGuard(): Promise<{ code: number; stderr: string }> { // script owns its own decision and is inert outside the real primary checkout. function runChecker(script: string, command: string): Promise<{ code: number; stderr: string }> { return new Promise((resolveResult) => { - const path = `${root}/bin/${script}`; - const child = spawn(process.platform === "win32" ? "bash" : path, [ - ...(process.platform === "win32" ? [path] : []), - "--command", - command, - ], { stdio: ["ignore", "ignore", "pipe"] }); + const invocation = firstmateShellInvocation( + `${root}/bin/${script}`, + ["--command", command], + ); + let child: ChildProcess; + try { + child = spawn(invocation.command, invocation.args, { + stdio: ["ignore", "ignore", "pipe"], + }); + } catch { + resolveResult({ code: 0, stderr: "" }); + return; + } let stderr = ""; - child.stderr.on("data", (chunk) => { + child.stderr?.on("data", (chunk) => { stderr += chunk.toString(); }); child.on("error", () => resolveResult({ code: 0, stderr: "" })); diff --git a/.pi/extensions/lib/fm-branch-dispatch.ts b/.pi/extensions/lib/fm-branch-dispatch.ts index dbbf58bb905..5687aa47879 100644 --- a/.pi/extensions/lib/fm-branch-dispatch.ts +++ b/.pi/extensions/lib/fm-branch-dispatch.ts @@ -1,4 +1,4 @@ -import { readdirSync, readFileSync } from "node:fs"; +import { lstatSync, readdirSync, readFileSync } from "node:fs"; import { runCommandAsync } from "./fm-async-exec.ts"; // Shared wake-dispatch handshake between the Pi watcher extension (the @@ -54,6 +54,16 @@ export interface UnreadWakeScope { * either mode. */ corrupted: boolean; + /** + * The exact "key" field of every decision-owned signal or stale row this + * scan excluded. Signal rows are marked by bin/fm-watch.sh; stale rows are + * decision-owned when their task has an open needs-decision or its current + * declaration is captain-held. fm-primary-pi-watch.ts cross-references these + * keys against the current trigger so its entire coalesced batch is forced + * to main. + */ + needsDecisionKeys: string[]; + taskByWakeKey: Record; } const EMPTY_SCOPE: UnreadWakeScope = { @@ -63,6 +73,8 @@ const EMPTY_SCOPE: UnreadWakeScope = { eligibleSeqs: [], eligibleTasks: [], corrupted: false, + needsDecisionKeys: [], + taskByWakeKey: {}, }; const UNSAFE_SCOPE: UnreadWakeScope = { status: "unsafe", @@ -71,6 +83,8 @@ const UNSAFE_SCOPE: UnreadWakeScope = { eligibleSeqs: [], eligibleTasks: [], corrupted: true, + needsDecisionKeys: [], + taskByWakeKey: {}, }; // scopeForUnreadWake is the single owner of branch-eligibility classification @@ -86,6 +100,12 @@ const UNSAFE_SCOPE: UnreadWakeScope = { // (fm-primary-pi-watch.ts forces every check-kind TRIGGER to main), so nothing // starves by being left behind. // +// A signal row whose payload is "needs-decision:"-prefixed, or a stale row +// for a task with an open needs-decision or a current captain-held declaration, +// gets the identical treatment: excluded from eligibleSeqs, never a scan veto, +// and forced to main on its own triggering close (fm-primary-pi-watch.ts's +// offerWakeToBranch). Heartbeat handling remains independent. +// // That applies to a heartbeat review too, and it is the whole point: a // heartbeat used to be deferred to main merely because some unrelated check // row happened to be sitting unread, which put a routine fleet review in the @@ -102,6 +122,71 @@ const UNSAFE_SCOPE: UnreadWakeScope = { // this repo's fm_wake_append could never have produced (an unknown kind, or a // line that fails the structural tab-field check) also still vetoes the whole // scan - that is queue corruption, not an everyday mixed queue. +function statusLineVerb(line: string): string { + const beforeColon = line.split(":", 1)[0].split("[", 1)[0].trim(); + const words = beforeColon.split(/\s+/); + if (!words.some((word) => word.startsWith("corr="))) return beforeColon; + return words.filter((word, index) => index === 0 || !/^corr=[0-9a-f]{16}$/i.test(word)).join(" "); +} + +function decisionKey(line: string): string | null { + const colon = line.indexOf(":"); + const beforeColon = colon < 0 ? line : line.slice(0, colon); + const beforeMatch = beforeColon.match(/\[key=([^\]]*)\]/); + const noteMatch = beforeMatch || colon < 0 ? null : line.slice(colon + 1).trimStart().match(/^\[key=([^\]]*)\]/); + const key = (beforeMatch ?? noteMatch)?.[1] ?? "default"; + return /^[A-Za-z0-9._-]+$/.test(key) ? key : null; +} + +function statusLineNote(line: string): string { + const colon = line.indexOf(":"); + if (colon < 0) return line; + const note = line.slice(colon + 1).trimStart(); + if (/\[key=[^\]]*\]/.test(line.slice(0, colon))) return note; + const match = note.match(/^\[key=([A-Za-z0-9._-]+)\]/); + return match ? note.slice(match[0].length).trimStart() : note; +} + +interface StaleDecisionCacheEntry { + version: string; + config: string; + decisionOwned: boolean; +} + +const staleDecisionCache = new Map(); + +function statusFileVersion(path: string): string | null { + try { + const stat = lstatSync(path); + if (stat.isSymbolicLink()) throw new Error("status path is a symbolic link"); + return `${stat.dev}:${stat.ino}:${stat.size}:${stat.mtimeMs}:${stat.ctimeMs}`; + } catch (error) { + if ((error as NodeJS.ErrnoException).code === "ENOENT") return null; + throw error; + } +} + +function hasOpenNeedsDecision( + lines: readonly string[], + resolveVerb: string, + heldVerb: string, + reservedPrefixes: readonly string[], +): boolean { + const open = new Map(); + for (const line of lines) { + const verb = statusLineVerb(line); + if (!["needs-decision", "blocked", resolveVerb, heldVerb].includes(verb)) continue; + const key = decisionKey(line); + if (!key) continue; + const note = statusLineNote(line); + const reservedPrefix = reservedPrefixes.find((prefix) => key.startsWith(prefix)); + if (reservedPrefix && !(note.startsWith(reservedPrefix) && note.slice(reservedPrefix.length).includes(":"))) continue; + if (verb === "needs-decision" || verb === "blocked") open.set(key, verb); + else open.delete(key); + } + return [...open.values()].includes("needs-decision"); +} + export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWakeScope { let queue = ""; try { @@ -128,6 +213,8 @@ export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWak if (project) { metadata.set(task, project); taskByKey.set(task, task); + taskByKey.set(`${task}.status`, task); + taskByKey.set(`${task}.turn-ended`, task); if (window) { metadata.set(window, project); taskByKey.set(window, task); @@ -140,6 +227,14 @@ export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWak const eligibleSeqs: string[] = []; const eligibleTasks = new Set(); + const needsDecisionKeys: string[] = []; + const staleDecisionOwnership = new Map(); + const resolveVerb = process.env.FM_CLASSIFY_RESOLVE_VERB || "resolved"; + const heldVerb = process.env.FM_CLASSIFY_CAPTAIN_HELD_VERB || "captain-held"; + const reservedPrefixes = (process.env.FM_CLASSIFY_RESERVED_KEY_PREFIXES || "pending-reply-") + .split(/\s+/) + .filter(Boolean); + const decisionConfig = `${resolveVerb}\0${heldVerb}\0${reservedPrefixes.join("\0")}`; for (const line of rows) { const fields = line.split("\t"); if (fields.length < 5 || !/^[0-9]+$/.test(fields[1])) return UNSAFE_SCOPE; @@ -159,11 +254,59 @@ export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWak let project = ""; let task = ""; if (kind === "signal") { + const payload = fields[4] ?? ""; + if (/^needs-decision:/.test(payload)) { + // Main-owned exactly like a check-kind row above: a needs-decision + // status append surfaced through the actionable signal path is + // excluded from what the branch may claim without vetoing the scan + // (docs/pi-supervision-branch.md "Autonomy"). + needsDecisionKeys.push(key); + continue; + } task = key.replace(/\.(?:status|turn-ended)$/, ""); project = metadata.get(task) ?? ""; } else if (kind === "stale") { task = taskByKey.get(key) ?? taskByKey.get(key.replace(/^fm-/, "")) ?? ""; project = metadata.get(key) ?? metadata.get(key.replace(/^fm-/, "")) ?? ""; + if (task) { + const statusPath = `${state}/${task}.status`; + if (!staleDecisionOwnership.has(statusPath)) { + let version: string | null; + try { + version = statusFileVersion(statusPath); + } catch { + return UNSAFE_SCOPE; + } + let decisionOwned = false; + if (version) { + const cached = staleDecisionCache.get(statusPath); + if (cached?.version === version && cached.config === decisionConfig) { + decisionOwned = cached.decisionOwned; + } else { + let statusLines: string[]; + try { + statusLines = readFileSync(statusPath, "utf8").split(/\r?\n/).filter((line) => /\S/.test(line)); + if (statusFileVersion(statusPath) !== version) return UNSAFE_SCOPE; + } catch { + return UNSAFE_SCOPE; + } + decisionOwned = hasOpenNeedsDecision(statusLines, resolveVerb, heldVerb, reservedPrefixes) || + statusLineVerb(statusLines.at(-1) ?? "") === heldVerb; + staleDecisionCache.set(statusPath, { version, config: decisionConfig, decisionOwned }); + if (staleDecisionCache.size > 512) { + staleDecisionCache.delete(staleDecisionCache.keys().next().value!); + } + } + } else { + staleDecisionCache.delete(statusPath); + } + staleDecisionOwnership.set(statusPath, decisionOwned); + } + if (staleDecisionOwnership.get(statusPath)) { + needsDecisionKeys.push(key); + continue; + } + } } else { // A kind fm_wake_append never emits: structural corruption, not an // ordinary main-only row. @@ -189,6 +332,8 @@ export function scopeForUnreadWake(state: string, heartbeat: boolean): UnreadWak eligibleSeqs, eligibleTasks: [...eligibleTasks], corrupted: false, + needsDecisionKeys, + taskByWakeKey: Object.fromEntries(taskByKey), }; } diff --git a/.pi/extensions/lib/fm-operational-input.ts b/.pi/extensions/lib/fm-operational-input.ts index 4884b8c8365..4070684c6a4 100644 --- a/.pi/extensions/lib/fm-operational-input.ts +++ b/.pi/extensions/lib/fm-operational-input.ts @@ -21,17 +21,23 @@ export type FirstmateCurrentOperationalKind = type OperationalInputCommand = "encode" | "classify" | "kind"; +export function firstmateShellInvocation( + script: string, + args: readonly string[], +): { command: string; args: string[] } { + return process.platform === "win32" + ? { command: "bash", args: [script, ...args] } + : { command: script, args: [...args] }; +} + // The one owner of how each command is invoked and how its exit status and // stdout become an answer, shared by the synchronous and awaited callers // below so the two can never drift. -function operationalInputInvocation( +function operationalInputArgs( command: OperationalInputCommand, kind?: FirstmateCurrentOperationalKind, -): { command: string; args: string[] } { - const args = command === "encode" ? [command, kind ?? ""] : [command]; - return process.platform === "win32" - ? { command: "bash", args: [operationalInputScript, ...args] } - : { command: operationalInputScript, args }; +): string[] { + return command === "encode" ? [command, kind ?? ""] : [command]; } function operationalInputAnswer( @@ -48,13 +54,20 @@ function runOperationalInputCommand( content: string, kind?: FirstmateCurrentOperationalKind, ): string | undefined { - const invocation = operationalInputInvocation(command, kind); - const result = spawnSync(invocation.command, invocation.args, { - encoding: "utf8", - input: content, - maxBuffer: 1024 * 1024, - }); - return operationalInputAnswer(command, result.status, result.stdout); + const invocation = firstmateShellInvocation( + operationalInputScript, + operationalInputArgs(command, kind), + ); + try { + const result = spawnSync(invocation.command, invocation.args, { + encoding: "utf8", + input: content, + maxBuffer: 1024 * 1024, + }); + return operationalInputAnswer(command, result.status, result.stdout ?? ""); + } catch { + return undefined; + } } function encodeFailure(kind: FirstmateCurrentOperationalKind): Error { @@ -89,7 +102,10 @@ export async function encodeFirstmateOperationalInputWith( kind: FirstmateCurrentOperationalKind, content: string, ): Promise { - const invocation = operationalInputInvocation("encode", kind); + const invocation = firstmateShellInvocation( + operationalInputScript, + operationalInputArgs("encode", kind), + ); const result = await run(invocation.command, invocation.args, { input: content }); const encoded = operationalInputAnswer("encode", result.status, result.stdout); if (encoded === undefined) throw encodeFailure(kind); diff --git a/AGENTS.md b/AGENTS.md index a2c6f15a980..a853518a25f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,5 +1,8 @@ # Firstmate +This is the supervisor contract for primary firstmates and persistent secondmates. +Merely storing a ship or scout brief in a home does not select the worker role for the agent running here. + You are the first mate. The user is the captain. This file is your entire job description. @@ -69,7 +72,7 @@ config/crew-harness crewmate harness override; LOCAL, gitignored; absent or "de config/crew-dispatch.json optional crewmate dispatch profiles; LOCAL, gitignored; firstmate-maintained but human-editable natural-language rules that choose a per-task harness/model/effort profile (section 4). Inherited by secondmate homes config/secondmate-harness harness the PRIMARY uses to launch SECONDMATE agents, optionally followed by a model and effort token on the same line (" [] []"; section 4); LOCAL, gitignored; absent or "default" harness falls back to config/crew-harness then firstmate's own. The primary's own setting; NOT inherited into secondmate homes (secondmates do not spawn secondmates) config/backlog-backend backlog backend override; LOCAL, gitignored; absent or "tasks-axi" = default tasks-axi backend, "manual" = force routine backlog updates to hand-editing; inherited by secondmate homes (section 10) -config/backend runtime session-provider backend override for new tasks; LOCAL, gitignored; absent = falls through to runtime auto-detection (the runtime firstmate itself is executing inside), then tmux; tmux is the verified reference backend (docs/tmux-backend.md), while herdr, zellij, orca, and cmux are experimental spawn backends (docs/herdr-backend.md, docs/zellij-backend.md, docs/orca-backend.md, docs/cmux-backend.md) - herdr and cmux can also be selected by runtime auto-detection, zellij and orca never are (always explicit), and codex-app is not accepted; see docs/codex-app-backend.md; inherited by secondmate homes under the primary-authoritative contract in secondmate-provisioning +config/backend runtime session-provider backend override for new tasks; LOCAL, gitignored; absent = falls through to runtime auto-detection (the runtime firstmate itself is executing inside), then tmux; tmux is the verified reference backend (docs/tmux-backend.md), herdr has its own required CI lane (docs/herdr-backend.md), while zellij, orca, and cmux remain experimental with no dedicated real-backend CI lane (docs/zellij-backend.md, docs/orca-backend.md, docs/cmux-backend.md) - herdr and cmux can also be selected by runtime auto-detection, zellij and orca never are (always explicit), and codex-app is not accepted; see docs/codex-app-backend.md; inherited by secondmate homes under the primary-authoritative contract in secondmate-provisioning config/calm Pi Calm presentation preference; LOCAL, gitignored, and not inherited; see docs/configuration.md "Pi Calm preference" config/supervision-branch-model config/supervision-branch-effort Pi supervision-branch model and reasoning-effort pins written by /supervision-model; LOCAL, gitignored, independently settable, and not inherited; see docs/configuration.md "Pi supervision branch model and effort" config/startup-memory-budget primary-authoritative per-home startup-memory budget; LOCAL, gitignored, materialized as 7,500 estimated tokens by locked primary bootstrap and inherited into secondmate homes; see docs/configuration.md "Startup memory budget" @@ -120,6 +123,7 @@ state/ runtime records and signals; gitignored procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line decision-bindings/ private records marking a captured-answer source as feeding the keyed-answer intake, with a legacy origin on pre-collapse records; written only by bin/fm-captain-hold.sh bind, dropped by unbind and by source retirement (section 13; docs/captain-hold-lifecycle.md) + reconcile-requests/ private open obligations to re-check a captain call whose board selection was `reconcile`; written only by bin/fm-captain-hold.sh, retired by its verify-then-decide outcomes or a normal answer that settles the call (section 13; docs/captain-hold-lifecycle.md) when/ private condition->action watch specs, their trust bindings, and single-fire markers; written only by bin/fm-procevent-when.sh (section 13's process-event-sources trigger) inbox/ captain notes captured out of band by bin/fm-inbox.sh, including the voice handover's queued requests; each note appends one `check` wake and stays pending until acknowledged with `bin/fm-inbox.sh drain --ack `, which moves it to inbox/handled/ (docs/voice-relay.md) x-inbox/ generated Relay pending mention payloads; fmx-respond drains it (section 14) @@ -199,7 +203,7 @@ A silent bootstrap section needs no action; for any printed actionable diagnosti ## 4. Harness and runtime dispatch Load `harness-adapters` before every spawn or recovery and before trust handling, skill invocation, interrupt, exit, resume, or adapter verification. -The verified harnesses are `claude`, `codex`, `copilot`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, and `cursor`, plus `gemini` and `muse` for crewmates and scouts only; never dispatch on an unverified adapter. +The verified harnesses are `claude`, `codex`, `copilot`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, and `omp`, plus `muse`, `gemini`, and `rovo` for crewmates and scouts only; never dispatch on an unverified adapter. If static `config/crew-harness` or `config/secondmate-harness` names an unverified adapter, report it and fall back only to a verified adapter rather than launching it. `docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation. @@ -366,7 +370,7 @@ Require the matching `resolved` event, forbid `--yes`, and require the worker to Resume fleet supervision immediately after the decision lands. Judge validation by the currently attributed run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. -Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed or checks-passed is done; failed or cancelled is failed. +Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed or checks-passed is done; failed or cancelled is failed exactly as `bin/fm-crew-state.sh` prints it - only that state line reclassifies an orphaned ci monitor after green checks as held-for-merge done, or a terminal failed record with the daemon unreachable as unknown, never the raw run record. A worker hand-editing, committing, aborting, or restarting during an active validation run duplicates pipeline ownership outside the supersession sequence above; steer it back to the gate response flow. The worker reports the PR when CI first becomes green rather than waiting for merge monitoring to finish. @@ -450,7 +454,7 @@ The skill owns the daemon procedure; these safety facts remain inline: ### Stuck-worker trigger -Load `stuck-crewmate-recovery` after a stale wake, looping or confused pane, answered-by-brief question, unresponsive worker, or failed steer. +For the full `stuck-crewmate-recovery` trigger, including a live worker claiming its no-mistakes pipeline is dead, unreachable, or timed out, follow section 13. ## 9. Escalation and captain etiquette @@ -553,7 +557,7 @@ These skills are not captain-invocable; load them only at their precise triggers - `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata. - `project-management` - load before adding, creating, removing, or initializing a project. Cloning or registering a project is add intake and uses the same trigger. -- `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer. +- `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer, and whenever a live worker reports its no-mistakes pipeline dead, unreachable, or timed out. - `secondmate-provisioning` - load before creating, seeding, validating, launching, handing backlog to, recovering, pushing inherited local material into, or retiring a secondmate home, and before editing `data/secondmates.md`. - `captain-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any `RECORD DIVERGENCE` line from the wake drain. - `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), and on any `procevent ` check wake. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a87e47e0432..4898822d207 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -21,19 +21,19 @@ Contributions use ordinary GitHub pull requests. ## Repo conventions - This repo is a template for running a firstmate orchestrator agent. - `AGENTS.md` is the agent's main job description and names when to load bundled firstmate skills; `CLAUDE.md` is a real `@AGENTS.md` pointer to it, and `.claude/skills` is a symlink to `.agents/skills`. + [`AGENTS.md`](AGENTS.md) owns the supervisor contract, role boundary, and bundled firstmate skill triggers; `CLAUDE.md` is a real `@AGENTS.md` pointer to it, and `.claude/skills` is a symlink to `.agents/skills`. - Only shared material is tracked: `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `bin/`, `.agents/skills/`, and `skills/`. `.agents/skills/` holds agent-loaded skills that assume a live firstmate home and carry `metadata.internal: true` so installers such as [skills.sh](https://skills.sh) hide them from discovery; `skills/` holds standalone, installer-facing public skills for external agents, and those skills never depend on a live firstmate home or private fleet state (see the README's "Two-tier skill layout"). Everything personal to one captain's fleet (`.env`, `data/`, `state/`, `config/`, `projects/`, `.no-mistakes/`) is gitignored; never commit it. The root `.tasks.toml` is tracked `tasks-axi` config for `data/backlog.md`; compatible `tasks-axi` is the default backend for routine backlog mutations, with the compatibility definition owned by [`docs/configuration.md`](docs/configuration.md) ("Backlog backend"). A local `config/backlog-backend=manual` opt-out forces firstmate's routine backlog updates to hand-editing and stays gitignored; validated secondmate handoffs still delegate through `tasks-axi mv`. - A local `config/backend` file explicitly overrides runtime auto-detection for new task endpoints and stays gitignored; spawn-supported values are `tmux` plus experimental `herdr`, `zellij`, `orca`, and `cmux`, while `codex-app` is documented only in `docs/codex-app-backend.md`. + A local `config/backend` file explicitly overrides runtime auto-detection for new task endpoints and stays gitignored; spawn-supported values are `tmux`, `herdr` (which has its own required CI lane), and `zellij`, `orca`, and `cmux`, which remain experimental with no dedicated real-backend CI lane, while `codex-app` is documented only in `docs/codex-app-backend.md`. It does not make `data/` tracked. - Shell helpers in `bin/` are plain Bash; semantic policy owners may use tracked `.mjs` modules, and Windows-native installers or bridges may use tracked `.ps1` scripts. Each starts with a usage header comment; keep it accurate when you change behavior. Test scripts and helpers in `tests/` are plain bash too. - `bin/fm-lint.sh` must pass: it is the single owner of the lint definition (the shellcheck file set, config, pinned shellcheck version, and pinned actionlint workflow lint), and both CI and the optional no-mistakes configuration run its no-argument full-analysis path. - Its header and `--help` output own the exact local lint modes and flags. + `bin/fm-lint.sh` must pass: it is the single owner of the lint definition (the shellcheck file set, config, pinned shellcheck version, and pinned actionlint workflow lint), and both CI and the optional no-mistakes configuration invoke it with no arguments. + Its header and `--help` output own the exact local lint modes, file-set selection, and analysis flags. A malformed `.github/workflows/*.yml`, including a self-broken `ci.yml`, fails that local lint path before merge because a broken workflow cannot report its own breakage. It pins one exact shellcheck version and one exact actionlint version and refuses to run under any other. Print the shellcheck pin with `bin/fm-lint.sh --required-version` and the actionlint pin with `bin/fm-lint-workflows.sh --required-version`. @@ -56,8 +56,8 @@ When supervising live crewmates, keep firstmate's own long validation or build c When no-mistakes is used optionally, follow the installed version's SKILL.md and live `axi` help instead of duplicating its mechanics here. Firstmate's wrapper still matters: crewmates route every `ask-user` finding to firstmate, which applies `ask-user-authority`. Crewmates never pass `--yes` or `-y` because either flag bypasses that check and any required captain escalation. -`.no-mistakes.yaml` publishes test evidence to the orphan `no-mistakes/evidence` branch, which shares no history with code branches, pins the gate's lint command to `bin/fm-lint.sh`, matching the Linux CI lint job, and pins its test command to `bin/fm-test-run.sh --changed`. -Local no-mistakes Test is intent-targeted and must not re-run every `tests/*.test.sh`; `.github/workflows/ci.yml` owns the broad behavior suite plus platform-specific compatibility lanes. +[`docs/configuration.md`](docs/configuration.md#gate-defaults-no-mistakesyaml) owns the tracked `.no-mistakes.yaml` gate defaults. +The `firstmate-coding-guidelines` skill owns the rule that local no-mistakes Test stays intent-targeted rather than configuring `commands.test`. Verify the same way the gate does: reach for `bin/fm-test-run.sh` with the subjects you care about rather than chaining `bash tests/a.test.sh && bash tests/b.test.sh`, because a list of script paths gets the same bounded concurrency as `--changed`. The pipeline publishes that evidence itself, so never hand-commit `.no-mistakes/` paths onto a feature branch; CI rejects them as tracked personal fleet paths. @@ -89,10 +89,11 @@ tmp=$(mktemp -d) && printf 'done: smoke\n' > "$tmp/smoke.status" && FM_STATE_OVE ``` `bin/fm-test-run.sh` is the single owner of behavior-suite selection, portable CI lane composition, bounded concurrency admission, per-script timing markers, family totals, the coverage guard, and the optional JSON timing artifact. +Suites registered with `fm_test_run_cases` support the shared named-case and case-listing interface documented in `tests/lib.sh`; use it through the behavior runner when investigating one failure. +The public [`reconcile-firstmate-upstream`](skills/reconcile-firstmate-upstream/SKILL.md) skill owns the bounded, resumable local-check and early-PR workflow for upstream reconciliation, while existing GitHub Actions lanes retain broad coverage and merge-readiness checks. Its header and `--help` own the flags, family labels, lanes, and changed-file map; this section only documents the entry points. `bin/fm-test-isolation-proof.sh` remains the single owner of the portable candidate proof and reusable family proof harness; see `docs/fm-test-isolation-proof.md`. Portable shard balance evidence lives in `docs/fm-test-portable-shards.md`. -Local no-mistakes Test stays intent-targeted and must not wire `commands.test` to `--all` or a `tests/*.test.sh` walk. Family selection is the ordinary local path; `--all` is deliberate full regression only. CI owns broad regression across required portable parallel shards, the portable serial lane's separate-runner shards, the Herdr lane, lint, invariants, the coverage guard, and stock macOS Bash compatibility in [`.github/workflows/ci.yml`](.github/workflows/ci.yml). CI also installs pinned Pi and OpenCode packages in a non-credentialed compatibility job so package API drift and missing runner CLIs fail pull requests explicitly. diff --git a/README.md b/README.md index 3f6a35bef7f..a5fee090d5c 100644 --- a/README.md +++ b/README.md @@ -40,7 +40,7 @@ firstmate is not a model, not a harness, not a skill, not an MCP server, and not firstmate is an agent distro for running a crew of agents. An agent distro is a portable directory of instructions, skills, tooling, policies, and state conventions that turns a general-purpose agent into a specialized one. There is no app to install: the cloned repo is the distro - `AGENTS.md`, bundled firstmate skills, and helper scripts that any terminal coding agent can follow. -Launching a supported harness inside it instantiates your first mate - and makes you the captain. +Launching a supported harness inside it for your primary session instantiates your first mate - and makes you the captain. firstmate runs on Linux, macOS, and Windows. GitHub Copilot CLI is supported for both the primary first mate session and crewmate workers, alongside the other [supported harnesses](#recommended-harnesses). @@ -48,7 +48,7 @@ GitHub Copilot CLI is supported for both the primary first mate session and crew ## Features - **One liaison** - you talk only to the first mate; it dispatches, supervises, escalates only real decisions, and reports plain outcomes. -- **A visible crew** - every crewmate works in its own tmux window, experimental herdr/zellij tab, cmux workspace, or Orca terminal you can watch or type into; the first mate reconciles. +- **A visible crew** - every crewmate works in its own tmux window, Herdr tab, or experimental zellij tab, cmux workspace, or Orca terminal you can watch or type into; the first mate reconciles. - **Disposable worktrees** - each task runs in a clean [treehouse](https://github.com/kunchenguid/treehouse) git worktree, or an Orca-managed worktree when `backend=orca`, so parallel work on one repo never collides. - **Two task shapes** - ship tasks deliver authorized changes; scout tasks leave standalone investigation reports when the intake contract warrants separate research. - **Explicit project modes** - each project ships via `no-mistakes`, `direct-PR`, or `local-only`, with an optional `+yolo` merge-autonomy flag. @@ -65,7 +65,7 @@ Full detail on every feature lives in [docs/architecture.md](docs/architecture.m ### Requirements - Linux, macOS, or Windows; Windows setup uses the PowerShell installer below. -- A verified primary agent harness: Claude Code, GitHub Copilot CLI, Grok, Pi, `pi-signed`, Codex, OpenCode, or Cursor Agent CLI. +- A verified primary agent harness: Claude Code, GitHub Copilot CLI, Grok, Pi, `pi-signed`, Oh My Pi (`omp`), Codex, OpenCode, or Cursor Agent CLI. - Git and the GitHub CLI, authenticated through `gh auth login`. - The CLI and dependencies for your selected runtime backend; tmux is the reference default. @@ -79,6 +79,7 @@ Claude Code uses a tracked Stop hook for tokenless watcher re-arm and rewake, Gr All three have verified turn-end guard paths when launched with their documented setup. Pick whichever one matches your subscription and workflow. +Oh My Pi (`omp`), a Pi fork, is verified as a primary with the same extension-owned watcher model as Pi and a stronger turn-end guard: its blocking `session_stop` hook compels a continuation instead of requesting one. Codex and OpenCode are also verified and supported as primary harnesses; Codex uses bounded foreground checkpoints, and OpenCode uses a TUI plugin, so both carry more harness-specific supervision tradeoffs than the three co-primaries. Cursor Agent CLI is verified as a primary too, using a tracked project-scope `.cursor/hooks.json` whose `stop` hook parks on the watcher between turns, closest in shape to Claude Code's. Launch it with `--trust`, or none of its project hooks load; it also has no turn-end hook in headless `cursor-agent -p`, so run the primary session interactively. @@ -129,6 +130,16 @@ pi FM_PI_HARNESS=pi-signed pi-signed ``` +**Oh My Pi** + +```sh +omp +# or, when starting from inside a Claude Code pane +FM_OMP_HARNESS=omp omp +``` + +Start `omp` with this checkout as its working directory: it auto-discovers the tracked `.omp/extensions/*.ts` files with no trust dialog, and naming them with `-e` as well would load each twice. + For Grok, `--trust` is needed once per clone so project hooks and the turn-end guard load; `/hooks-trust` inside Grok works too. For Pi, approve the project trust prompt once per clone on first launch so the tracked `.pi/extensions/*.ts` files auto-load. Pi's `/calm` toggle hides supported transcript chrome, including canonically classified Firstmate operational user rows, and uses a Calm-only animated working boat during active runs while preserving all model context and session data. @@ -231,7 +242,7 @@ Firstmate's skills live in two separate places with different audiences: - [docs/voice-relay.md](docs/voice-relay.md) - the optional spoken interface: setup on both machines, measured round-trip cost, what a spoken answer may read, and what this build does not do yet. - [docs/wedge-alarm.md](docs/wedge-alarm.md) - configure the active alert for an away-mode escalation delivery that gets stuck. - [docs/tmux-backend.md](docs/tmux-backend.md) - current setup and limits for the tmux reference backend. -- [docs/herdr-backend.md](docs/herdr-backend.md) - current setup, safety boundaries, and limits for the experimental Herdr backend. +- [docs/herdr-backend.md](docs/herdr-backend.md) - current setup, CI coverage, safety boundaries, and limits for the Herdr backend. - [docs/zellij-backend.md](docs/zellij-backend.md) - current setup and limits for the experimental Zellij backend. - [docs/orca-backend.md](docs/orca-backend.md) - current setup and limits for the experimental Orca backend. - [docs/cmux-backend.md](docs/cmux-backend.md) - current setup, socket security, and limits for the experimental cmux backend. @@ -240,10 +251,10 @@ Firstmate's skills live in two separate places with different audiences: - [docs/gitlab-merge-watch.md](docs/gitlab-merge-watch.md) - maintainer verification for watching and merging GitLab merge requests on arbitrary instances. - [docs/turnend-guard.md](docs/turnend-guard.md) - the primary session's current "no turn ends blind" backstop, scope, loop safety, and compatibility limits. - [docs/verification/supervision.md](docs/verification/supervision.md) - active maintainer verification for session-start, guard, continuity, and wedge integrations. -- [docs/supervision-protocols/](docs/supervision-protocols/) - rendered primary-harness watcher protocols for Claude, Codex, GitHub Copilot CLI, OpenCode, Pi and `pi-signed`, Grok, Cursor, and unknown harness fallback. +- [docs/supervision-protocols/](docs/supervision-protocols/) - rendered primary-harness watcher protocols for Claude, Codex, GitHub Copilot CLI, OpenCode, Pi and `pi-signed`, omp, Grok, Cursor, and unknown harness fallback. - [docs/scripts.md](docs/scripts.md) - the `bin/` toolbelt reference. - [docs/documentation-audiences.md](docs/documentation-audiences.md) - documentation audiences and the machine-checked placement boundary. -- [`AGENTS.md`](AGENTS.md) - the distro's always-loaded operating contract and routing index for conditional procedures. +- [`AGENTS.md`](AGENTS.md) - the supervisor contract, role boundary, and routing index for conditional procedures. - [CONTRIBUTING.md](CONTRIBUTING.md) - how to contribute, including the dev/test commands. ## Contributing diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index 8713e79fac8..c26731dbde2 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -671,7 +671,7 @@ fm_backend_herdr_presentation_lock_namespace() { fm_backend_herdr_presentation_lock_namespace_mode() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - stat -f '%Lp' "$1" 2>/dev/null + /usr/bin/stat -f '%Lp' "$1" 2>/dev/null else stat -c '%a' "$1" 2>/dev/null fi @@ -679,7 +679,7 @@ fm_backend_herdr_presentation_lock_namespace_mode() { fm_backend_herdr_presentation_lock_namespace_uid() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - stat -f '%u' "$1" 2>/dev/null + /usr/bin/stat -f '%u' "$1" 2>/dev/null else stat -c '%u' "$1" 2>/dev/null fi diff --git a/bin/backends/tmux.sh b/bin/backends/tmux.sh index eb7a0106829..70ad8fc4925 100644 --- a/bin/backends/tmux.sh +++ b/bin/backends/tmux.sh @@ -172,7 +172,10 @@ fm_backend_tmux_classify_process_name() { # [argv0] -> agent|shell|other # cannot carry it either: ~/.local/bin/muse-bin- has no `muse` path # COMPONENT, so the fm_harness_path_name fallback below never fires for it. muse|muse-bin-*) printf 'agent' ;; - *claude*|*codex*|copilot|copilot.exe|*opencode*|*grok*|*kimi*|pi|pi-signed|pi-launcher|Pi) printf 'agent' ;; + # omp (Oh My Pi) is anchored for the same reason as muse: its live process + # name is the bare word `omp` (verified, omp 18.1.11) and a glob would claim + # unrelated commands such as ompd or comp. + *claude*|*codex*|copilot|copilot.exe|*opencode*|*grok*|*kimi*|*rovo*|pi|pi-signed|pi-launcher|Pi|omp) printf 'agent' ;; zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; *) if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then diff --git a/bin/fm-backlog-receive.sh b/bin/fm-backlog-receive.sh index 15d9bde99ae..46cf06783fd 100755 --- a/bin/fm-backlog-receive.sh +++ b/bin/fm-backlog-receive.sh @@ -57,7 +57,7 @@ list_keys() { # lock_age() { local modified now if [ "$(uname 2>/dev/null)" = Darwin ]; then - modified=$(stat -f '%m' "$1" 2>/dev/null) || return 1 + modified=$(/usr/bin/stat -f '%m' "$1" 2>/dev/null) || return 1 else modified=$(stat -c '%Y' "$1" 2>/dev/null) || return 1 fi diff --git a/bin/fm-backlog-transition-lib.sh b/bin/fm-backlog-transition-lib.sh index 4946bdd716e..42b6d8e8642 100644 --- a/bin/fm-backlog-transition-lib.sh +++ b/bin/fm-backlog-transition-lib.sh @@ -22,20 +22,22 @@ # # SCOPE. fm_backlog_transition_applies is the single gate. It excludes # secondmates (persistent agents are never backlog items, AGENTS.md section 10), -# homes whose configured backlog backend is manual and homes that keep no -# backlog file at all. Those return-1 exemptions are never errors; an -# unresolvable configured data directory or incompatible tasks-axi instead -# returns 2 so callers refuse before mutation. +# and homes whose configured backlog backend is manual. Markdown homes that +# keep no backlog file at all are likewise exempt. Those return-1 exemptions +# are never errors; a home on any other configured backend has no markdown +# file requirement at all. An unresolvable configured data directory or +# incompatible tasks-axi instead returns 2 so callers refuse before mutation. # -# ADDRESSING. Every mutation call passes `--file /backlog.md` so the -# change lands in the home that owns the task regardless of the caller's -# working directory, and runs from that data directory's parent so the same -# home's `.tasks.toml` supplies done_keep and the archive path. Row probes pass -# `--file` only for the markdown backend and otherwise run from the addressing -# root so backend-owned state remains discoverable. The parent of the data -# directory is the addressing root rather than FM_HOME, so a home whose data -# directory is relocated keeps its backlog and its archive together. A root -# with no `.tasks.toml` gets tasks-axi's built-in defaults. +# ADDRESSING. The markdown backend owns the explicit `--file /backlog.md` +# on every mutation and row probe so the change lands in the home that owns +# the task regardless of the caller's working directory. Every other +# configured backend is addressed from that data directory's parent - the +# addressing root - with no markdown file override, so the same home's +# `.tasks.toml` supplies its backend, done_keep, and the archive path and +# backend-owned state remains discoverable. The parent of the data directory +# is the addressing root rather than FM_HOME, so a home whose data directory +# is relocated keeps its backlog and its archive together. A root with no +# `.tasks.toml` gets tasks-axi's built-in defaults. # # CRASH RECOVERY. Only teardown needs a durable record: it removes the meta and # with it the completion links, so a process killed between the two halves would @@ -175,7 +177,7 @@ fm_backlog_data_relative() { # } fm_backlog_transition_applies() { # - local config=$1 data authorized_data=$2 kind=$3 file + local config=$1 data authorized_data=$2 kind=$3 file root FM_BACKLOG_TRANSITION_SKIP= if [ "$kind" = secondmate ]; then FM_BACKLOG_TRANSITION_SKIP="secondmates are not backlog items" @@ -189,13 +191,16 @@ fm_backlog_transition_applies() { # FM_BACKLOG_TRANSITION_ERROR="data directory cannot be resolved: $2" return 2 fi - file=$(fm_backlog_file "$data") - if [ ! -e "$file" ] && [ ! -L "$file" ]; then - FM_BACKLOG_TRANSITION_SKIP="this home keeps no backlog at $file" - return 1 - fi - if ! fm_backlog_record_present "$file" "backlog file" "$authorized_data"; then - return 2 + root=$(fm_backlog_root "$data") || return 2 + if [ "$(fm_tasks_axi_backend "$root")" = markdown ]; then + file=$(fm_backlog_file "$data") + if [ ! -e "$file" ] && [ ! -L "$file" ]; then + FM_BACKLOG_TRANSITION_SKIP="this home keeps no backlog at $file" + return 1 + fi + if ! fm_backlog_record_present "$file" "backlog file" "$authorized_data"; then + return 2 + fi fi if ! fm_tasks_axi_compatible; then FM_BACKLOG_TRANSITION_ERROR="automatic backlog transitions require tasks-axi $FM_TASKS_AXI_MIN or newer with the required update and mv features" @@ -204,6 +209,83 @@ fm_backlog_transition_applies() { # return 0 } +# Run `tasks-axi` with an optional FM_TASKS_AXI_TIMEOUT bound. A caller that +# holds a lock across the call - the spawn commit and its preservation +# read-back run under the per-task meta lock - sets the bound, so an +# unresponsive tasks-axi cannot hold that lock open indefinitely; a timed-out +# call exits 124, or 137 when the kill-after had to fire (GNU timeout's own +# status for a KILL-forced expiry), and the callers treat either as the bound +# expiring and report the timeout as the reason through their existing error +# plumbing. GNU timeout is used where it exists, +# gtimeout where coreutils ships under that name, and a small perl watchdog +# elsewhere (a stock macOS host has perl but no timeout variant; perl is +# already a hard dependency of this library's byte validators, so the +# fallback adds no new tool). Every bounded path forces termination: a +# tasks-axi that ignores SIGTERM must not outlive the bound, since an +# unbounded call under the lock is exactly the hang the bound exists to +# prevent - so the GNU variants carry a kill-after of one further bound +# (TERM at the bound, KILL after that grace) and the watchdog kills the +# same way. When a bound was requested but no bounding mechanism exists at +# all, the call fails closed instead of running unbounded. Must be the last +# command of a subshell: the exec keeps the tasks-axi process exactly where +# the plain call sat, and the bound kills the child, not the caller. +fm_tasks_axi_timeout_expired() { # + case $1 in + 124 | 137) return 0 ;; + esac + return 1 +} + +fm_tasks_axi() { + local bound=${FM_TASKS_AXI_TIMEOUT:-} + if [ -z "$bound" ]; then + exec tasks-axi "$@" + fi + if command -v timeout >/dev/null 2>&1; then + exec timeout -k "$bound" "$bound" tasks-axi "$@" + elif command -v gtimeout >/dev/null 2>&1; then + exec gtimeout -k "$bound" "$bound" tasks-axi "$@" + elif command -v perl >/dev/null 2>&1; then + # Fork, run tasks-axi in the child, and poll waitpid(WNOHANG) until the + # child exits or the bound expires: the same contract as + # `timeout $bound tasks-axi ...`. Expiry kills the child with TERM, waits + # one further bound of grace, then KILL, and exits 124 so the callers' + # timeout plumbing reports it. Polling rather than alarm+die keeps the + # bound off perl's platform-dependent syscall-restart signal semantics. + exec perl -MPOSIX=WNOHANG -e ' + my $bound = shift; + exit 127 unless defined $bound && $bound =~ /\A[0-9]+\z/; + my $pid = fork; + exit 127 unless defined $pid; + if ($pid == 0) { exec @ARGV; exit 127 } + my $step = 0.05; + my $elapsed = 0; + while (1) { + my $done = waitpid $pid, WNOHANG; + exit(($? & 127) ? 128 + ($? & 127) : $? >> 8) if $done == $pid; + exit 127 if $done == -1; + if ($elapsed >= $bound) { + kill "TERM", $pid; + my $grace = 0; + my $gone = waitpid $pid, WNOHANG; + while ($gone == 0 && $grace < $bound) { + select undef, undef, undef, $step; + $grace += $step; + $gone = waitpid $pid, WNOHANG; + } + kill "KILL", $pid if $gone == 0; + waitpid $pid, 0; + exit 124; + } + select undef, undef, undef, $step; + $elapsed += $step; + } + ' -- "$bound" tasks-axi "$@" + fi + printf 'fm_tasks_axi: cannot bound tasks-axi within %ss: none of timeout, gtimeout, or perl is available\n' "$bound" >&2 + exit 127 +} + # Print one row's `tasks-axi show` output (plus stderr) from the backlog root, # with `--file` only for the markdown backend; the exit status is tasks-axi's. # Extra flags (such as --full) are passed through. @@ -213,9 +295,9 @@ fm_backlog_row_show() { # [flag...] file=$(fm_backlog_file "$data") || return 1 root=$(fm_backlog_root "$data") || return 1 if [ "$(fm_tasks_axi_backend "$root")" = markdown ]; then - (cd "$root" 2>/dev/null && tasks-axi show "$id" "$@" --file "$file" 2>&1) + (cd "$root" 2>/dev/null && fm_tasks_axi show "$id" "$@" --file "$file" 2>&1) else - (cd "$root" 2>/dev/null && tasks-axi show "$id" "$@" 2>&1) + (cd "$root" 2>/dev/null && fm_tasks_axi show "$id" "$@" 2>&1) fi } @@ -232,7 +314,7 @@ fm_backlog_row_list() { # [flag...] } fm_backlog_row_probe() { # - local data authorized_data=$1 file id=$2 out state held blocked hold_kind command_status + local data authorized_data=$1 file id=$2 out state held blocked hold_kind command_status root if ! data=$(fm_backlog_data_absolute "$1"); then FM_BACKLOG_ROW_RESULT=error FM_BACKLOG_ROW_STATE= @@ -243,18 +325,20 @@ fm_backlog_row_probe() { # FM_BACKLOG_ROW_STATE= FM_BACKLOG_ROW_HOLD_KIND= FM_BACKLOG_ROW_ERROR= - file=$(fm_backlog_file "$data") || { + root=$(fm_backlog_root "$data") || { FM_BACKLOG_ROW_ERROR=$FM_BACKLOG_TRANSITION_ERROR return 1 } - if ! fm_backlog_record_present "$file" "backlog file" "$authorized_data"; then - FM_BACKLOG_ROW_ERROR=$FM_BACKLOG_TRANSITION_ERROR - return 1 + if [ "$(fm_tasks_axi_backend "$root")" = markdown ]; then + file=$(fm_backlog_file "$data") || { + FM_BACKLOG_ROW_ERROR=$FM_BACKLOG_TRANSITION_ERROR + return 1 + } + if ! fm_backlog_record_present "$file" "backlog file" "$authorized_data"; then + FM_BACKLOG_ROW_ERROR=$FM_BACKLOG_TRANSITION_ERROR + return 1 + fi fi - fm_backlog_root "$data" >/dev/null || { - FM_BACKLOG_ROW_ERROR=$FM_BACKLOG_TRANSITION_ERROR - return 1 - } out=$(fm_backlog_row_show "$data" "$id") command_status=$? if [ "$command_status" -ne 0 ]; then @@ -262,8 +346,13 @@ fm_backlog_row_probe() { # FM_BACKLOG_ROW_RESULT=not_found else FM_BACKLOG_ROW_ERROR=$(printf '%s\n' "$out" | sed -n '1p') - [ -n "$FM_BACKLOG_ROW_ERROR" ] \ - || FM_BACKLOG_ROW_ERROR="tasks-axi show $id failed with no output" + if [ -z "$FM_BACKLOG_ROW_ERROR" ]; then + if fm_tasks_axi_timeout_expired "$command_status" && [ -n "${FM_TASKS_AXI_TIMEOUT:-}" ]; then + FM_BACKLOG_ROW_ERROR="tasks-axi show $id did not finish within ${FM_TASKS_AXI_TIMEOUT}s" + else + FM_BACKLOG_ROW_ERROR="tasks-axi show $id failed with no output" + fi + fi fi return "$command_status" fi @@ -285,24 +374,38 @@ fm_backlog_row_probe() { # } # Run one tasks-axi mutation against 's backlog, capturing its first -# output line in FM_BACKLOG_TRANSITION_ERROR on failure. +# output line in FM_BACKLOG_TRANSITION_ERROR on failure. The markdown backend +# keeps its explicit /backlog.md file and presence requirement; every +# other configured backend is addressed by the root's own tasks-axi +# configuration, so passing the markdown-era path would write the wrong store. fm_backlog_mutate() { # [flag...] - local data authorized_data=$1 file verb=$2 id=$3 out command_status + local data authorized_data=$1 file verb=$2 id=$3 out command_status root if ! data=$(fm_backlog_data_absolute "$1"); then FM_BACKLOG_TRANSITION_ERROR="data directory cannot be resolved: $1" return 1 fi shift 3 FM_BACKLOG_TRANSITION_ERROR= - file=$(fm_backlog_file "$data") || return 1 - fm_backlog_record_present "$file" "backlog file" "$authorized_data" || return 1 - out=$(cd "$(fm_backlog_root "$data")" 2>/dev/null && tasks-axi "$verb" "$id" \ - --file "$file" "$@" 2>&1) - command_status=$? + root=$(fm_backlog_root "$data") || return 1 + if [ "$(fm_tasks_axi_backend "$root")" != markdown ]; then + out=$(cd "$root" 2>/dev/null && fm_tasks_axi "$verb" "$id" "$@" 2>&1) + command_status=$? + else + file=$(fm_backlog_file "$data") || return 1 + fm_backlog_record_present "$file" "backlog file" "$authorized_data" || return 1 + out=$(cd "$root" 2>/dev/null && fm_tasks_axi "$verb" "$id" \ + --file "$file" "$@" 2>&1) + command_status=$? + fi [ "$command_status" -ne 0 ] || return 0 FM_BACKLOG_TRANSITION_ERROR=$(printf '%s\n' "$out" | sed -n '1p') - [ -n "$FM_BACKLOG_TRANSITION_ERROR" ] \ - || FM_BACKLOG_TRANSITION_ERROR="tasks-axi $verb $id failed with no output" + if [ -z "$FM_BACKLOG_TRANSITION_ERROR" ]; then + if fm_tasks_axi_timeout_expired "$command_status" && [ -n "${FM_TASKS_AXI_TIMEOUT:-}" ]; then + FM_BACKLOG_TRANSITION_ERROR="tasks-axi $verb $id did not finish within ${FM_TASKS_AXI_TIMEOUT}s" + else + FM_BACKLOG_TRANSITION_ERROR="tasks-axi $verb $id failed with no output" + fi + fi return "$command_status" } diff --git a/bin/fm-bearings-board.sh b/bin/fm-bearings-board.sh index 2c302fcd395..6a8d82c7e2e 100755 --- a/bin/fm-bearings-board.sh +++ b/bin/fm-bearings-board.sh @@ -11,22 +11,58 @@ # fm-bearings-board.sh build # fm-bearings-board.sh path # -# build Validate the payload and inject it into a fresh copy of the shipped -# template at the stable board path. Establish or resume the Lavish -# session on that board BEFORE binding and arming its answer source, -# so a registered poll can never race a session that does not exist. +# build Validate the payload, drop the Captain's Call cards whose subject +# already landed, give every surviving decision card the standard +# reconcile choice, and inject the result into a fresh copy of the +# shipped template at the stable board path. Establish the Lavish +# session on that board and PROVE it is live BEFORE binding and +# arming its answer source, so a registered poll can never race a +# session that does not exist or attach to one that has ended. # Bind to the keyed-answer intake (bin/fm-captain-hold.sh) ALWAYS # precedes arm, so the board can never produce an answer that has # nowhere to go (captain-hold-lifecycle's ordering rule, enforced # here rather than left to agent memory). Output starts with # `board: `, then includes lavish-axi's session output and # the remaining status: +# session: live | reopened # served: # bound: # armed: (first registration) # already-armed: (registration already present) +# listening: (only when a replacement was needed) +# Every dropped card is named on stderr as a `dropped-landed-card:` +# line, so a rebuild states what it removed instead of quietly +# shrinking Captain's Call. # path Print the stable board path for this home. # +# A LIVE SESSION IS PROVED, NEVER ASSUMED. `lavish-axi ` exits 0 even +# when it refuses to reopen a session the captain ended from the browser, +# reporting `status: user-ended` with the same session id, so exit status alone +# cannot tell a live board from a dead one. build requires the server's fresh +# session listing to show the canonical board open and refuses rather than +# arming an ended session. After a reopen it retires the pre-reopen source +# generation through the guarded adapter path, arms a fresh registration, and +# accepts only the replacement listener as live. A registered board with no +# live owner also gets a replacement before build returns, because +# `already-armed` is not the same fact as `listening`. +# +# CAPTAIN'S CALL HYGIENE. A decision card is dropped when its work item, PR, or +# structured artifact/version subject appears among the payload's own landed +# rows, or when `bin/fm-captain-hold.sh open` reports the task is no longer an +# open captain call. A newer published version also supersedes a version card. +# A task whose state cannot be established is kept, because a call wrongly +# hidden is worse than a card wrongly shown. Cleanup is therefore a normal +# rebuild effect rather than a committed migration or direct state mutation. +# +# THE RECONCILE CHOICE. Every decision card carries the standard `reconcile` +# option, injected here so the guarantee does not depend on the composer's +# memory, and the payload validator reserves that value across every card type. +# The validator's reservation scope must equal the adapter's reconcile +# classification scope, which is all card types because the captured payload +# carries no card type. Its meaning, and the reason it can never reach the +# keyed-answer intake as a blind close, are owned by +# docs/captain-hold-lifecycle.md. +# # Validation is fail-closed: the payload must be valid JSON with # schema=fm-bearings-board.v1 and every renderer-consumed field must satisfy # the fm-bearings-board.v1 types and item invariants below. Every fleet row and @@ -79,6 +115,14 @@ validate_payload() { # or (.[$name] | type == "string" and test("^https://[A-Za-z0-9](?:[A-Za-z0-9.-]*[A-Za-z0-9])?(?::[0-9]{1,5})?(?:[/?#][^[:space:]]*)?$")); + def version: type == "string" and test("^(0|[1-9][0-9]{0,8})\\.(0|[1-9][0-9]{0,8})\\.(0|[1-9][0-9]{0,8})$"); + def optional_subject: + (has("subject") | not) + or (.subject + | type == "object" + and (keys | sort) == ["artifact", "version"] + and (.artifact | slug(128)) + and (.version | version)); def call_item: type == "object" and (.key | slug(128)) @@ -96,12 +140,16 @@ validate_payload() { # and (optional_string("decide")) and (optional_string("detail")) and (optional_https_url("pr_url")) + and optional_subject + and (if has("subject") then .type == "decision" else true end) and (optional_string("freeform_hint")) and ((has("close") | not) or (.close == "done" or .close == "release")) and ((has("allow_freeform") | not) or (.allow_freeform | type == "boolean")) and ((has("recommend_value") | not) or ((.recommend_value | slug(128)) - and (.recommend_value as $recommend | [.options[].value] | index($recommend) != null))) + and (.recommend_value as $recommend + | ([.options[].value] | index($recommend) != null)))) + and ([.options[].value] | index("reconcile") == null) and (if .type == "merge" then (.risk | nonempty_string) else true end); def underway_item: type == "object" and repo_marker and (.id | nonempty_string) @@ -109,7 +157,8 @@ validate_payload() { # def landed_item: type == "object" and repo_marker and (.id | nonempty_string) and (.what | nonempty_string) and (.owner | nonempty_string) - and optional_https_url("pr_url"); + and optional_https_url("pr_url") + and optional_subject; def charted_item: type == "object" and repo_marker and (.id | slug(128)) and (.title | nonempty_string) and (.reason | type == "string") @@ -136,8 +185,161 @@ validate_payload() { # ' "$1" >/dev/null } +# --- Lavish session liveness ------------------------------------------------- +# Verified against lavish-axi 0.1.61. `lavish-axi ` EXITS 0 even when it +# refuses to reopen a session the captain ended from the browser, reporting +# `status: user-ended` and the same session id, so an exit-code check alone +# cannot tell a live board from a dead one. The establish status is an initial +# signal only; the server's fresh session listing must also show the canonical +# board open before the build may bind or arm its source. + +board_realpath() { # + perl -MCwd=realpath -e '$p = realpath($ARGV[0]); defined($p) or exit 1; print "$p\n"' "$1" 2>/dev/null +} + +lavish_status_field() { # + printf '%s\n' "$1" | sed -n 's/^[[:space:]]*status:[[:space:]]*//p' | head -1 | tr -d '"' +} + +# The server's own listing, keyed on the canonical artifact path. Rows are +# `,,"",`, and only a live session is listed `open`. +lavish_session_listed_open() { # + local listing + listing=$(lavish-axi 2>/dev/null) || return 1 + printf '%s\n' "$listing" | awk -v path="$1" ' + { line = $0; sub(/^[[:space:]]+/, "", line) } + index(line, path ",") == 1 { + rest = substr(line, length(path) + 2) + split(rest, field, ",") + if (field[1] == "open") { found = 1 } + } + END { exit found ? 0 : 1 } + ' +} + +lavish_board_live() { # + lavish_session_listed_open "$2" +} + +# Establish the board session and PROVE it is live before anything arms a poll +# on it. A session the captain ended is reopened once - the captain asked for +# this board, which is exactly the attention `--reopen` exists for - and a +# session that is still not live after that refuses the build rather than +# arming a poll that can never attach. +establish_board_session() { # + local board=$1 real out status version + BOARD_SESSION_REOPENED=0 + real=$(board_realpath "$board") || fail "cannot resolve the board path: $board" + out=$(lavish-axi "$board") || fail "cannot establish the board Lavish session" + printf '%s\n' "$out" + if lavish_board_live "$out" "$real"; then + printf 'session: live\n' + return 0 + fi + out=$(lavish-axi "$board" --reopen) || fail "cannot reopen the ended board Lavish session" + printf '%s\n' "$out" + if lavish_board_live "$out" "$real"; then + BOARD_SESSION_REOPENED=1 + printf 'session: reopened\n' + return 0 + fi + status=$(lavish_status_field "$out") + version=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') + fail "the board Lavish session is not live after reopening it (lavish-axi ${version:-version-unknown} reported status ${status:-none}); refusing to arm a poll on an ended session" +} + +# --- Captain's Call hygiene --------------------------------------------------- +# A held decision whose subject already shipped is not a live call, so it is +# dropped here instead of being carded again. All checks use exact structured +# identities; unknown subject state keeps the card. + +decision_card_is_stale() { # + local task=$1 landed=$2 rc=0 + if [ "$landed" = 1 ]; then + printf 'structured subject already landed\n' + return 0 + fi + "$SCRIPT_DIR/fm-captain-hold.sh" open "$task" --distinguish-absent >/dev/null 2>&1 || rc=$? + # 1 is a definite "no longer an open captain call". 2 is "cannot tell", 3 is + # absent from this backlog, and a call wrongly hidden is worse than a card + # wrongly shown, so both uncertain and absent cards stay. + if [ "$rc" -eq 1 ]; then + printf 'no longer an open captain call\n' + return 0 + fi + return 1 +} + +# Drop every stale decision card, then give every surviving decision card the +# standard reconcile choice. Injecting it here is what makes "every decision +# card offers reconcile" a property of the board rather than of the composer's +# memory; the validator prevents duplicate decision options. +effective_payload() { # + local data=$1 dest=$2 landed_keys key reason drop='' tmp landed=0 + landed_keys=$(jq -c ' + def version_parts: split(".") | map(tonumber); + . as $payload + | [$payload.captains_call[] + | select(.type == "decision") + | . as $card + | select( + ($payload.landed | any(.id == $card.key)) + or (($card.pr_url? != null) and ($payload.landed | any(.pr_url? == $card.pr_url))) + or (($card.subject? != null) and ($payload.landed | any( + (.subject? != null) + and (.subject.artifact == $card.subject.artifact) + and ((.subject.version | version_parts) >= ($card.subject.version | version_parts))))) + ) + | .key] + ' "$data") || return 1 + while IFS= read -r key; do + [ -n "$key" ] || continue + landed=0 + if jq -e --arg key "$key" 'index($key) != null' <<< "$landed_keys" >/dev/null; then + landed=1 + fi + reason=$(decision_card_is_stale "$key" "$landed") || continue + printf 'dropped-landed-card: %s (%s)\n' "$key" "$reason" >&2 + drop=$drop$key$'\n' + done < <(jq -r '.captains_call[]? | select(.type == "decision") | .key' "$data") + tmp=$(printf '%s' "$drop" | jq -R -s 'split("\n") | map(select(length > 0))') || return 1 + jq --argjson dropped "$tmp" ' + .captains_call = [ + .captains_call[] + | . as $card + | select($card.type != "decision" or (($dropped | index($card.key)) == null)) + | if .type == "decision" + then .options += [{ + value: "reconcile", + label: "Reconcile", + hint: "Re-check the latest state, then close this with evidence or keep it open with a note" + }] + else . end + ]' "$data" > "$dest" || return 1 +} + +# The OWNER column bin/fm-procevent.sh already publishes: live, none, +# orphaned, or uncertain. Empty means the source is not registered at all. +source_owner() { # + "$SCRIPT_DIR/fm-procevent.sh" list 2>/dev/null \ + | awk -v id="$1" 'NR > 1 && $1 == id { print $3 }' +} + +# A replacement listener is started detached, so it claims the source shortly +# after reconcile returns. Wait for that claim rather than reporting the race. +await_source_owner() { # + local owner i=0 + while [ "$i" -lt 50 ]; do + owner=$(source_owner "$1") + [ "$owner" != live ] || { printf '%s\n' "$owner"; return 0; } + sleep 0.1 + i=$((i + 1)) + done + printf '%s\n' "${owner:-none}" +} + command_build() { - local data=${1-} board json tmp sid extracted + local data=${1-} board json tmp sid extracted effective owner version pre_reopen_owner [ "$#" -eq 1 ] || { usage >&2; exit 2; } command -v jq >/dev/null 2>&1 || fail "jq is required" [ -f "$data" ] || fail "board data does not exist: $data" @@ -147,7 +349,14 @@ command_build() { [ "$(grep -cxF "$PLACEHOLDER" "$TEMPLATE")" -eq 1 ] \ || fail "board template does not carry exactly one data slot: $TEMPLATE" - json=$(jq -c . "$data") || fail "cannot compact the board data" + effective=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-bearings-payload.XXXXXX") \ + || fail "cannot stage the board payload" + if ! effective_payload "$data" "$effective"; then + rm -f -- "$effective" + fail "cannot reconcile the board payload against landed work" + fi + json=$(jq -c . "$effective") || { rm -f -- "$effective"; fail "cannot compact the board data"; } + rm -f -- "$effective" # `<` never appears in JSON syntax outside strings, so escaping every # occurrence keeps the payload valid JSON while making inert. json=${json///dev/null 2>&1 || fail "lavish-axi is not installed" - lavish-axi "$board" || fail "cannot establish the board Lavish session" - printf 'served: %s\n' "$board" - sid=$("$SCRIPT_DIR/fm-procevent-lavish.sh" source-id "$board") \ || fail "cannot derive the board source id" + pre_reopen_owner=$(source_owner "$sid") + establish_board_session "$board" + if [ "$BOARD_SESSION_REOPENED" = 1 ]; then + "$SCRIPT_DIR/fm-procevent-lavish.sh" retire "$board" >/dev/null \ + || fail "cannot retire the pre-reopen source generation (observed owner: ${pre_reopen_owner:-none})" + fi + if ! lavish_session_listed_open "$(board_realpath "$board")"; then + version=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') + fail "the board Lavish session is not listed open immediately before arming (lavish-axi ${version:-version-unknown}); refusing to arm a poll on observed state not-open" + fi + printf 'served: %s\n' "$board" + "$SCRIPT_DIR/fm-captain-hold.sh" bind "$sid" >/dev/null \ || fail "cannot bind the board source to the keyed-answer intake" printf 'bound: %s\n' "$sid" - if "$SCRIPT_DIR/fm-procevent.sh" list | awk 'NR > 1 { print $1 }' | grep -Fxq "$sid"; then + owner=$(source_owner "$sid") + if [ "$BOARD_SESSION_REOPENED" = 1 ]; then + "$SCRIPT_DIR/fm-procevent-lavish.sh" arm "$board" >/dev/null \ + || fail "cannot arm a fresh board source after reopening" + printf 'armed: %s\n' "$sid" + owner=$(source_owner "$sid") + elif [ -n "$owner" ]; then printf 'already-armed: %s\n' "$sid" else "$SCRIPT_DIR/fm-procevent-lavish.sh" arm "$board" >/dev/null \ || fail "cannot arm the board as a process-event source" printf 'armed: %s\n' "$sid" + owner=$(source_owner "$sid") + fi + # Registered is not listening. A board whose source has no live owner gets a + # replacement started now rather than at the next supervision cycle, which is + # what keeps a rebuilt board from sitting silent behind `already-armed`. + if [ "$owner" != live ]; then + "$SCRIPT_DIR/fm-procevent.sh" reconcile >/dev/null 2>&1 || true + owner=$(await_source_owner "$sid") + if [ "$owner" != live ]; then + fail "source $sid is not listening after reconcile (observed owner: ${owner:-none})" + fi + printf 'listening: live\n' fi } diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 1931f1defa1..6ba09493d58 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -792,7 +792,7 @@ secondmate_liveness_one() { # [ -n "$target" ] || target="$window" agent_state=$(fm_backend_agent_state "$backend" "$target" 2>/dev/null) || agent_state=unreadable case "$harness" in - claude|codex|copilot|opencode|pi|pi-signed|grok|kimi) ;; + claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|omp) ;; *) case "$agent_state" in dead|missing) agent_state=unverified-harness ;; esac ;; @@ -938,7 +938,7 @@ x_mode_write_if_changed() { [ "$parent" != "$dest" ] || return 1 [ -d "$parent" ] && [ ! -L "$parent" ] || return 1 if [ "$(uname)" = Darwin ]; then - parent_device=$(stat -f %d "$parent" 2>/dev/null) || return 1 + parent_device=$(/usr/bin/stat -f %d "$parent" 2>/dev/null) || return 1 else parent_device=$(stat -c %d "$parent" 2>/dev/null) || return 1 fi @@ -1100,7 +1100,7 @@ crew_dispatch_validate() { return 0 fi err=$(jq -r ' - def verified($h): ["claude","codex","copilot","opencode","pi","pi-signed","grok","kimi","cursor","muse"] | index($h); + def verified($h): ["claude","codex","copilot","opencode","pi","pi-signed","grok","kimi","cursor","muse","rovo","omp"] | index($h); def effort_ok($h; $e): if $e == null then true elif ($e | type) != "string" then false @@ -1108,8 +1108,9 @@ crew_dispatch_validate() { elif $h == "codex" then (["low","medium","high","xhigh"] | index($e)) elif $h == "copilot" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "grok" then (["low","medium","high"] | index($e)) - elif $h == "pi" or $h == "pi-signed" then (["low","medium","high","xhigh","max"] | index($e)) + elif $h == "pi" or $h == "pi-signed" or $h == "omp" then (["low","medium","high","xhigh","max"] | index($e)) elif $h == "muse" then (["low","medium","high","xhigh","max"] | index($e)) + elif $h == "rovo" then (["low","medium","high","max"] | index($e)) elif $h == "opencode" or $h == "kimi" or $h == "cursor" then false else true end; diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index b45a8f5ea78..5cb0a80f5e3 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -64,6 +64,9 @@ # it carries the AGENTS.md authoring bar (widely useful knowledge only, pointers # over copied detail) and defers self-governance recognition and insertion to # fm-ensure-agents-md.sh's contract. +# Scaffolds carry no role scope: fm-spawn.sh supplies fm_brief_worker_role from +# fm-dod-lib.sh to every ship/scout launch brief, so this file never becomes a +# second owner of a contract that must stay current across relaunches. # Refuses to overwrite an existing brief. set -eu @@ -387,8 +390,17 @@ The report is the only thing that survives, so anything worth keeping must be in A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work. Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. 7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving - every lane/home, so restarting it kills other lanes' in-flight pipeline runs. On ANY no-mistakes - daemon error, append \`blocked: {the daemon error}\` and stop; only firstmate manages the daemon. + every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate + manages the daemon. + Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and + \`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append + \`blocked: {the daemon error}\` and stop even when the local run record still says running or + fixing, because that record can be stale after the daemon exits. A run record failed with a + daemon error is also a real block. + Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep + going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error: + the daemon accepts \`respond\` immediately and runs the round in the background, so a killed or + timed-out call was only waiting for a read while the run kept working. $INBOX_SECTION @@ -469,8 +481,17 @@ $ASK_USER_BLOCK A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work. Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=]\` if you opened it with one) as you resume. 7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving - every lane/home, so restarting it kills other lanes' in-flight pipeline runs. On ANY no-mistakes - daemon error, append \`blocked: {the daemon error}\` and stop; only firstmate manages the daemon. + every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate + manages the daemon. + Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and + \`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append + \`blocked: {the daemon error}\` and stop even when the local run record still says running or + fixing, because that record can be stale after the daemon exits. A run record failed with a + daemon error is also a real block. + Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep + going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error: + the daemon accepts \`respond\` immediately and runs the round in the background, so a killed or + timed-out call was only waiting for a read while the run kept working. $INBOX_SECTION diff --git a/bin/fm-busy-event.sh b/bin/fm-busy-event.sh index 0abcab8ee39..17fde457810 100755 --- a/bin/fm-busy-event.sh +++ b/bin/fm-busy-event.sh @@ -103,7 +103,7 @@ LOCK="$REC.lock" # gets a non-numeric token. Detect the platform once and pick the right form, # exactly as bin/fm-watch.sh does. if [ "$(uname)" = Darwin ]; then - lock_mtime() { stat -f %m "$1" 2>/dev/null; } + lock_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else lock_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-busy-lib.sh b/bin/fm-busy-lib.sh index 3b5d23c2494..eb2012d5cfd 100755 --- a/bin/fm-busy-lib.sh +++ b/bin/fm-busy-lib.sh @@ -29,6 +29,7 @@ # task's recorded harness classifies unknown, so one adapter's writer can # never classify another adapter): # pi-ext Pi/pi-signed per-task extension (agent_start/agent_settled) +# omp-ext omp (Oh My Pi) per-task extension (agent_start/agent_end without willContinue) # opencode-plugin OpenCode per-task plugin (session.status) # claude-hook Claude lifecycle hooks (UserPromptSubmit/Stop/StopFailure/SessionEnd) # gemini-hook Gemini agent hooks (BeforeAgent opens; AfterAgent and @@ -42,7 +43,7 @@ # fm-interrupt the legacy Claude fm-send --key Escape idle event # fm-recovery a documented recovery reset after relaunch # Classifier-only sources (never written into a record): -# endpoint-gone, herdr-native, grok-regex, muse-session-log, +# endpoint-gone, herdr-native, grok-regex, rovo-regex, muse-session-log, # cursor-transcript, missing, malformed, gen-mismatch, source-mismatch, # kimi-unverified, codex-unverified, capture-failed, no-target # @@ -53,16 +54,18 @@ # 3. a valid, gen-matching, source-trusted record -> its state and source # 4. no record at all: herdr's native busy verdict is trusted as busy # (generation state is sufficient for busy, not for idle), then the -# muse session-log and cursor transcript pull sources, then the Grok-only -# temporary regex fallback classifies a grok task from its rendered tail, -# then unknown missing +# muse session-log and cursor transcript pull sources, then the Grok/Rovo +# temporary regex fallbacks classify a grok or rovo task from its +# rendered tail, then unknown missing # 5. malformed, stale, or untrusted records -> unknown, never a fallback -# The Grok arm is the ONLY rendered-text classification that survives the -# redesign, because Grok's structured lifecycle was not credited-live-verified -# in the approved audit; it is scoped to harness=grok and can never classify -# another adapter. The delivery guards in bin/fm-composer-lib.sh match rendered -# footers for submit acknowledgement and away-mode supervisor injection only; -# neither is a recorded worker state source. +# Grok and Rovo are the ONLY rendered-text classifications that survive the +# redesign, because neither's structured lifecycle was credited-live-verified +# in the approved audit (Rovo's clean ACP stopReason lives outside the TUI +# path firstmate drives, see references/harness/rovo.md); each is scoped to +# its own harness= and can never classify another adapter. The delivery +# guards in bin/fm-composer-lib.sh match rendered footers for submit +# acknowledgement and away-mode supervisor injection only; neither is a +# recorded worker state source. # # The muse pull source is semantic, not rendered: it folds muse's own durable # session event log. It has no writer, no arm, and no gen, because @@ -197,6 +200,7 @@ fm_busy_sources_for_harness() { # opencode*) adapter=opencode-plugin ;; gemini*) adapter=gemini-hook ;; pi|pi-signed) adapter=pi-ext ;; + omp) adapter=omp-ext ;; kimi*) fm_busy_kimi_verified || { printf ''; return 0; } adapter='kimi-wire kimi-hook' @@ -836,6 +840,19 @@ fm_busy_grok_tail_busy() { | grep -qiE "${FM_BUSY_REGEX:-${FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT:-Ctrl\\+c:cancel}}" } +# fm_busy_rovo_tail_busy: the Rovo-only temporary rendered-tail fallback. +# Consumes the tail on stdin; 0 when Rovo's verified animated busy line +# matches (the "Rovo is thinking..." text rendered while a turn is running, +# verified live on rovo 202609.1.2; both observed glyph variants share this +# literal text). rovo has no turn-end hook - its eventHooks fire at tool +# granularity only - so this fallback, like Grok's, is the only source; it is +# never armed as a semantic writer (fm_busy_sources_for_harness trusts +# nothing for rovo). FM_BUSY_ROVO_REGEX overrides the signature. +fm_busy_rovo_tail_busy() { + grep -v '^[[:space:]]*$' | tail -12 \ + | grep -qiE "${FM_BUSY_ROVO_REGEX:-Rovo is thinking}" +} + # fm_busy_classify: semantic classification for a task whose endpoint the # caller has already established as present. Prints " ": # busy|idle|unknown plus the producing source (see header). Never probes @@ -942,6 +959,28 @@ fm_busy_classify() { # [tail40] fi return 0 ;; + rovo*) + if [ -z "$tail40" ]; then + if command -v fm_backend_capture >/dev/null 2>&1; then + tail40=$(fm_backend_capture "$backend" "$target" 40 2>/dev/null) || { + printf 'unknown capture-failed' + return 0 + } + else + printf 'unknown capture-failed' + return 0 + fi + fi + # This fallback is best-effort: a long turn can scroll the busy marker + # out of the captured tail, so its absence means "can't tell," never + # definitive idle - matching the muse and cursor arms above. + if printf '%s' "$tail40" | fm_busy_rovo_tail_busy; then + printf 'busy rovo-regex' + else + printf 'unknown rovo-regex' + fi + return 0 + ;; esac printf 'unknown missing' } diff --git a/bin/fm-captain-hold.sh b/bin/fm-captain-hold.sh index e6083668d79..3ebc13cb923 100755 --- a/bin/fm-captain-hold.sh +++ b/bin/fm-captain-hold.sh @@ -24,13 +24,17 @@ # [--title ] [--repo <repo>] [--origin <origin-id>] [--until YYYY-MM-DD] # fm-captain-hold.sh answer <task-id> --decision-file <path> [--release] # fm-captain-hold.sh answers [<legacy-origin> | --any-origin] --source <provenance> (keyed answers on stdin) +# fm-captain-hold.sh reconcile-requests --source-id <source-id> --source <provenance> (task ids on stdin) # fm-captain-hold.sh bind <source-id> [<legacy-origin> | --any-origin] # fm-captain-hold.sh unbind <source-id> # fm-captain-hold.sh binding <source-id> # fm-captain-hold.sh complete <origin-id> (--none | <task-id>...) # fm-captain-hold.sh verify <origin-id> -# fm-captain-hold.sh open <task-id> +# fm-captain-hold.sh open <task-id> [--identity] [--distinguish-absent] # fm-captain-hold.sh diverged +# fm-captain-hold.sh reconcile list +# fm-captain-hold.sh reconcile close <task-id> --evidence-file <path> +# fm-captain-hold.sh reconcile note <task-id> --note-file <path> # # `hold` places an existing task under an active captain hold, or creates the # task first when no work item exists to hold (--title required to create; the @@ -83,6 +87,24 @@ # keeps closing its rows; `--any-origin` and the stored `(any)` marker mean # what an absent origin means and are accepted for the same reason. # +# RECONCILE IS RESERVED AT THIS INTAKE, NOT FILTERED IN A CHANNEL. +# The exact answer value `reconcile` means "go re-check reality", never "the +# captain answered". `answers` matches it before it reads the close mode, +# visibly refuses it, and never passes it to `answer`, so no channel and no +# card-declared mode can turn it into a close, release, or request. A separate +# `reconcile-requests` intake verifies a captured source's binding before it +# records a durable request under `state/reconcile-requests/`. +# +# `reconcile` is the verify-then-decide half. Both outcomes require the pending +# request created by the captain's board selection. `close` is the moot outcome: +# it requires the evidence that made the call moot, writes a `reconciled` resolution record +# under a `Reconciliation evidence:` label so it can never read as the +# captain's words, and closes the task. `note` is the still-active outcome: it +# appends one dated `Captain hold reconciled:` note and leaves the hold in +# place. A normal answer also retires the request because the call is settled. +# `list` is the read-only enumeration. +# docs/captain-hold-lifecycle.md owns the semantics. +# # A channel's ONLY job is to turn whatever it received into those keyed lines # and pipe them here. It must never map keys to tasks, build decision records, # choose a close mode beyond what its card declared, or close anything itself. @@ -119,27 +141,48 @@ # `decisions_reviewed=1` and `decision_keys=` keys, and an inventory entry that # names no existing task resolves through the legacy `<origin>-decision-<entry>` # identity, so pre-collapse metadata written by fm-decision-hold.sh verifies -# unchanged. An entry that exists as a task id is always that task. +# unchanged. An entry that exists as a task id is always that task. On the +# Beads backend an attested legacy markdown id that resolves to no task is +# accepted through the migrated row fm-hold-migration produced, found by the +# authoritative evidence first: a row whose notes carry the marker line +# "migrated from data/backlog.md id <legacy id>", alone or followed by +# " on <date>". Only when no row carries that line is the legacy id tried under +# the configured beads prefix, and that name-only guess is accepted solely for +# a single row still held for the captain; two such rows refuse rather than +# attest, and `complete` names each prefix-resolved row beside its attested +# legacy id so the guess stays auditable. # # `open` is the read-only predicate a mechanical closer asks before it may # retire a task's row: is this task still an open captain call? Exit 0 means it # is (not Done, hold kind captain), 1 means it is not, and 2 means the answer # could not be established, so a caller that must never close a live call can -# treat "cannot tell" as its own case instead of as a no. It prints nothing on -# 0 or 1 and mutates nothing. bin/fm-teardown.sh asks it before its automatic +# treat "cannot tell" as its own case instead of as a no. With +# `--distinguish-absent`, an absent local task returns 3 instead of 1. +# It prints nothing on these predicate results and mutates nothing, unless +# `--identity` asks it to print this call's +# LIFECYCLE identity, which it does on an exit 0 only. That identity - the +# hold-set stamp and the count of recorded answers - is what distinguishes two +# successive calls on one task id: re-holding released work starts a new +# lifecycle without necessarily touching the task's status log, so a consumer +# that bounds repeated work per call cannot use the task id alone. +# bin/fm-teardown.sh asks it before its automatic # backlog close and, on 0, returns the row to Queued with its deliverable # recorded instead (bin/fm-backlog-transition-lib.sh owns that transition), so -# holding the very work item a question gates is safe; `answer` remains the -# only act that closes a captain call. +# holding the very work item a question gates is safe; only `answer` with the +# captain's words or evidence-backed `reconcile close` closes the call. +# bin/fm-watch.sh asks it when an ordinary +# crew task reaches a due stale alarm - its open backlog hold need not appear in +# the task's last status line - and on a 0 bounds repeated alarms from new pane +# hashes for the decision. # # `diverged` is the read-only guard over the seam between the two records of # one captain call. See "record divergence" beside command_diverged below. # # Resolution records: the block written into the body names this script, the -# decision digest, and a `Resolution mode:` of answered, released, or repaired. -# Records written by the retired fm-decision-hold.sh (routed, declined, -# answered, repaired) are recognized everywhere a record is read, so nothing -# already closed needs rewriting. +# decision digest, and a `Resolution mode:` of answered, released, repaired, or +# reconciled. Records written by the retired fm-decision-hold.sh (routed, +# declined, answered, repaired) are recognized everywhere a record is read, so +# nothing already closed needs rewriting. # # Parent channel: inside a secondmate home a task held for the captain, and its # answer, are captain-facing facts the moment they are recorded, so `hold` @@ -170,10 +213,6 @@ DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" # shellcheck source=bin/fm-backlog-transition-lib.sh # shellcheck disable=SC1091 . "$SCRIPT_DIR/fm-backlog-transition-lib.sh" -# Resolve the configured backlog once for diagnostics; keep startup non-fatal so -# commands retain their existing read-error handling. -CAPTAIN_BACKLOG_FILE=$(fm_backlog_file "$DATA" 2>/dev/null) \ - || CAPTAIN_BACKLOG_FILE="${DATA%/}/backlog.md" # shellcheck source=bin/fm-wake-lib.sh # shellcheck disable=SC1091 . "$SCRIPT_DIR/fm-wake-lib.sh" @@ -181,12 +220,14 @@ CAPTAIN_BACKLOG_FILE=$(fm_backlog_file "$DATA" 2>/dev/null) \ # shellcheck disable=SC1091 . "$SCRIPT_DIR/fm-parent-channel-lib.sh" +PARENT_HOLD_PUBLISHED=0 publish_parent_hold() { # <task-id> <occurrence> <verb> <note> local id=$1 occurrence=$2 verb=$3 note=$4 rc=0 + PARENT_HOLD_PUBLISHED=0 fm_parent_channel_report "$FM_HOME" "$STATE" \ "$verb [key=captain-hold-$id-$occurrence]: captain hold $id: $(fm_parent_channel_clean_note "$note")" || rc=$? case "$rc" in - 0|1) ;; + 0|1) PARENT_HOLD_PUBLISHED=1 ;; *) printf 'actionable: task %s is held for the captain in this home but that did not reach the parent channel (rc=%s)\n' "$id" "$rc" >&2 ;; esac } @@ -241,6 +282,13 @@ acquire_task_control_lock() { # <task-id> CAPTAIN_CONTROL_LOCK_HELD=1 } +release_task_control_lock() { + [ "$CAPTAIN_CONTROL_LOCK_HELD" = 1 ] || return 0 + fm_lock_release "$CAPTAIN_CONTROL_LOCK" + CAPTAIN_CONTROL_LOCK_HELD=0 + CAPTAIN_CONTROL_LOCK= +} + sha256_text() { # <text> if command -v shasum >/dev/null 2>&1; then printf '%s' "$1" | shasum -a 256 | awk '{print $1}' @@ -278,13 +326,20 @@ load_decision() { # <path>; sets DECISION_TEXT and DECISION_DIGEST # Mutations address the configured data directory's backlog from its root, the # way bin/fm-backlog-transition-lib.sh addresses every transition, so a home -# with a relocated data directory keeps one backlog. +# with a relocated data directory keeps one backlog. The explicit --file file +# belongs to the markdown backend only; a non-markdown backend is addressed by +# the root's own tasks-axi configuration, exactly like the transition library's +# mutate path. tasks_axi() { local data file root data=$(fm_backlog_data_absolute "$DATA") || fail "data directory cannot be resolved: $DATA" - file=$(fm_backlog_file "$data") || fail "$FM_BACKLOG_TRANSITION_ERROR" root=$(fm_backlog_root "$data") || fail "$FM_BACKLOG_TRANSITION_ERROR" - (cd "$root" && tasks-axi "$@" --file "$file") + if [ "$(fm_tasks_axi_backend "$root")" = markdown ]; then + file=$(fm_backlog_file "$data") || fail "$FM_BACKLOG_TRANSITION_ERROR" + (cd "$root" && tasks-axi "$@" --file "$file") + else + (cd "$root" && tasks-axi "$@") + fi } require_tasks_axi() { @@ -376,6 +431,7 @@ body_has_resolution_record() { # <task-body> case "$1" in *"Resolution recorded by fm-captain-hold."*"Captain decision:"*) return 0 ;; *"Resolution recorded by fm-decision-hold."*"Captain decision:"*) return 0 ;; + *"Resolution recorded by fm-captain-hold."*"Reconciliation evidence:"*) return 0 ;; esac return 1 } @@ -414,16 +470,28 @@ recorded_resolution_mode() { # <task-body> printf '%s' "$rest" } +closed_answer_replay_mode_compatible() { # <mode> <task-body> + case "$1" in + answered|repaired|routed) return 0 ;; + esac + return 1 +} + +# The record's label is what keeps an evidence-backed reconciliation from +# reading as the captain's own words. `reconciled` closes a call that went moot +# and carries verified evidence; every other mode carries what the captain said. resolution_block() { # <mode> - printf 'Resolution recorded by fm-captain-hold.\nDecision digest: %s\nResolution mode: %s\n\nCaptain decision:\n%s\n' \ - "$DECISION_DIGEST" "$1" "$DECISION_TEXT" + local label='Captain decision:' + [ "$1" != reconciled ] || label='Reconciliation evidence:' + printf 'Resolution recorded by fm-captain-hold.\nDecision digest: %s\nResolution mode: %s\n\n%s\n%s\n' \ + "$DECISION_DIGEST" "$1" "$label" "$DECISION_TEXT" } # Durable state of one captain call: an active captain hold (annotations # surviving even when a date gate has expired) or a recorded captain answer. verify_hold_durable() { # <task-id> local id=$1 show state hold_kind body - show=$(task_show "$id") || fail "captain-held task $id is absent from $CAPTAIN_BACKLOG_FILE" + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -436,23 +504,219 @@ verify_hold_durable() { # <task-id> fail "captain-held task $id is neither held for the captain nor closed with a recorded captain answer" } +# --- migrated legacy-id resolution on the Beads backend --------------------- +# +# A home that moved its backlog from markdown to Beads no longer carries the +# legacy hold ids a scout report attested: the migration rehomed every held +# row under a prefixed fm- id and recorded its markdown identity in the row's +# notes as "migrated from data/backlog.md id <legacy id>", alone or followed by +# " on <date>" (fm-hold-migration wrote the dated form on 2026-09-04). When an +# attested legacy id resolves to no task, the beads backend accepts the row the +# migration produced, found by scanning the configured graph's notes for either +# form of that marker line, and only when no row carries the marker by +# prepending the configured prefix to the legacy id - a name-only guess, so it +# is accepted solely for a row still held for the captain and only when it is +# the single such row. A markdown home keeps its legacy rows verbatim, so its +# exact-id resolution is unchanged. + +CAPTAIN_MIGRATION_SCAN_LOADED=0 +CAPTAIN_MIGRATION_SCAN_JSON= +NL_SEP=$'\n' + +# Section-aware [beads] extraction from a .tasks.toml: only keys inside the +# [beads] section, comments stripped. Prints "<key> <value>" lines. +captain_beads_toml_entries() { # <toml-file> + [ -f "$1" ] || return 0 + LC_ALL=C awk ' + function trim(v) { sub(/^[[:space:]]+/, "", v); sub(/[[:space:]]+$/, "", v); return v } + BEGIN { inbeads = 0 } + { + line = $0 + sub(/[[:space:]]*#.*/, "", line) + line = trim(line) + if (line ~ /^\[[^]]+\]$/) { inbeads = (line == "[beads]"); next } + if (!inbeads) next + if (line ~ /^(prefix|path|binary)[[:space:]]*=/) { + key = line + sub(/[[:space:]]*=.*/, "", key) + sub(/^[^=]*=[[:space:]]*/, "", line) + gsub(/^"|"$/, "", line); gsub(/^'\''|'\''$/, "", line) + printf "%s %s\n", key, line + } + } + ' "$1" +} + +captain_beads_setting() { # <entries-output> <setting> + printf '%s\n' "$1" | sed -n "s/^$2 //p" | head -1 +} + +# Read the configured beads graph's row listing for a migration-note scan. +# The listing is deliberately re-read per unresolvable key: the cache below +# lives and dies with the command-substitution subshell every resolve_entry +# call site runs in, so it cannot persist across keys - bounded by a scout +# report's handful of attested ids. Returns 0 when the listing loads, and 2 +# with the reason on stderr when the graph cannot be read. +captain_migration_scan_load() { # <resolved-data-dir> + local data=$1 root entries bd_bin bd_path + [ "$CAPTAIN_MIGRATION_SCAN_LOADED" = 1 ] && return 0 + root=$(fm_backlog_root "$data") || { + printf 'fm-captain-hold: the configured data directory cannot be resolved for a migration scan: %s\n' "$FM_BACKLOG_TRANSITION_ERROR" >&2 + return 2 + } + if [ "$(fm_tasks_axi_backend "$root")" != beads ]; then + CAPTAIN_MIGRATION_SCAN_LOADED=1 + return 0 + fi + entries=$(captain_beads_toml_entries "$root/.tasks.toml") + bd_bin=$(captain_beads_setting "$entries" binary) + bd_path=$(captain_beads_setting "$entries" path) + bd_bin=${bd_bin:-bd} + if [ -z "$bd_path" ]; then + printf 'fm-captain-hold: the beads backend carries no graph path in %s, so a migrated hold cannot be found\n' "$root/.tasks.toml" >&2 + return 2 + fi + # A relative [beads] path resolves against the backlog root, the same rule + # every other .tasks.toml path consumer uses, never against the process CWD. + case "$bd_path" in + /*) ;; + *) bd_path="$root/$bd_path" ;; + esac + command -v "$bd_bin" >/dev/null 2>&1 || { + printf 'fm-captain-hold: the beads binary %s is not on PATH, so a migrated hold cannot be found\n' "$bd_bin" >&2 + return 2 + } + command -v jq >/dev/null 2>&1 || { + printf 'fm-captain-hold: jq is required to scan the beads graph for a migrated hold\n' >&2 + return 2 + } + local bd_err + bd_err=$(mktemp "${TMPDIR:-/tmp}/fm-captain-hold-bd.XXXXXX") || { + printf 'fm-captain-hold: cannot stage the beads graph read diagnostics\n' >&2 + return 2 + } + if ! CAPTAIN_MIGRATION_SCAN_JSON=$(BEADS_DIR="$bd_path" "$bd_bin" list --all --json 2>"$bd_err"); then + printf 'fm-captain-hold: reading the beads graph at %s failed (%s), so a migrated hold cannot be found\n' \ + "$bd_path" "$(sanitize_field "$(head -c 200 "$bd_err" | tr '\n' ' ')")" >&2 + rm -f "$bd_err" + return 2 + fi + rm -f "$bd_err" + CAPTAIN_MIGRATION_SCAN_LOADED=1 + return 0 +} + +# Resolve one attested legacy id to the migrated row that carries it on the +# beads backend. Prints "<row id> <how>" and returns 0 when exactly one +# migration matches, returns 1 when none does, and returns 2 with the reason on +# stderr when the scan itself cannot run or is ambiguous. The marker note is the +# authoritative evidence and is scanned first; the bare configured prefix is a +# guess, so it only runs when no marker line matches any identity and it accepts +# a row solely when that row is itself still held for the captain. +resolve_migrated_entry() { # <origin-or-empty> <entry> + local origin=$1 entry=$2 data root entries prefix derived show + local candidate candidate_matches prefixed matches count prefixed_matches prefixed_count + data=$(fm_backlog_data_absolute "$DATA") || { + printf 'fm-captain-hold: the migrated hold of %s cannot be resolved: %s\n' \ + "$entry" "${FM_BACKLOG_TRANSITION_ERROR:-the configured data directory $DATA cannot be resolved}" >&2 + return 2 + } + root=$(fm_backlog_root "$data") || { + printf 'fm-captain-hold: the migrated hold of %s cannot be resolved: %s\n' \ + "$entry" "${FM_BACKLOG_TRANSITION_ERROR:-the configured data directory $DATA cannot be resolved}" >&2 + return 2 + } + [ "$(fm_tasks_axi_backend "$root")" = beads ] || return 1 + # Every identity this entry could have been migrated under: the raw entry, + # and - for a pre-collapse channel key - the derived legacy identity its + # origin would have minted, because fm-hold-migration recorded the DERIVED + # id in each migrated row's marker note. + CAPTAIN_MIGRATION_IDENTITIES=$entry + if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then + derived=$(legacy_hold_id "$origin" "$entry") + if [ "$derived" != "$entry" ]; then + CAPTAIN_MIGRATION_IDENTITIES="$CAPTAIN_MIGRATION_IDENTITIES $derived" + fi + fi + captain_migration_scan_load "$data" || return 2 + matches= + if [ -n "$CAPTAIN_MIGRATION_SCAN_JSON" ]; then + for candidate in $CAPTAIN_MIGRATION_IDENTITIES; do + candidate_matches=$(printf '%s\n' "$CAPTAIN_MIGRATION_SCAN_JSON" | jq -r \ + --arg exact "migrated from data/backlog.md id $candidate" \ + --arg dated "migrated from data/backlog.md id $candidate on " \ + '.[] | select(((.notes // "") | split("\n")) | any(. == $exact or startswith($dated))) | .id' 2>/dev/null) || { + printf 'fm-captain-hold: the beads graph scan for the migrated hold of %s could not be parsed\n' "$candidate" >&2 + return 2 + } + matches="${matches}${matches:+$NL_SEP}${candidate_matches}" + done + count=$(printf '%s\n' "$matches" | sed '/^$/d' | wc -l | tr -d ' ') + case "$count" in + 0) : ;; + 1) printf '%s migrated-note' "$(printf '%s\n' "$matches" | sed '/^$/d' | sed -n 1p)"; return 0 ;; + *) + printf 'fm-captain-hold: the migrated hold of %s is ambiguous: %s rows carry its marker line (identities tried: %s)\n' \ + "$entry" "$count" "$(printf '%s' "$CAPTAIN_MIGRATION_IDENTITIES" | tr ' ' ',')" >&2 + return 2 + ;; + esac + fi + # No marker line anywhere: a mechanical migration keeps the legacy id under + # the configured prefix, but that name alone is evidence of nothing, so only + # a row still held for the captain - and only one of them - is accepted. + entries=$(captain_beads_toml_entries "$root/.tasks.toml") + prefix=$(captain_beads_setting "$entries" prefix) + [ -n "$prefix" ] || return 1 + prefixed_matches= + for candidate in $CAPTAIN_MIGRATION_IDENTITIES; do + case "$prefix" in + *-) prefixed="$prefix$candidate" ;; + *) prefixed="$prefix-$candidate" ;; + esac + show=$(task_show "$prefixed" 2>/dev/null) || continue + [ "$(show_field_value "$show" hold_kind)" = captain ] || continue + prefixed_matches="${prefixed_matches}${prefixed_matches:+$NL_SEP}$prefixed" + done + prefixed_count=$(printf '%s\n' "$prefixed_matches" | sed '/^$/d' | wc -l | tr -d ' ') + case "$prefixed_count" in + 0) return 1 ;; + 1) printf '%s migrated-prefix' "$prefixed_matches"; return 0 ;; + esac + printf 'fm-captain-hold: the migrated hold of %s is ambiguous: %s captain-held rows carry the configured prefix (identities tried: %s)\n' \ + "$entry" "$prefixed_count" "$(printf '%s' "$CAPTAIN_MIGRATION_IDENTITIES" | tr ' ' ',')" >&2 + return 2 +} + # Resolve one inventory entry or channel key to the task that carries it: the -# exact task id when it exists, else the legacy derived identity. -resolve_entry() { # <origin-or-empty> <entry>; prints the resolved id or fails - local origin=$1 entry=$2 legacy +# exact task id when it exists, else the legacy derived identity, else - on the +# beads backend - the migrated row the markdown-to-beads hold migration wrote. +# Prints "<resolved id> <how>", where <how> is exact, legacy, migrated-note or +# migrated-prefix, so a caller can record which evidence carried the attestation. +resolve_entry() { # <origin-or-empty> <entry>; prints "<id> <how>" or fails + local origin=$1 entry=$2 legacy migrated rc if task_show "$entry" >/dev/null 2>&1; then - printf '%s' "$entry" + printf '%s exact' "$entry" return 0 fi if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then legacy=$(legacy_hold_id "$origin" "$entry") if task_show "$legacy" >/dev/null 2>&1; then - printf '%s' "$legacy" + printf '%s legacy' "$legacy" return 0 fi - fail "no captain-held task $entry and no legacy identity $legacy in $CAPTAIN_BACKLOG_FILE" fi - fail "no captain-held task $entry in $CAPTAIN_BACKLOG_FILE" + rc=0 + migrated=$(resolve_migrated_entry "$origin" "$entry") || rc=$? + case "$rc" in + 0) printf '%s' "$migrated"; return 0 ;; + 2) return 2 ;; + esac + if [ -n "$origin" ] && [ "$origin" != "$BINDING_ANY" ]; then + legacy=$(legacy_hold_id "$origin" "$entry") + fail "no captain-held task $entry and no migrated hold for it in this home's configured backlog (data directory $DATA); the nearest legacy identity $legacy also resolves to nothing" + fi + fail "no captain-held task $entry and no migrated hold for it in this home's configured backlog (data directory $DATA)" } body_hold_set_timestamp() { # <decoded-task-body> @@ -666,7 +930,7 @@ command_answer() { load_decision "$decision_file" acquire_task_control_lock "$id" require_tasks_axi - show=$(task_show "$id") || fail "captain-held task $id is absent from $CAPTAIN_BACKLOG_FILE" + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" state=$(show_field "$show" state) hold_kind=$(show_field_value "$show" hold_kind) body=$(show_field "$show" body) @@ -681,15 +945,15 @@ command_answer() { [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ] \ || fail "captain-held task $id records a different captain decision" recorded_mode=$(recorded_resolution_mode "$body" || true) - [ "$recorded_mode" != released ] \ - || fail "task $id records this answer with mode released; a closed task cannot replay that release" + closed_answer_replay_mode_compatible "$recorded_mode" "$body" \ + || fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a captain-answer replay" [ "$release" = 0 ] \ || fail "task $id records this answer with mode ${recorded_mode:-unknown}; --release cannot reopen a closed task" remove_interrupted_answer_stamp "$id" if [ "$recorded_mode" = repaired ]; then - publish_parent_hold "$id" $((occurrence - 1)) resolved "answered (repaired)" + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) "answered (repaired)" else - publish_parent_hold "$id" $((occurrence - 1)) resolved answered + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) answered fi printf 'answered: %s\n' "$id" return 0 @@ -706,7 +970,7 @@ command_answer() { [ "$(show_field "$show" state)" = "done" ] || fail "recording the answer reopened closed task $id" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" - publish_parent_hold "$id" "$occurrence" resolved "answered (repaired)" + publish_parent_resolution_then_retire "$id" "$occurrence" "answered (repaired)" printf 'repaired: %s\n' "$id" return 0 fi @@ -723,13 +987,14 @@ command_answer() { recorded_mode=$(recorded_resolution_mode "$body" || true) case "$recorded_mode" in released) [ "$release" = 1 ] || fail "task $id records this answer as a release; retry with --release" ;; - answered) [ "$release" = 0 ] || fail "task $id records this answer as a close; retry without --release" ;; + answered|routed) [ "$release" = 0 ] || fail "task $id records this answer as a close; retry without --release" ;; + *) fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a captain-answer replay" ;; esac if ! close_answered "$id" "$release"; then fail "could not close answered captain-held task $id" fi remove_interrupted_answer_stamp "$id" - publish_parent_hold "$id" $((occurrence - 1)) resolved "$outcome" + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) "$outcome" printf '%s: %s\n' "$outcome" "$id" return 0 fi @@ -741,7 +1006,7 @@ command_answer() { show=$(task_show "$id") || fail "task $id disappeared after closing" body_has_resolution_record "$(show_field "$show" body)" \ || fail "captain-held task $id did not retain its durable resolution record" - publish_parent_hold "$id" "$occurrence" resolved "$outcome" + publish_parent_resolution_then_retire "$id" "$occurrence" "$outcome" printf '%s: %s\n' "$outcome" "$id" return 0 fi @@ -754,7 +1019,7 @@ command_answer() { [ "$recorded_mode" = released ] && [ "$release" = 1 ] \ || fail "task $id records this answer with mode ${recorded_mode:-unknown}; replay requires matching --release" remove_interrupted_answer_stamp "$id" - publish_parent_hold "$id" $((occurrence - 1)) resolved released + publish_parent_resolution_then_retire "$id" $((occurrence - 1)) released printf 'released: %s\n' "$id" return 0 fi @@ -852,6 +1117,10 @@ sanitize_field() { # <text> printf '%s' "$1" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177' | cut -c1-512 } +sanitize_reconcile_provenance() { + printf '%s' "$1" | tr '\n\r\t' ' ' | LC_ALL=C tr -d '\000-\037\177' | cut -c1-1024 +} + command_answers() { local origin='' source='' row rest key answer label mode id show state hold_kind body digest legacy_digest legacy_key local recorded_digest recorded_mode occurrence tmp err closed=0 skipped=0 reason release_flag tab=$'\t' @@ -890,6 +1159,11 @@ command_answers() { answer=$(sanitize_field "${answer:-}") [ -n "$answer" ] || continue label=$(sanitize_field "${label:-}") + if [ "$answer" = "$RECONCILE_VALUE" ]; then + printf 'refused: %s (reconcile requests require a bound captured source)\n' "$key" + skipped=$((skipped + 1)) + continue + fi release_flag='' case "${mode:-}" in ''|done) : ;; @@ -900,7 +1174,16 @@ command_answers() { continue ;; esac - if ! id=$(resolve_entry "$origin" "$key" 2>/dev/null); then + resolve_rc=0 + id=$(resolve_entry "$origin" "$key" 2>"$err") || resolve_rc=$? + id=${id%% *} + if [ "$resolve_rc" = 2 ]; then + reason=$(tr -d '\n' < "$err") + printf 'skipped: %s (migrated-hold scan refused%s)\n' "$key" "${reason:+: $reason}" + skipped=$((skipped + 1)) + continue + fi + if [ "$resolve_rc" -ne 0 ]; then printf 'skipped: %s (no captain-held task with that id)\n' "$key" skipped=$((skipped + 1)) continue @@ -930,14 +1213,15 @@ command_answers() { && { [ "$recorded_digest" = "$digest" ] \ || { case "$body" in *"Resolution recorded by fm-decision-hold."*) true ;; *) false ;; esac \ && [ -n "$legacy_digest" ] && [ "$recorded_digest" = "$legacy_digest" ]; }; }; then - if { [ -z "$release_flag" ] && [ "$state" = "done" ] && [ "$recorded_mode" != released ]; } \ + if { [ -z "$release_flag" ] && [ "$state" = "done" ] \ + && closed_answer_replay_mode_compatible "$recorded_mode" "$body"; } \ || { [ "$release_flag" = --release ] && [ "$state" != "done" ] \ && [ "$hold_kind" != captain ] && [ "$recorded_mode" = released ]; }; then occurrence=$(resolution_record_count "$body") case "$recorded_mode" in - repaired) publish_parent_hold "$id" "$occurrence" resolved "answered (repaired)" ;; - released) publish_parent_hold "$id" "$occurrence" resolved released ;; - *) publish_parent_hold "$id" "$occurrence" resolved answered ;; + repaired) publish_parent_resolution_then_retire "$id" "$occurrence" "answered (repaired)" ;; + released) publish_parent_resolution_then_retire "$id" "$occurrence" released ;; + *) publish_parent_resolution_then_retire "$id" "$occurrence" answered ;; esac printf 'closed: %s\n' "$id" closed=$((closed + 1)) @@ -972,8 +1256,276 @@ command_answers() { [ "$skipped" -eq 0 ] } +# --- reconcile: verify latest state, then close with evidence or annotate ---- +# +# The semantics are owned by docs/captain-hold-lifecycle.md; this section owns +# the durable record and the two terminal operations that retire it. Nothing +# here closes a captain call on the strength of a reconcile alone: `close` +# demands the evidence that made the call moot, and `note` leaves it open. + +RECONCILE_DIR="$STATE/reconcile-requests" +RECONCILE_SCHEMA=fm-reconcile-request.v1 +RECONCILE_VALUE=reconcile + +reconcile_request_path() { printf '%s/%s.request\n' "$RECONCILE_DIR" "$1"; } + +# Idempotent per task: a repeated reconcile keeps the one request and its +# original timestamp, so a re-delivered board answer never resets the clock on +# an obligation that is already open. +reconcile_request_record() { # <task-id> <provenance> + local id=$1 source=$2 path tmp + path=$(reconcile_request_path "$id") + [ ! -e "$path" ] || return 0 + (umask 077; mkdir -p "$RECONCILE_DIR") || return 1 + [ -d "$RECONCILE_DIR" ] && [ ! -L "$RECONCILE_DIR" ] || return 1 + tmp=$(umask 077; mktemp "$RECONCILE_DIR/.request.XXXXXX") || return 1 + if { + printf 'schema=%s\n' "$RECONCILE_SCHEMA" + printf 'task=%s\n' "$id" + printf 'requested=%s\n' "${FM_CAPTAIN_HOLD_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)}" + printf 'source=%s\n' "$(sanitize_reconcile_provenance "$source")" + } > "$tmp" && chmod 0600 "$tmp" && mv -f -- "$tmp" "$path"; then + return 0 + fi + rm -f -- "$tmp" + return 1 +} + +reconcile_request_read() { # <task-id>; sets RECONCILE_REQUESTED/RECONCILE_SOURCE + local id=$1 path schema task + path=$(reconcile_request_path "$id") + [ -f "$path" ] && [ ! -L "$path" ] || return 1 + schema=$(sed -n 's/^schema=//p' "$path" | head -1) + [ "$schema" = "$RECONCILE_SCHEMA" ] || fail "reconcile request has an incompatible schema: $path" + task=$(sed -n 's/^task=//p' "$path" | head -1) + [ "$task" = "$id" ] || fail "reconcile request names a different task: $path" + RECONCILE_REQUESTED=$(sed -n 's/^requested=//p' "$path" | head -1) + RECONCILE_SOURCE=$(sed -n 's/^source=//p' "$path" | head -1) +} + +reconcile_request_retire() { # <task-id> + rm -f -- "$(reconcile_request_path "$1")" \ + || fail "could not retire the pending reconcile request for $1" +} + +publish_parent_resolution_then_retire() { # <task-id> <occurrence> <note> + local id=$1 occurrence=$2 note=$3 request + request=$(reconcile_request_path "$id") + publish_parent_hold "$id" "$occurrence" resolved "$note" + if [ -e "$request" ] && [ "$PARENT_HOLD_PUBLISHED" != 1 ]; then + fail "could not publish the answered captain-held task $id to its parent" + fi + reconcile_request_retire "$id" +} + +command_reconcile_requests() { + local source_id='' source='' origin row id note provenance show created=0 skipped=0 tab=$'\t' + while [ "$#" -gt 0 ]; do + case "$1" in + --source-id) shift; source_id=${1:-} ;; + --source) shift; source=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_source_id "$source_id" + [ -n "$source" ] || fail "--source provenance is required" + origin=$(read_binding "$source_id") || fail "cannot verify the binding for source $source_id" + [ -n "$origin" ] || fail "source $source_id is not bound; no reconcile requests were created" + require_tasks_axi + while IFS= read -r row; do + id=${row%%"$tab"*} + note='' + case "$row" in *"$tab"*) note=${row#*"$tab"} ;; esac + [ -n "$id" ] || continue + case "$id" in + *[!A-Za-z0-9._-]*) printf 'refused: %s (invalid task id)\n' "$id"; skipped=$((skipped + 1)); continue ;; + esac + [ "${#id}" -le 128 ] \ + || { printf 'refused: %s (task id is too long)\n' "$id"; skipped=$((skipped + 1)); continue; } + acquire_task_control_lock "$id" + show=$(task_show "$id") || true + if [ -z "$show" ]; then + printf 'refused: %s (absent)\n' "$id" + skipped=$((skipped + 1)) + elif [ "$(show_field "$show" state)" = "done" ]; then + printf 'refused: %s (already closed)\n' "$id" + skipped=$((skipped + 1)) + elif [ "$(show_field_value "$show" hold_kind)" != captain ]; then + printf 'refused: %s (not held for the captain)\n' "$id" + skipped=$((skipped + 1)) + else + provenance=$source + [ -z "$note" ] || provenance="$source; captain note: $(sanitize_field "$note")" + if reconcile_request_record "$id" "$provenance"; then + printf 'reconcile: %s\n' "$id" + created=$((created + 1)) + else + printf 'refused: %s (cannot record the reconcile request)\n' "$id" + skipped=$((skipped + 1)) + fi + fi + release_task_control_lock || fail "cannot release task control for $id" + done + printf 'reconcile-requests: created=%s skipped=%s\n' "$created" "$skipped" + [ "$skipped" -eq 0 ] +} + +command_reconcile() { + local action=${1:-} + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + case "$action" in + list) reconcile_list "$@" ;; + close) reconcile_close "$@" ;; + note) reconcile_note "$@" ;; + *) usage >&2; exit 2 ;; + esac +} + +reconcile_list() { + local path id count=0 + [ "$#" -eq 0 ] || { usage >&2; exit 2; } + [ -d "$RECONCILE_DIR" ] || { printf 'reconcile-requests: 0\n'; return 0; } + for path in "$RECONCILE_DIR"/*.request; do + [ -e "$path" ] || continue + id=${path##*/}; id=${id%.request} + RECONCILE_REQUESTED='' + RECONCILE_SOURCE='' + reconcile_request_read "$id" || continue + printf '%s\trequested=%s\tsource=%s\n' "$id" "$RECONCILE_REQUESTED" "$RECONCILE_SOURCE" + count=$((count + 1)) + done + printf 'reconcile-requests: %s\n' "$count" +} + +# The moot outcome. The evidence is what closes the call, and the `reconciled` +# resolution mode is what keeps the record from claiming the captain answered. +reconcile_close() { + local id=${1:-} evidence_file='' show state hold_kind body occurrence recorded_mode + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + while [ "$#" -gt 0 ]; do + case "$1" in + --evidence-file) shift; evidence_file=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_slug task-id "$id" + [ -n "$evidence_file" ] || fail "--evidence-file is required; a moot call closes on evidence, never on assertion" + load_decision "$evidence_file" + acquire_task_control_lock "$id" + reconcile_request_read "$id" \ + || fail "task $id has no pending board-created reconcile request" + require_tasks_axi + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + state=$(show_field "$show" state) + hold_kind=$(show_field_value "$show" hold_kind) + body=$(show_field "$show" body) + occurrence=$(( $(resolution_record_count "$body") + 1 )) + if [ "$state" = "done" ]; then + # An exact retry finishes an interrupted close and stays idempotent; a + # different evidence text on an already closed call is refused. + body_has_resolution_record "$body" \ + || fail "task $id is already closed with no resolution record; use answer to record what closed it" + [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ] \ + || fail "task $id records a different resolution; it cannot be reconciled again" + [ "$(recorded_resolution_mode "$body" || true)" = reconciled ] \ + || fail "task $id was not closed by reconciliation" + occurrence=$(resolution_record_count "$body") + remove_interrupted_answer_stamp "$id" + publish_parent_hold "$id" "$occurrence" resolved reconciled + [ "$PARENT_HOLD_PUBLISHED" = 1 ] \ + || fail "could not publish the reconciled captain-held task $id to its parent" + reconcile_request_retire "$id" + printf 'reconciled: %s\n' "$id" + return 0 + fi + [ "$hold_kind" = captain ] \ + || fail "task $id is not held for the captain; there is no captain call to reconcile" + if body_has_resolution_record "$body" \ + && [ "$(recorded_decision_digest "$body" || true)" = "$DECISION_DIGEST" ]; then + recorded_mode=$(recorded_resolution_mode "$body" || true) + [ "$recorded_mode" = reconciled ] \ + || fail "task $id records this resolution with mode ${recorded_mode:-unknown}; it is not a reconciliation retry" + occurrence=$(resolution_record_count "$body") + else + write_resolution_record "$id" reconciled "$body" + fi + close_answered "$id" 0 || fail "could not close reconciled captain-held task $id" + remove_interrupted_answer_stamp "$id" + show=$(task_show "$id") || fail "task $id disappeared after closing" + body_has_resolution_record "$(show_field "$show" body)" \ + || fail "captain-held task $id did not retain its durable resolution record" + publish_parent_hold "$id" "$occurrence" resolved reconciled + [ "$PARENT_HOLD_PUBLISHED" = 1 ] \ + || fail "could not publish the reconciled captain-held task $id to its parent" + reconcile_request_retire "$id" + printf 'reconciled: %s\n' "$id" +} + +# The still-active outcome. The hold survives, so the call stays the captain's +# and stays on Captain's Call, now carrying what the re-check found. +reconcile_note() { + local id=${1:-} note_file='' note show body stamp tmp note_digest marker + [ "$#" -ge 1 ] || { usage >&2; exit 2; } + shift + while [ "$#" -gt 0 ]; do + case "$1" in + --note-file) shift; note_file=${1:-} ;; + *) usage >&2; exit 2 ;; + esac + shift + done + validate_slug task-id "$id" + [ -n "$note_file" ] || fail "--note-file is required; leaving a call open records what the re-check found" + [ -f "$note_file" ] || fail "note file does not exist: $note_file" + note=$(cat "$note_file") + [ -n "$note" ] || fail "note file must not be empty" + [ "$(printf '%s' "$note" | LC_ALL=C wc -c | tr -d ' ')" -le 8192 ] \ + || fail "note file exceeds 8192 bytes" + acquire_task_control_lock "$id" + reconcile_request_read "$id" \ + || fail "task $id has no pending board-created reconcile request" + require_tasks_axi + command_open "$id" \ + || fail "task $id is not an open captain call; a note cannot keep a closed call open" + show=$(task_show "$id") || fail "captain-held task $id is absent from this home's configured backlog (data directory $DATA)" + body=$(decode_shown_value "$(show_field "$show" body)") \ + || fail "could not decode the existing body for $id" + note_digest=$(sha256_text "$note") + marker="Reconcile request: $RECONCILE_REQUESTED | $RECONCILE_SOURCE | note digest: $note_digest" + case "$body" in + *"$marker"*) + reconcile_request_retire "$id" \ + || fail "could not retire the applied reconcile request for $id" + command_open "$id" || fail "recording the reconcile note released captain-held task $id" + printf 'still-open: %s\n' "$id" + return 0 + ;; + esac + stamp=${FM_CAPTAIN_HOLD_NOW:-$(date -u +%Y-%m-%dT%H:%M:%SZ)} + tmp=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-captain-hold-note.XXXXXX") \ + || fail "cannot stage the reconcile note" + if ! printf '%s\n\nCaptain hold reconciled: %s\n%s\n%s\n' "$body" "$stamp" "$marker" "$note" > "$tmp"; then + rm -f -- "$tmp" + fail "cannot stage the reconcile note for $id" + fi + if ! tasks_axi update "$id" --body-file "$tmp" --archive-body >/dev/null; then + rm -f -- "$tmp" + fail "could not record the reconcile note on $id" + fi + rm -f -- "$tmp" + reconcile_request_retire "$id" \ + || fail "could not retire the applied reconcile request for $id" + command_open "$id" || fail "recording the reconcile note released captain-held task $id" + printf 'still-open: %s\n' "$id" +} + command_complete() { - local origin=${1:-} meta previous='' supplied='' keys='' entry key status_file open raw_open has_meta=0 transfer_rc + local origin=${1:-} meta previous='' supplied='' keys='' entry key status_file open raw_open has_meta=0 transfer_rc resolved + local resolved_how attested_by_prefix='' [ "$#" -ge 2 ] || { usage >&2; exit 2; } validate_slug origin-id "$origin" shift @@ -1004,7 +1556,16 @@ command_complete() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - verify_hold_durable "$(resolve_entry "$origin" "$entry")" + if ! resolved=$(resolve_entry "$origin" "$entry"); then + # resolve_entry has already refused on stderr naming the entry. + exit 1 + fi + resolved_how=${resolved##* } + resolved=${resolved%% *} + verify_hold_durable "$resolved" + if [ "$resolved_how" = migrated-prefix ]; then + attested_by_prefix="${attested_by_prefix}${attested_by_prefix:+ }$entry=$resolved" + fi done <<EOF $(printf '%s\n' "$keys" | tr ',' '\n') EOF @@ -1042,11 +1603,12 @@ $raw_open EOF fi fi - printf 'complete: %s captain-call inventory reviewed%s\n' "$origin" "${keys:+ ($keys)}" + printf 'complete: %s captain-call inventory reviewed%s%s\n' "$origin" "${keys:+ ($keys)}" \ + "${attested_by_prefix:+ [attested through the configured prefix: $attested_by_prefix]}" } command_verify() { - local origin=${1:-} meta reviewed keys entry key open + local origin=${1:-} meta reviewed keys entry key open resolved [ "$#" -eq 1 ] || { usage >&2; exit 2; } validate_slug origin-id "$origin" meta="$STATE/$origin.meta" @@ -1058,7 +1620,11 @@ command_verify() { if [ -n "$keys" ]; then while IFS= read -r entry; do [ -n "$entry" ] || continue - verify_hold_durable "$(resolve_entry "$origin" "$entry")" + if ! resolved=$(resolve_entry "$origin" "$entry"); then + # resolve_entry has already refused on stderr naming the entry. + exit 1 + fi + verify_hold_durable "${resolved%% *}" done <<EOF $(printf '%s\n' "$keys" | tr ',' '\n') EOF @@ -1195,12 +1761,23 @@ EOF } # Still an open captain call? Exit 0 yes, 1 no, 2 cannot tell (see the header). -# A row this home does not carry holds no captain call, so an absent task is a -# plain no; every other read failure is a 2, printed to stderr, because a -# mechanical closer must never read "cannot tell" as permission to close. -command_open() { # <task-id> - local id=${1:-} data state - [ "$#" -eq 1 ] || { usage >&2; exit 2; } +# A row this home does not carry is 3 when the caller requests the distinction; +# every other read failure is a 2, printed to stderr, because a mechanical +# closer must never read "cannot tell" as permission to close. +command_open() { # <task-id> [--identity] [--distinguish-absent] + local id='' identity=0 distinguish_absent=0 data state show shown_body + while [ "$#" -gt 0 ]; do + case "$1" in + --identity) identity=1 ;; + --distinguish-absent) distinguish_absent=1 ;; + -*) usage >&2; exit 2 ;; + *) + [ -z "$id" ] || { usage >&2; exit 2; } + id=$1 + ;; + esac + shift + done case "$id" in ''|*[!A-Za-z0-9._-]*) printf 'fm-captain-hold: task id must be a non-empty privacy-safe slug: %s\n' "$id" >&2 @@ -1213,11 +1790,24 @@ command_open() { # <task-id> if fm_backlog_row_probe "$data" "$id"; then state=${FM_BACKLOG_ROW_STATE%% *} if [ "$state" != "done" ] && [ "$FM_BACKLOG_ROW_HOLD_KIND" = captain ]; then + if [ "$identity" -eq 1 ]; then + show=$(task_show "$id") || { + printf 'fm-captain-hold: captain call %s is open but its record could not be read\n' "$id" >&2 + exit 2 + } + shown_body=$(show_field "$show" body) + printf '%s#%s\n' \ + "$(body_hold_set_timestamp "$(decode_shown_value "$shown_body")")" \ + "$(resolution_record_count "$shown_body")" + fi return 0 fi return 1 fi - [ "$FM_BACKLOG_ROW_RESULT" != not_found ] || return 1 + if [ "$FM_BACKLOG_ROW_RESULT" = not_found ]; then + [ "$distinguish_absent" = 0 ] || return 3 + return 1 + fi printf 'fm-captain-hold: %s\n' "$FM_BACKLOG_ROW_ERROR" >&2 exit 2 } @@ -1226,6 +1816,7 @@ case "${1:-}" in hold) shift; command_hold "$@" ;; answer) shift; command_answer "$@" ;; answers) shift; command_answers "$@" ;; + reconcile-requests) shift; command_reconcile_requests "$@" ;; bind) shift; command_bind "$@" ;; unbind) shift; command_unbind "$@" ;; binding) shift; command_binding "$@" ;; @@ -1233,6 +1824,7 @@ case "${1:-}" in verify) shift; command_verify "$@" ;; open) shift; command_open "$@" ;; diverged) shift; command_diverged "$@" ;; + reconcile) shift; command_reconcile "$@" ;; -h|--help) usage ;; *) usage >&2; exit 2 ;; esac diff --git a/bin/fm-classify-lib.sh b/bin/fm-classify-lib.sh index 506f398cce9..fdaf0b23474 100755 --- a/bin/fm-classify-lib.sh +++ b/bin/fm-classify-lib.sh @@ -418,6 +418,14 @@ _fm_decision_key_transition_allowed() { # <key> <note> return 0 } +_fm_is_pending_reply_escalation() { # <key> <note> + case "$1" in pending-reply-*) ;; *) return 1 ;; esac + case "$2" in + pending-reply-missed:*|pending-reply-delivery-unknown:*|pending-reply-recovery-delivery-failed:*|pending-reply-recovery-delivery-unknown:*) return 0 ;; + *) return 1 ;; + esac +} + _fm_decision_fold_line() { # <open-set> <status-line> <resolve-verb> <held-verb> local open=$1 line=$2 resolve=$3 held=$4 verb key note # Blank-line guard. A `case` glob answers "does this line hold any non-space @@ -644,9 +652,9 @@ _fm_open_decisions_file_ident() { # <file> -> strongest available identity return fi if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - ident=$(LC_ALL=C stat -f '%d:%i' "$f" 2>/dev/null) || return 1 - epoch=$(LC_ALL=C stat -f '%B' "$f" 2>/dev/null) || epoch=0 - if [ "$epoch" != 0 ]; then birth=$(LC_ALL=C stat -f '%FB' "$f" 2>/dev/null) || birth=''; else birth=''; fi + ident=$(LC_ALL=C /usr/bin/stat -f '%d:%i' "$f" 2>/dev/null) || return 1 + epoch=$(LC_ALL=C /usr/bin/stat -f '%B' "$f" 2>/dev/null) || epoch=0 + if [ "$epoch" != 0 ]; then birth=$(LC_ALL=C /usr/bin/stat -f '%FB' "$f" 2>/dev/null) || birth=''; else birth=''; fi else ident=$(LC_ALL=C stat -c '%d:%i' "$f" 2>/dev/null) || return 1 epoch=$(LC_ALL=C stat -c '%W' "$f" 2>/dev/null) || epoch=0 @@ -663,7 +671,7 @@ _fm_status_file_size() { # <status-file> return fi if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%z' "$f" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%z' "$f" 2>/dev/null else LC_ALL=C stat -c '%s' "$f" 2>/dev/null fi @@ -672,7 +680,7 @@ _fm_status_file_size() { # <status-file> _fm_status_file_mtime() { # <status-file> local f=$1 if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%m' "$f" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%m' "$f" 2>/dev/null else LC_ALL=C stat -c '%Y' "$f" 2>/dev/null fi @@ -1063,7 +1071,7 @@ status_presentation_marker_parse() { _status_observed_path_state() { if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%HT:%p' "$1" 2>/dev/null + LC_ALL=C /usr/bin/stat -f '%HT:%p' "$1" 2>/dev/null else LC_ALL=C stat -c '%F:%f' "$1" 2>/dev/null fi @@ -1555,12 +1563,16 @@ window_to_task() { # Capture the bytes of an append-only status log at or after <start-offset> under # one size-and-identity snapshot. -# The record form prints `<endpoint>\t<identity>\t<events>` and returns 0 when +# The record form produces `<endpoint>\t<identity>\t<events>` and returns 0 when # the span has actionable events, joining every such event in source order with # ` ; ` so callers report the complete captured span before committing it. +# With optional <record-var>, it assigns that record instead of printing it; with +# optional <needs-decision-var>, it also assigns 1 when the span newly surfaces a +# needs-decision, captain-held declaration, or pending-reply escalation, otherwise +# 0. This side-band classification never changes the event text. # It returns 1 after a successful classification with no actionable event; an -# existing log still prints its committable endpoint and identity, while an absent -# log is the ordinary empty case and prints no record. +# existing log still produces its committable endpoint and identity, while an absent +# log is the ordinary empty case and produces no record. # It returns 2 with no committable endpoint when an existing status object cannot # be classified. # The simpler wrapper prints only the event field, and the predicate discards the @@ -1615,9 +1627,9 @@ _fm_status_open_decision_origins() { # <status-file> printf '%s' "$origins" } -status_span_first_actionable_record() { # <status-file> <start-offset> - local f=$1 start=${2:-0} size ident cur_ident scratch chunk_file full_file prefix_file - local line verb key origins='' folded=0 rc=1 failed=0 prefix_lines=0 line_number=0 live_line='' events='' _line _key +status_span_first_actionable_record() { # <status-file> <start-offset> [record-var] [needs-decision-var] + local f=$1 start=${2:-0} output_var=${3-} needs_var=${4-} size ident cur_ident scratch chunk_file full_file prefix_file result + local line verb key origins='' folded=0 rc=1 failed=0 prefix_lines=0 line_number=0 live_line='' events='' _line _key _fm_span_needs_decision=0 [ -e "$f" ] || { [ -L "$f" ] && return 2; return 1; } [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 2 ident=$(_fm_open_decisions_file_ident "$f") || return 2 @@ -1626,7 +1638,16 @@ status_span_first_actionable_record() { # <status-file> <start-offset> case "$size" in ''|*[!0-9]*) return 2 ;; esac case "$start" in ''|*[!0-9]*) start=0 ;; esac [ "$start" -le "$size" ] || start=0 - [ "$start" -lt "$size" ] || { printf '%s\t%s' "$size" "$ident"; return 1; } + if [ "$start" -ge "$size" ]; then + result="${size}"$'\t'"${ident}" + if [ -n "$output_var" ]; then + printf -v "$output_var" '%s' "$result" + [ -z "$needs_var" ] || printf -v "$needs_var" '%s' 0 + else + printf '%s' "$result" + fi + return 1 + fi scratch=$(_fm_status_span_scratch "$f") || return 2 chunk_file="${scratch}.span"; full_file="${scratch}.full"; prefix_file="${scratch}.prefix" _fm_status_read_span "$f" "$start" "$((size - start))" > "$chunk_file" 2>/dev/null \ @@ -1638,6 +1659,13 @@ status_span_first_actionable_record() { # <status-file> <start-offset> while IFS= read -r line || [ -n "$line" ]; do line_number=$((line_number + 1)) case "$line" in *[![:space:]]*) ;; *) continue ;; esac + if status_is_captain_held "$line"; then + # A transfer closes the status-log decision and remains non-actionable to + # stale classification. The side-band marker lets signal routing surface + # the captain-owned hold without changing that established stale verdict. + _fm_span_needs_decision=1 + continue + fi status_is_captain_relevant "$line" || continue verb=$(status_line_verb "$line") case "$verb" in @@ -1645,12 +1673,14 @@ status_span_first_actionable_record() { # <status-file> <start-offset> key=$(_fm_decision_key "$line") || { [ -n "$events" ] && events="${events} ; " events="${events}${line}" + [ "$verb" = needs-decision ] && _fm_span_needs_decision=1 rc=0 continue } _fm_decision_key_transition_allowed "$key" "$(status_line_note "$line")" || { [ -n "$events" ] && events="${events} ; " events="${events}reconciliation-required: ${line}" + [ "$verb" = needs-decision ] && _fm_span_needs_decision=1 rc=0 continue } @@ -1674,6 +1704,10 @@ EOF [ -n "$live_line" ] && [ "$((prefix_lines + line_number))" -eq "$live_line" ] || continue [ -n "$events" ] && events="${events} ; " events="${events}${line}" + if [ "$verb" = needs-decision ] || { [ "$verb" = blocked ] && + _fm_is_pending_reply_escalation "$key" "$(status_line_note "$line")"; }; then + _fm_span_needs_decision=1 + fi rc=0 ;; *) @@ -1685,7 +1719,13 @@ EOF done < "$chunk_file" rm -f "$chunk_file" "$full_file" "$prefix_file" [ "$failed" -eq 0 ] || return 2 - if [ "$rc" -eq 0 ]; then printf '%s\t%s\t%s' "$size" "$ident" "$events"; else printf '%s\t%s' "$size" "$ident"; fi + if [ "$rc" -eq 0 ]; then result="${size}"$'\t'"${ident}"$'\t'"${events}"; else result="${size}"$'\t'"${ident}"; fi + if [ -n "$output_var" ]; then + printf -v "$output_var" '%s' "$result" + [ -z "$needs_var" ] || printf -v "$needs_var" '%s' "$_fm_span_needs_decision" + else + printf '%s' "$result" + fi return "$rc" } diff --git a/bin/fm-claude-stop-autoarm.sh b/bin/fm-claude-stop-autoarm.sh index 762b1a3dcf4..5f8b2d14a75 100755 --- a/bin/fm-claude-stop-autoarm.sh +++ b/bin/fm-claude-stop-autoarm.sh @@ -18,8 +18,8 @@ # - AFK: while state/.afk exists the away daemon owns the watcher and triage; # this hook exits 0 and NEVER rewakes the primary (checked again at # translation time so a mid-cycle AFK transition is honored). -# - Need: arms only while work is in flight (state/*.meta) or X mode has a -# relay poll to run (state/x-watch.check.sh); an idle home exits 0. +# - Need: arms only while the home needs supervision, as +# bin/fm-supervision-lib.sh defines it; an idle home exits 0. # - Single-flight: Claude does not dedupe async hooks, so exactly one # GENERATION owner arms per event epoch: the epoch ledger's monotonic # sequence is the claim generation, every firing defers (exit 0) to a live @@ -73,7 +73,6 @@ FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" -GRACE=${FM_GUARD_GRACE:-300} OWNER_LOCK="$STATE/.claude-autoarm.lock" FAILURE_NOTICE="$STATE/.claude-autoarm-failure-notified" FAILURE_ALARM="$STATE/.claude-autoarm-failure-alarmed" @@ -94,6 +93,13 @@ esac # shellcheck source=bin/fm-hook-host-lib.sh . "$SCRIPT_DIR/fm-hook-host-lib.sh" +# fm-watch.sh touches the liveness beacon once per cycle, immediately before +# its terminal wait, so a healthy watcher's beacon can legitimately age up to +# FM_POLL seconds between touches (docs/turnend-guard.md "Guard grace and the +# poll cadence"). fm_poll_derived_grace (bin/fm-wake-lib.sh) is the single +# owner of that max(300, poll+60) derivation. +GRACE=${FM_GUARD_GRACE:-$(fm_poll_derived_grace)} + # Consume the Stop payload once. The decisions below are state-based; the # payload is read so a slow writer can never wedge on a full pipe, and its host # is inspected before anything else runs. @@ -129,7 +135,7 @@ fi # --- AFK: the away daemon owns the watcher and triage; never rewake ---------- [ -e "$STATE/.afk" ] && exit 0 -# --- need: in-flight work or an X-mode relay poll ---------------------------- +# --- need: whatever bin/fm-supervision-lib.sh counts as supervision need ------ need_supervision() { fm_supervision_needed "$STATE" "$GRACE" } @@ -217,9 +223,9 @@ while [ "$attempt" -lt "$AUTOARM_ATTEMPTS" ]; do attempt=$((attempt + 1)) OUT=$(mktemp "$STATE/.claude-autoarm-output.XXXXXX") || OUT= if [ -n "$OUT" ]; then - "$SCRIPT_DIR/fm-watch-arm.sh" >"$OUT" 2>&1 || true + FM_GUARD_GRACE="$GRACE" "$SCRIPT_DIR/fm-watch-arm.sh" >"$OUT" 2>&1 || true else - "$SCRIPT_DIR/fm-watch-arm.sh" >/dev/null 2>&1 || true + FM_GUARD_GRACE="$GRACE" "$SCRIPT_DIR/fm-watch-arm.sh" >/dev/null 2>&1 || true fi # AFK may have appeared mid-cycle: the daemon owns triage now, so suppress diff --git a/bin/fm-composer-lib.sh b/bin/fm-composer-lib.sh index 07b3b02fffb..cdea8d98abd 100644 --- a/bin/fm-composer-lib.sh +++ b/bin/fm-composer-lib.sh @@ -290,7 +290,7 @@ fm_composer_strip_ghost() { # Matching a footer to confirm a keystroke landed is a different question from # asking what a worker is doing, and the two must not be conflated. # Delivery-only rendered busy footers per harness. claude/codex: "esc to -# interrupt"; opencode: "esc interrupt"; pi: "Working..."; grok: "Ctrl+c:cancel". +# interrupt"; opencode: "esc interrupt"; pi: "Working..."; omp: "Working…"; grok: "Ctrl+c:cancel". # Claude's current spinner has a rotating glyph and word, but every active-turn # line has an ellipsis followed by a parenthesized elapsed duration. Keep this # signature separate from the shared default because that shape is not generic @@ -311,11 +311,28 @@ fm_composer_strip_ghost() { # part of that union for the same reason the others are: without it a cursor # submit could never be acknowledged, because cursor parks its terminal cursor # outside its composer and the composer verdict is therefore always `unknown`. -FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working\.\.\.|Ctrl\+c:cancel|ctrl\+c to stop' +FM_DELIVERY_BUSY_REGEX_DEFAULT='esc (to )?interrupt|Working(\.\.\.|…)|Ctrl\+c:cancel|ctrl\+c to stop' FM_DELIVERY_CLAUDE_BUSY_REGEX_DEFAULT='esc to interrupt|…[[:space:]]+\([0-9]+[smh]' FM_DELIVERY_CODEX_BUSY_REGEX_DEFAULT='esc to interrupt' FM_DELIVERY_OPENCODE_BUSY_REGEX_DEFAULT='esc interrupt' FM_DELIVERY_PI_BUSY_REGEX_DEFAULT='Working\.\.\.' +# omp (Oh My Pi) renders its TUI busy line as `Working…` with U+2026 HORIZONTAL +# ELLIPSIS, not Pi's three ASCII dots (verified byte-level on omp 18.1.2, +# re-verified live on 18.1.11 through the Herdr backend). Only the TUI form is +# accepted: every supervised omp pane is the TUI, and the three-dot spelling its +# headless -p mode writes to stderr never reaches a pane. The status row's +# leading braille spinner plus elapsed cell (`⠧ 11s`) is the second, independent +# busy signal, so no single vendor string is load-bearing; its idle form is a +# static identity glyph with no elapsed time. +# The spinner is an alternation of omp 18.1.11's unicode-preset frames (its +# `status` set ⣾⣽⣻⢿⡿⣟⣯⣷ and `activity` set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the +# build that rendered the live `⠧`), declared once for the busy regex and the +# status-row furniture rule below. It is deliberately NOT a bracket range over +# the braille block: GNU grep rejects a range between multibyte endpoints +# ("Invalid collation character"), so `[⠁-⣿]` compiled on macOS and failed +# every omp busy and furniture read on Linux CI. +FM_OMP_SPINNER_FRAMES_RE='(⠋|⠙|⠹|⠸|⠼|⠴|⠦|⠧|⠇|⠏|⣾|⣽|⣻|⢿|⡿|⣟|⣯|⣷)' +FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT='Working…|^[[:space:]]*'"$FM_OMP_SPINNER_FRAMES_RE"'[[:space:]]+[0-9]+[smh]' FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT='Ctrl\+c:cancel' # cursor-agent's busy footer. The TOKEN is matched, not the spinner verb: the # same version rendered both `Working` and `Running` beside its braille spinner @@ -338,6 +355,7 @@ fm_busy_lines_match() { # [harness] codex) regex=$FM_DELIVERY_CODEX_BUSY_REGEX_DEFAULT ;; opencode) regex=$FM_DELIVERY_OPENCODE_BUSY_REGEX_DEFAULT ;; pi|pi-signed) regex=$FM_DELIVERY_PI_BUSY_REGEX_DEFAULT ;; + omp) regex=$FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT ;; grok) regex=$FM_DELIVERY_GROK_BUSY_REGEX_DEFAULT ;; kimi) regex=$FM_DELIVERY_KIMI_BUSY_REGEX_DEFAULT ;; cursor) regex=$FM_DELIVERY_CURSOR_BUSY_REGEX_DEFAULT ;; @@ -375,6 +393,25 @@ FM_COMPOSER_IDLE_RE_DEFAULT='^Type a message\.\.\.$|^Ask anything\.\.\.|^Plan, s # ("Build · GPT-5.5 Fast OpenAI · high"). It is composer furniture, not typed # text, and only the run's LAST row is ever matched against it. FM_COMPOSER_LEFTBAR_FOOTER_RE_DEFAULT='^(Build|Plan)[[:space:]]+·[[:space:]]+' +# omp (Oh My Pi) draws a one-row status line directly BELOW its borderless +# composer: an identity or spinner cell, then middle-dot separated model, path, +# git, and context cells. Verified live through Herdr on omp 18.1.11: +# ` π · ◔ GPT-6-Astra · 🌳 …-workspace · ⑂ detached · ◫ 15.4%/272K ⟲ · (sub)` +# idle under the unicode preset, ` 󰵗 · qwen3:8b · … · 36.7%/41K` under +# nerd, and ` ⠧ 11s · …` while busy. Without this rule the bare composer's +# wrap region walks straight into that row and an idle omp pane reads +# `pending`, the false verdict that skipped the doorbell on the first live omp +# worker. A row is omp status furniture when it opens with omp's identity cell +# then a middle dot (`π` under the unicode preset, `󰵗` under nerd: the +# `icon.omp` of those omp 18.1.11 presets, never an arbitrary short token, so +# a wrapped typed row such as `fix · tests` stays composer input; the ascii +# preset's `pi` is deliberately absent because that preset's `sep.dot` is +# ` - `, so its status row never carries a middle dot and a `pi ·` alternative +# could only ever match typed text), when it opens with one of omp's spinner +# frames then an elapsed cell, or when it carries the context-usage cell after +# a middle dot. It is consulted only as the boundary BELOW a bare composer, +# never on the composer row itself. +FM_COMPOSER_OMP_STATUS_RE_DEFAULT='^[[:space:]]*(π|󰵗)[[:space:]]+·[[:space:]]|^[[:space:]]*'"$FM_OMP_SPINNER_FRAMES_RE"'[[:space:]]+[0-9]+[smh]([[:space:]]|$)|[[:space:]]·[[:space:]].*[0-9]+(\.[0-9]+)?%/[0-9]+K' # The bounded row window adapters should capture for a composer read. One # shared policy (previously three per-backend variables that had drifted to @@ -928,6 +965,13 @@ _fm_composer_classify_bare_row() { # <screen> <styled> <row> printf '%s' "$state" } +# _fm_composer_row_is_omp_status: 0 when the trimmed row is omp's status line +# (FM_COMPOSER_OMP_STATUS_RE_DEFAULT above) - composer furniture that sits +# below a bare composer and must bound its wrap region exactly as an edge does. +_fm_composer_row_is_omp_status() { # <trimmed-row> + fm_composer_idle_matches "$1" "${FM_COMPOSER_OMP_STATUS_RE:-$FM_COMPOSER_OMP_STATUS_RE_DEFAULT}" sensitive +} + # _fm_composer_wrap_region_ok: 0 when every row STRICTLY BELOW <glyph-row> # through <cursor-row> is non-blank and carries no structural edge - the # contiguity proof that those rows are the bare composer's wrapped input @@ -941,6 +985,7 @@ _fm_composer_wrap_region_ok() { # <plain-screen> <glyph-row> <cursor-row> fm_composer_normalize_trim_var trimmed [ -n "$trimmed" ] || return 1 if fm_composer_row_has_edge "$trimmed"; then return 1; fi + if _fm_composer_row_is_omp_status "$trimmed"; then return 1; fi if fm_composer_leading_shell_glyph_var glyph "$trimmed"; then return 1; fi row=$((row + 1)) done @@ -1077,6 +1122,7 @@ _fm_composer_select_cursorless() { fm_composer_normalize_trim_var trimmed [ -n "$trimmed" ] || break fm_composer_row_has_edge "$trimmed" && break + _fm_composer_row_is_omp_status "$trimmed" && break FM_COMPOSER_SELECTED_LAST=$next next=$((next + 1)) done diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index 0b3ec94f091..de54245ad22 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -63,7 +63,7 @@ FM_SHARED_CAPTAIN_MODE="444" # The declared inheritable set (space-separated, config-dir-relative item paths). # Extend here to inherit more of the primary's local config; override via the # environment only in tests. Items must not contain whitespace. -FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context}" +FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context launch-env-allowlist}" # Items whose value is a home-SESSION enablement decision rather than durable # local configuration. They are inherited at the launch convergence point, where @@ -93,9 +93,17 @@ fm_config_inherit_items() { printf '%s\n' "$FM_SHARED_CAPTAIN_REL" } +fm_config_source_present() { + perl -MErrno=ENOENT -e ' + if (lstat $ARGV[0]) { print 1 } + elsif ($! == ENOENT) { print 0 } + else { die "error: cannot inspect configuration source at $ARGV[0]: $!\n" } + ' -- "$1" +} + fm_inherit_file_mode() { if [ "$(uname)" = Darwin ]; then - stat -f %Lp "$1" 2>/dev/null + /usr/bin/stat -f %Lp "$1" 2>/dev/null else stat -c %a "$1" 2>/dev/null fi @@ -103,7 +111,7 @@ fm_inherit_file_mode() { fm_inherit_file_device() { if [ "$(uname)" = Darwin ]; then - stat -f %d "$1" 2>/dev/null + /usr/bin/stat -f %d "$1" 2>/dev/null else stat -c %d "$1" 2>/dev/null fi @@ -111,7 +119,7 @@ fm_inherit_file_device() { fm_inherit_file_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi @@ -175,12 +183,13 @@ destination_allows_inherited_item() { # so this writes nothing there. It emits concise stderr diagnostics only for # notable events: a guard skip or a copy/remove error. A source item that is # present is copied only when its content differs (idempotent: a re-run never -# churns mtimes). A source item that is absent is mirrored as a missing +# churns mtimes). A source item proven absent is mirrored as a missing # destination item, so clearing the primary's value clears it downstream too -# (primary-authoritative). The destination dir is created lazily, only when there -# is actually something to write, so a primary with no inherited config item set is a -# complete no-op (it leaves the secondmate home exactly as it was - the -# backward-compatible path). When FM_CONFIG_INHERIT_REPORT points at a writable +# (primary-authoritative). Inspection errors or existing nonregular sources +# leave that destination item unchanged and report an error; inaccessible paths +# and dangling source links must never silently remove an inherited grant. +# The destination dir is created lazily, only when there is something to copy; +# absence on both sides is a no-op. When FM_CONFIG_INHERIT_REPORT points at a writable # file, one tab-separated line per item is appended there: # <item> <status> <reason> # Status is pushed, unchanged, skipped, or error. Skipped items are warnings and @@ -440,7 +449,7 @@ propagate_secondmate_inheritance() { } propagate_inheritable_config() { - local src_config=$1 dest_config=$2 item src dest reason rc + local src_config=$1 dest_config=$2 item src dest source_present reason rc [ -n "$src_config" ] || return 1 [ -n "$dest_config" ] || return 1 rc=0 @@ -454,6 +463,13 @@ propagate_inheritable_config() { fi src="$src_config/$item" dest="$dest_config/$item" + if ! source_present=$(fm_config_source_present "$src"); then + reason="cannot inspect primary source" + warn_inheritable_config_error "$item" "$src" "$reason" + record_inheritable_config_result "$item" error "$reason" + rc=1 + continue + fi # This one scalar config is consumed as a local safety boundary, so reject # every unsafe or malformed source/destination artifact before the generic # byte-copy behavior below can treat it as ordinary inherited material. @@ -514,6 +530,11 @@ propagate_inheritable_config() { else record_inheritable_config_result "$item" unchanged "" fi + elif [ "$source_present" = 1 ]; then + reason="primary source is not a regular file" + warn_inheritable_config_error "$item" "$src" "$reason" + record_inheritable_config_result "$item" error "$reason" + rc=1 elif [ -e "$dest" ] || [ -L "$dest" ]; then if ! destination_allows_inherited_item "$dest_config" "$item"; then reason=$(inheritable_config_skip_reason) diff --git a/bin/fm-control-lib.sh b/bin/fm-control-lib.sh index 5a9a291982f..85c7f0e3d59 100644 --- a/bin/fm-control-lib.sh +++ b/bin/fm-control-lib.sh @@ -37,7 +37,7 @@ # `resume` is deliberately NOT a verb. It is not deterministic across the # verified adapters: codex, grok, and gemini resume only from a session id printed at # exit, opencode resumes the most recent session for the cwd with --continue, -# and claude, copilot, pi, pi-signed, and kimi have no verified pane-resume +# and claude, copilot, pi, pi-signed, omp, and kimi have no verified pane-resume # contract at all. `relaunch` covers the same need deterministically for every adapter, # because the brief on disk - not a harness-private session - is the durable # instruction. @@ -63,7 +63,7 @@ fm_control_verb_allowed() { # <verb> # than guessed at, exactly as a spawn on it would be. fm_control_harness_supported() { # <harness> case "${1-}" in - claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse) return 0 ;; + claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) return 0 ;; esac return 1 } @@ -74,12 +74,13 @@ fm_control_harness_supported() { # <harness> # harness= that way), which is why the spawn adapters match `claude*`, `muse*`, # and friends. This is the one place that prefix rule is stated. `pi` and # `pi-signed` are exact because a `pi*` prefix would swallow the signed adapter, -# and an unrecognized value returns nonzero rather than being guessed into a -# family. +# `omp` is exact because an `omp*` prefix would claim unrelated commands, and an +# unrecognized value returns nonzero rather than being guessed into a family. fm_control_harness_family() { # <recorded-harness> case "${1-}" in pi) printf 'pi' ;; pi-signed) printf 'pi-signed' ;; + omp) printf 'omp' ;; claude*) printf 'claude' ;; codex*) printf 'codex' ;; copilot*) printf 'copilot' ;; @@ -89,21 +90,22 @@ fm_control_harness_family() { # <recorded-harness> cursor*) printf 'cursor' ;; gemini*) printf 'gemini' ;; muse*) printf 'muse' ;; + rovo*) printf 'rovo' ;; *) return 1 ;; esac } -# Which task kinds an adapter is verified to run. muse and gemini are -# crewmate/scout adapters only: neither has a primary supervision protocol, -# and bin/fm-spawn.sh refuses a --secondmate launch on either. The control plane -# asks this BEFORE it stops anything, so an incompatible relaunch target is +# Which task kinds an adapter is verified to run. muse, gemini, and rovo are +# crewmate/scout adapters only: none has a primary supervision protocol, +# and bin/fm-spawn.sh refuses a --secondmate launch on any of them. The control +# plane asks this BEFORE it stops anything, so an incompatible relaunch target is # refused while the current agent is still running rather than after it has # been stopped. fm_control_harness_supports_kind() { # <harness> <kind> local harness=${1-} kind=${2-} fm_control_harness_supported "$harness" || return 1 case "$harness" in - muse|gemini) [ "$kind" != secondmate ] || return 1 ;; + muse|gemini|rovo) [ "$kind" != secondmate ] || return 1 ;; esac return 0 } @@ -112,9 +114,13 @@ fm_control_harness_supports_kind() { # <harness> <kind> # whose Esc only moves focus to the scrollback; grok cancels on Ctrl+C. # gemini names its own key in the running turn's status row # (`(esc to cancel, <n>s)`), and a single Escape was verified to cancel it. +# rovo cancels on a single Escape too, printing "Agent cancelled" (verified, +# 202609.1.2). omp (Oh My Pi) shares Pi's single Escape, empty composer +# afterwards, and /quit exit (verified omp 18.1.2 in a PTY, re-verified 18.1.11 +# through Herdr). fm_control_interrupt_key() { # <harness> case "${1-}" in - claude|codex|opencode|pi|pi-signed|kimi|cursor|gemini|muse) printf 'Escape' ;; + claude|codex|opencode|pi|pi-signed|omp|kimi|cursor|gemini|muse|rovo) printf 'Escape' ;; copilot|grok) printf 'C-c' ;; *) return 1 ;; esac @@ -125,7 +131,7 @@ fm_control_interrupt_key() { # <harness> fm_control_interrupt_repeat() { # <harness> case "${1-}" in opencode) printf '2' ;; - claude|codex|copilot|pi|pi-signed|grok|kimi|cursor|gemini|muse) printf '1' ;; + claude|codex|copilot|pi|pi-signed|omp|grok|kimi|cursor|gemini|muse|rovo) printf '1' ;; *) return 1 ;; esac } @@ -146,7 +152,7 @@ fm_control_interrupt_repeat() { # <harness> fm_control_interrupt_clear_key() { # <harness> case "${1-}" in muse) printf 'C-u' ;; - claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini) ;; + claude|codex|copilot|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo) ;; *) return 1 ;; esac } @@ -158,7 +164,10 @@ fm_control_interrupt_ack_source() { # <harness> # after an interrupt was measured as variable - sometimes seconds, sometimes # not within 20 - so a cancellation claim built on it would be unreliable. # Normal turn completion is prompt, which is what the busy fold depends on. - claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini) printf 'none' ;; + # rovo's TUI prints "Agent cancelled" on Escape, but for parity with + # claude/cursor this stays 'none': the ack is a rendered string, not a + # recorded state source, and rovo has no busy wiring to confirm against. + claude|codex|copilot|opencode|pi|pi-signed|omp|grok|kimi|cursor|gemini|rovo) printf 'none' ;; *) return 1 ;; esac } @@ -166,8 +175,8 @@ fm_control_interrupt_ack_source() { # <harness> # The command that exits the agent from its own composer. fm_control_exit_command() { # <harness> case "${1-}" in - claude|copilot|opencode|grok|kimi|cursor|muse) printf '/exit' ;; - codex|pi|pi-signed|gemini) printf '/quit' ;; + claude|copilot|opencode|grok|kimi|cursor|muse|rovo) printf '/exit' ;; + codex|pi|pi-signed|omp|gemini) printf '/quit' ;; *) return 1 ;; esac } @@ -214,6 +223,7 @@ fm_control_harness_wiring_paths() { # <harness> <worktree> <state-dir> <id> claude) printf '%s\n' "$wt/.claude/settings.local.json" ;; opencode) printf '%s\n' "$wt/.opencode/plugins/fm-busy-state.js" ;; pi|pi-signed) printf '%s\n' "$state/$id.pi-ext.ts" ;; + omp) printf '%s\n' "$state/$id.omp-ext.ts" ;; grok) printf '%s\n' "$wt/.fm-grok-turnend" printf '%s\n' "$state/$id.grok-turnend-token" diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 0ecbb9670c9..f3a99c3e3e5 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -50,11 +50,27 @@ # the active step is ci, `axi status` alone cannot tell "still waiting on # checks" from "checks green, waiting on merge" (see nm_ci_checks_state) - # a ci-step log-tail check overrides working -> done once checks read -# green, so a green PR is never silently read as still-validating. +# green, so a green PR is never silently read as still-validating. And a +# terminal FAILED run whose only failure is the ci monitor step, after +# every substantive step completed and the ci log's last marker reads +# checks green, also reads done (held-for-merge), never failed: a monitor +# whose only remaining job is to observe a human merge decision must not +# convert the absence of that decision into a failure verdict +# (nm_failed_run_is_green_held_ci; 2026-09-05 jr-voice incident). In the +# coarse runs-ledger fallback (no steps table, no ci log), a terminal +# FAILED record whose daemon an explicit probe proves down reads unknown, +# never failed: an instrument failure must not read as work failure +# (nm_daemon_probe_down). # 3. Reconcile the status log: if its last line says needs-decision/blocked but # the run-step shows the run moved on, the log is deterministically stale and # is flagged superseded. A genuinely parked run plus a needs-decision log -# agree, and are reported as parked. +# agree, and are reported as parked. A `blocked:` line that reports a +# refused or missing daemon socket remains blocked even if an attributed +# run record is stale or terminal. Other daemon, timeout, or unreachability +# claims are superseded BECAUSE THE RUN IS ALIVE when the run is +# running/fixing with recent reported activity: a killed or timed-out drive +# call is not daemon death, so that claim is answered by steering the crew +# to reattach, not by escalating. # 4. No run for this crew (pre-validation, or kind=scout): fall back to the # recorded backend's pane busy state, then the status log's last line only # when its verb maps to a recognized run-state. Decision-only events such as @@ -315,6 +331,139 @@ log_reports_ci_ready() { esac } +# 0 when a status-log line reports positive daemon socket failure rather than a +# client-side timeout or generic unreachability. +log_reports_daemon_socket_down() { # <line> + local line + line=$(printf '%s' "$1" | tr '[:upper:]' '[:lower:]') + case "$line" in + *daemon*|*no-mistakes*) ;; + *) return 1 ;; + esac + case "$line" in + *"connection refused"*|*"connections refused"*|*"socket refused connection"*|*"socket refuses connection"*|*"socket refusing connection"*|*"socket missing"*|*"socket is missing"*|*"missing socket"*) return 0 ;; + esac + return 1 +} + +# 0 when a status-log line blames the pipeline's transport rather than the work. +# None of these claims alone is evidence the daemon died: a drive call is only +# waiting for a read while the fix round runs in the background. +log_claims_pipeline_unreachable() { # <line> + case "$(printf '%s' "$1" | tr '[:upper:]' '[:lower:]')" in + *daemon*|*timeout*|*"timed out"*|*unreachab*) return 0 ;; + esac + return 1 +} + +# Rows of the `active_steps[N]{...}:` table in the captured run output +# ($RUN_OUT), which the pipeline emits only while a step is actually running or +# fixing. Column order is deliberately not assumed: the header's own indentation +# bounds the block, and callers below read the table as text. +nm_active_steps_rows() { + printf '%s\n' "$RUN_OUT" | awk ' + /^[[:space:]]*active_steps\[[0-9]+\]\{/ { hdr = index($0, "active_steps"); inblock = 1; next } + inblock { + if ($0 ~ /^[[:space:]]*$/) { inblock = 0; next } + match($0, /[^ \t]/) + if (RSTART <= hdr) { inblock = 0; next } + print + } + ' +} + +# Rows of the `steps[N]{step,status,findings,duration_ms}:` table in the +# captured run output ($RUN_OUT) - the full per-step ledger, present on +# terminal runs too, unlike active_steps[] which the pipeline emits only while +# a step is actually running or fixing. Column order is deliberately not +# assumed: the header's own indentation bounds the block, and callers below +# read the table as text. +nm_steps_rows() { + printf '%s\n' "$RUN_OUT" | awk ' + /^[[:space:]]*steps\[[0-9]+\]\{/ { hdr = index($0, "steps"); inblock = 1; next } + inblock { + if ($0 ~ /^[[:space:]]*$/) { inblock = 0; next } + match($0, /[^ \t]/) + if (RSTART <= hdr) { inblock = 0; next } + print + } + ' +} + +# 0 when the pipeline itself reports RECENT activity on an actively running or +# fixing step. The client prefixes a step's `last_activity` with `quiet` once no +# step log or native-agent lifecycle event has arrived for longer than its +# configured quiet warning, so its own recency verdict is the signal here rather +# than a second threshold invented in firstmate. Positive evidence is required: +# an absent table is not recency, so a run record that merely still says +# `running` while nothing executes it never reads as alive. +nm_run_activity_is_recent() { + local rows + rows=$(nm_active_steps_rows) + [ -n "$rows" ] || return 1 + ! printf '%s\n' "$rows" | grep -q 'quiet' +} + +# 0 when a terminal FAILED run's only failure is the ci monitor step and the +# ci log's last recognized marker reads checks green. Requires the exact +# shape, all on positive evidence: a steps[] table where every step completed +# except exactly `ci` failed (any other non-completed status, or a second +# failed step, disqualifies), plus nm_ci_checks_state=green (a genuinely red +# check, or an unreadable ci log, keeps the failure a failure). This is the +# orphaned-CI-monitor gap (2026-09-05 jr-voice): a run held for a captain +# merge decision polls until the shared daemon restarts under it and marks +# the run failed, although GitHub's own check state - the actual shippability +# authority - is green and every substantive step completed. +nm_failed_run_is_green_held_ci() { + local rows row rest step status saw_ci_failed + rows=$(nm_steps_rows) + [ -n "$rows" ] || return 1 + saw_ci_failed=0 + while IFS= read -r row; do + row=$(trim "$row") + step=$(trim "${row%%,*}") + rest=${row#*,} + status=$(strip_quotes "$(trim "${rest%%,*}")") + case "$status" in + completed) continue ;; + failed) + [ "$step" = ci ] || return 1 + saw_ci_failed=1 + continue + ;; + *) return 1 ;; + esac + done <<EOF +$rows +EOF + [ "$saw_ci_failed" = 1 ] || return 1 + [ "$(nm_ci_checks_state)" = green ] +} + +# Reclassify a terminal failed run as done (held-for-merge) when +# nm_failed_run_is_green_held_ci matches, surfacing the run's PR URL so the +# supervisor reads the concrete review-ready outcome instead of a failure. +nm_reclassify_failed_run_as_held_green() { + nm_failed_run_is_green_held_ci || return 1 + RUN_STATE="done" + RUN_DETAIL="checks green: PR held for merge (ci monitor ended)" + local pr_url + pr_url=$(strip_quotes "$(nm_field pr)") + [ -n "$pr_url" ] && RUN_DETAIL="$RUN_DETAIL: $pr_url" + return 0 +} + +# 0 when an explicit probe proves the shared daemon down: `no-mistakes daemon +# status` is the canonical down-probe (the same one fm-brief.sh hands crews +# before a blocked append) and exits non-zero when the daemon is not running. +# Bounded like every other CLI call; a probe that fails for any reason - +# refused socket, timeout, non-zero answer - means the daemon is not provably +# up, which is the only fact the coarse fallback needs. +nm_daemon_probe_down() { + fm_nm_run_checked "$WT" "$NM_TIMEOUT" daemon status >/dev/null || return 0 + return 1 +} + nm_ci_step_status() { local row rest row=$(printf '%s\n' "$RUN_OUT" | grep -E '^[[:space:]]*ci,[[:space:]]*"?(running|fixing)"?[[:space:]]*,' | head -1) @@ -464,7 +613,17 @@ if [ "$HAVE_RUN" = 1 ]; then case "$COARSE_STATUS" in running) RUN_STATE=working; RUN_DETAIL="validating (background run)" ;; completed) RUN_STATE="done"; RUN_DETAIL="run completed" ;; - failed) RUN_STATE=failed; RUN_DETAIL="run failed" ;; + failed) + # The ledger row is terminal but the coarse path has no steps table + # and no ci log, so the orphaned-monitor shape cannot be recognized + # here. With the daemon provably down, the row is unverified evidence + # from a dead instrument and must not read as work failure. + if nm_daemon_probe_down; then + RUN_STATE=unknown + RUN_DETAIL="no-mistakes daemon unreachable; last ledger record failed - unverified" + else + RUN_STATE=failed; RUN_DETAIL="run failed" + fi ;; cancelled) RUN_STATE=failed; RUN_DETAIL="run cancelled" ;; *) RUN_STATE=unknown; RUN_DETAIL="runs list status: $COARSE_STATUS" ;; esac @@ -481,7 +640,10 @@ if [ "$HAVE_RUN" = 1 ]; then case "$outcome" in passed) RUN_STATE="done"; RUN_DETAIL="run passed: PR merged/closed" ;; checks-passed) RUN_STATE="done"; RUN_DETAIL="checks green: PR ready for review" ;; - failed) RUN_STATE=failed; RUN_DETAIL="run failed" ;; + failed) + if nm_reclassify_failed_run_as_held_green; then :; else + RUN_STATE=failed; RUN_DETAIL="run failed" + fi ;; cancelled) RUN_STATE=failed; RUN_DETAIL="run cancelled" ;; *) RUN_STATE=unknown; RUN_DETAIL="outcome: $outcome" ;; esac @@ -505,7 +667,10 @@ if [ "$HAVE_RUN" = 1 ]; then ci) RUN_STATE=working; RUN_DETAIL="ci running" ;; running|fixing) RUN_STATE=working; RUN_DETAIL="validating ($status)" ;; completed) RUN_STATE="done"; RUN_DETAIL="run completed" ;; - failed) RUN_STATE=failed; RUN_DETAIL="run failed" ;; + failed) + if nm_reclassify_failed_run_as_held_green; then :; else + RUN_STATE=failed; RUN_DETAIL="run failed" + fi ;; cancelled) RUN_STATE=failed; RUN_DETAIL="run cancelled" ;; "") RUN_STATE=working; RUN_DETAIL="run active" ;; *) RUN_STATE=working; RUN_DETAIL="run active ($status)" ;; @@ -548,11 +713,29 @@ if [ "$HAVE_RUN" = 1 ]; then # Reconcile the status log. A needs-decision/blocked log line that the run-step # has moved past (anything but a genuinely parked run) is deterministically # stale: the gate resolved and the run resumed or finished. + # + # A refused or missing daemon socket is positive daemon-down evidence and + # outranks any attributed run record, including a terminal one left behind + # after the daemon stopped. Other blocked claims caused by a timed-out drive + # call are contradicted only when the run reports recent + # activity; the answer is then to steer the crew to reattach without touching + # the shared daemon. case "$LOG_VERB" in needs-decision|blocked) + if [ "$LOG_VERB" = blocked ] \ + && log_reports_daemon_socket_down "$LOG_LINE"; then + emit blocked status-log "$(status_line_note "$LOG_LINE")${SEP}daemon socket down despite attributed run record" + fi if [ "$RUN_STATE" != parked ]; then if [ "$RUN_STATE" = working ]; then - RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded by active run" + if [ "$LOG_VERB" = blocked ] \ + && log_claims_pipeline_unreachable "$LOG_LINE" \ + && { [ "$RUN_STATUS" = running ] || [ "$RUN_STATUS" = fixing ]; } \ + && nm_run_activity_is_recent; then + RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded: run alive, not a daemon failure (steer reattach)" + else + RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded by active run" + fi else RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded (run $RUN_STATE)" fi diff --git a/bin/fm-dod-lib.sh b/bin/fm-dod-lib.sh index c5be1b1455c..07a7b46e242 100755 --- a/bin/fm-dod-lib.sh +++ b/bin/fm-dod-lib.sh @@ -22,6 +22,22 @@ # restating the rule. # Every heredoc here stays outside a command substitution: `VAR=$(cat <<EOF ...)` # breaks parsing of the whole file on Bash 3.2 (tests/fm-brief.test.sh). +# fm_brief_worker_role owns the ship/scout role scope. bin/fm-spawn.sh is its one +# emitter, supplying it to every ship/scout launch brief and never to a +# secondmate charter. Like fm_brief_intent_overlay it is a distinctly titled +# launch section that states its own precedence for Firstmate tasks, so a brief +# that authors its own role wording is superseded rather than duplicated. + +fm_brief_worker_role() { + cat <<'EOF' +# Current worker role contract +When this task works on Firstmate itself, this section supersedes every earlier brief instruction about your role and identity. +When this task works on Firstmate itself, the repository root `AGENTS.md` (also imported by `CLAUDE.md`) is the primary/secondmate supervisor's contract: follow this brief instead of that supervisor contract. +For that Firstmate task, do the assigned work yourself and report to firstmate; do not adopt the supervisor identity, delegate the task, run fleet supervision, or address the captain. +This exception preserves this brief's safety and authority boundaries and applicable contributor guidance, including `CONTRIBUTING.md` and `firstmate-coding-guidelines` for Firstmate changes. +Other projects retain their own instructions unchanged. +EOF +} # Return 0 when a Task subsection still consists only of its scaffold # placeholder. A missing file and legacy briefs carry no such placeholders. @@ -216,6 +232,12 @@ When the captain's intent refers to a report, decision, or PR ("do items 1, 2, 3 This replaces the no-mistakes skill's advice to enrich \`--intent\` with decisions and tradeoffs; that advice does not apply to Firstmate-dispatched work. Do not hand-edit, commit, or fix findings yourself while a run is active - the pipeline applies every fix. +One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain. +So background the drive call and poll \`no-mistakes axi status\` from a separate call instead of sitting in one blocking hold your harness will kill. +Where a harness's own command limit is not established, assume it bounds commands and use that same background-and-poll shape. +A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working. +Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real. + Two firstmate-specific rules layer on top of that guidance: - ask-user findings are never yours to answer: escalate to firstmate using rule 6's ask-user format and stop. Firstmate applies \`ask-user-authority\` and obtains any required captain decision. diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index 5ef8ffc35d1..f127f7d5211 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -1086,8 +1086,8 @@ case "$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" in ''|*[!0-9]*) FM_SNAPSHOT_SECON # pollute arithmetic input before failing. Select the platform syntax once. if [ "$(uname 2>/dev/null || true)" = Darwin ]; then SNAPSHOT_STAT_STYLE=bsd - file_mtime_epoch() { stat -f '%m' "$1" 2>/dev/null || true; } - file_mode_octal() { stat -f '%Lp' "$1" 2>/dev/null || true; } + file_mtime_epoch() { /usr/bin/stat -f '%m' "$1" 2>/dev/null || true; } + file_mode_octal() { /usr/bin/stat -f '%Lp' "$1" 2>/dev/null || true; } else SNAPSHOT_STAT_STYLE=gnu file_mtime_epoch() { stat -c '%Y' "$1" 2>/dev/null || true; } @@ -1440,7 +1440,7 @@ bounded_parent_activities_json() { # <status-file> stat_style=$6 . "$classify" if [ "$stat_style" = bsd ]; then - size=$(stat -f "%z" "$f" 2>/dev/null) || exit 3 + size=$(/usr/bin/stat -f "%z" "$f" 2>/dev/null) || exit 3 else size=$(stat -c "%s" "$f" 2>/dev/null) || exit 3 fi diff --git a/bin/fm-guard.sh b/bin/fm-guard.sh index 0b2a34a824e..9dd9b20a31b 100755 --- a/bin/fm-guard.sh +++ b/bin/fm-guard.sh @@ -5,9 +5,9 @@ # First, always warn if the firstmate primary checkout (FM_ROOT) is on a named # non-default branch, because that means firstmate-on-itself work landed in the # primary instead of an isolated worktree. -# Then, if a task is in flight (a state/<id>.meta exists) or X-mode relay -# polling is active (state/x-watch.check.sh exists) and supervision is not -# healthy, prints a loud, clearly delimited banner so the agent cannot skim past +# Then, if the home needs supervision (bin/fm-supervision-lib.sh owns that +# condition set) and that supervision is not healthy, prints a loud, clearly +# delimited banner so the agent cannot skim past # it in the tool output of whatever it was doing - the one channel every harness # has. Supervision health is MODEL-AWARE (fm_watcher_supervision_verdict in # bin/fm-wake-lib.sh): under the Claude Stop auto-arm model the watcher runs only @@ -27,7 +27,13 @@ # bounded). Independent alarms (queued wakes, worktree tangle) are never # suppressed by that dedup. Normal wake handling (watcher briefly down between a # wake and the next supervision resume) stays inside the grace window and stays -# silent. The queued-wakes warning stays silent for the supervision branch +# silent. The queued-wakes warning counts only the rows the calling actor can +# itself present or retire (fm_wake_actor_pending_count), so it is never an +# instruction to run a drain with nothing to present. A row reserved by a live +# supervision-branch grant is never a drain instruction for main; instead of +# going silent about a visibly non-empty queue, main gets a distinct advisory +# naming the branch as the holder and saying not to drain those rows. +# The ordinary warning also stays silent for the supervision branch # actor (FM_SUPERVISION_ACTOR=branch), because that actor runs guarded commands # while handling exactly the queued rows its grant covers and can drain nothing # else. Always exits 0: the guard warns, it never blocks. @@ -41,6 +47,7 @@ CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" WATCH="$SCRIPT_DIR/fm-watch.sh" GRACE=${FM_GUARD_GRACE:-300} queue_pending=false +queue_branch_held=false READ_ONLY=${FM_GUARD_READ_ONLY:-0} case "$READ_ONLY" in 1|true|TRUE|yes|YES) READ_ONLY=1 ;; *) READ_ONLY=0 ;; esac CONTINUE_LINE=${FM_GUARD_CONTINUE_LINE:-This is a supervision warning only; the guarded operation WILL still run.} @@ -158,11 +165,12 @@ if [ -n "$tangle_branch" ]; then fi # Compute supervision need and watcher-beacon freshness via the shared -# grace-based predicate (bin/fm-supervision-lib.sh). Act when work, an event -# source, or an X-mode relay poll needs supervision. +# grace-based predicate (bin/fm-supervision-lib.sh), which owns what needs +# supervision. fm_supervision_status "$STATE" "$GRACE" in_flight=$FM_SUP_IN_FLIGHT sources=$FM_SUP_SOURCES +checks=$FM_SUP_CHECKS needed=$FM_SUP_NEEDED beacon_desc=$FM_SUP_BEACON_DESC fm_watcher_supervision_verdict "$STATE" "$WATCH" "$GRACE" "$FM_HOME" "$FM_ROOT" @@ -176,7 +184,18 @@ if [ "$needed" = false ]; then exit 0 fi -[ -s "$FM_WAKE_QUEUE" ] && queue_pending=true +# Count only the rows this actor could actually present or retire, so the +# warning never sends an actor to a drain that provably has nothing for it. +# fm-wake-lib.sh owns that per-actor classification. A non-empty queue with +# nothing for main is the branch-held case: keep the raw pending signal visible +# there as its own advisory rather than dropping it. +if [ -s "$FM_WAKE_QUEUE" ]; then + if [ "$(fm_wake_actor_pending_count "$GUARD_ACTOR")" -gt 0 ]; then + queue_pending=true + elif [ "$GUARD_ACTOR" != branch ] && [ "$(fm_wake_actor_pending_count branch)" -gt 0 ]; then + queue_branch_held=true + fi +fi # No fresh watcher with tasks in flight is the dangerous state: emit a prominent, # bordered banner FIRST so it reads as an alarm, not a buried stderr line. Later @@ -216,6 +235,8 @@ if [ "$watcher_healthy" = false ]; then printf '● %s task(s) in flight, but %s.\n' "$in_flight" "$watcher_cause" elif [ "$sources" -gt 0 ]; then printf '● %s process-event source(s) registered, but %s.\n' "$sources" "$watcher_cause" + elif [ "$checks" -gt 0 ]; then + printf '● %s registered custom check(s), but %s.\n' "$checks" "$watcher_cause" else printf '● X-mode relay polling needs supervision, but %s.\n' "$watcher_cause" fi @@ -253,5 +274,7 @@ if "$queue_pending"; then elif [ "$GUARD_ACTOR" != branch ]; then echo "WARNING: queued wakes pending - drain them with bin/fm-wake-drain.sh before anything else." >&2 fi +elif "$queue_branch_held"; then + echo "NOTICE: wake rows held by the live supervision branch - it presents and acknowledges them; do not drain them from here." >&2 fi exit 0 diff --git a/bin/fm-harness.sh b/bin/fm-harness.sh index f9684b19152..f58e2f5b7a7 100755 --- a/bin/fm-harness.sh +++ b/bin/fm-harness.sh @@ -1,6 +1,6 @@ #!/usr/bin/env bash # Detect the agent harness this process tree runs on. -# Usage: fm-harness.sh print own harness: claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|unknown +# Usage: fm-harness.sh print own harness: claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp|unknown # fm-harness.sh crew print the effective CREWMATE harness # (config/crew-harness; "default" resolves to own) # fm-harness.sh secondmate print the harness the PRIMARY uses to launch @@ -73,6 +73,30 @@ detect_own() { # carrying the claude primary's value (claude-code_2-1-260_agent), so it is # an inherited launcher marker, not a Gemini identity. [ "${GEMINI_CLI:-}" = "1" ] && { echo gemini; return; } + # rovo (Atlassian Rovo CLI) sets ATLASSIAN_AGENT_TYPE=rovo, ROVODEV_CLI=1, and + # AGENT=rovodev_cli on its tool subprocesses (verified, rovo 202609.1.2). It does + # NOT scrub an inherited CLAUDECODE, so a rovo worker launched from a claude + # session carries both markers - this must be tested BEFORE the CLAUDECODE line, + # the same ordering hazard cursor documents above (see issue #3517). bin/fm-spawn.sh + # additionally clears foreign markers at rovo's launch boundary as defense in depth. + [ "${ATLASSIAN_AGENT_TYPE:-}" = "rovo" ] && { echo rovo; return; } + [ "${ROVODEV_CLI:-}" = "1" ] && { echo rovo; return; } + # omp (Oh My Pi) publishes NO harness-identity marker of its own: verified on + # omp 18.1.11 that PI_CODING_AGENT is absent from the binary and that the + # default profile sets neither PI_CODING_AGENT_DIR nor OMP_PROFILE in the + # process environment. FM_OMP_HARNESS=omp is therefore a Firstmate-OWNED + # launch marker, established by bin/fm-spawn.sh at the omp launch boundary + # (which also clears every foreign marker) and by the README's primary launch + # command. It is a PRECEDENCE override, never evidence on its own: it wins + # over an inherited CLAUDECODE only when an omp process is genuinely in the + # ancestry, so `FM_OMP_HARNESS=omp omp` started from a Claude pane identifies + # as omp, while the same variable leaking from an omp secondmate into that + # home's claude worker (whose ancestry holds no omp) changes nothing. The + # anchored ancestry arm below covers a plain hand-started `omp` by itself. + if [ "${FM_OMP_HARNESS:-}" = omp ] && ancestry_names_omp; then + echo omp + return + fi [ "${CLAUDECODE:-}" = "1" ] && { echo claude; return; } if [ "${PI_CODING_AGENT:-}" = "true" ]; then if [ "${FM_PI_HARNESS:-}" = pi-signed ]; then echo pi-signed; else echo pi; fi @@ -128,6 +152,7 @@ detect_own() { *opencode*) echo opencode; return ;; *grok*) echo grok; return ;; kimi) echo kimi; return ;; + rovo) echo rovo; return ;; # muse's installed launcher ~/.local/bin/muse execs ~/.local/bin/muse-bin-<version> # (verified in the published launcher, muse 0.1.0-R708.1), so the live process # name carries the version and CHANGES on every auto-update. Match the stable @@ -136,6 +161,15 @@ detect_own() { muse|muse-bin-*) echo muse; return ;; pi-signed) echo pi; return ;; pi) echo pi; return ;; + # omp is a Bun-compiled single binary whose process name is exactly `omp` + # (verified, omp 18.1.11: `ps -o comm=` reports omp from both its `!` + # bash path and the model's bash tool). Anchored, never *omp*, so ompd, + # comp, and similar unrelated commands are not misread as this harness. + # It sits above the node*|python* interpreter fallback deliberately: the + # optional claude-bridge extension runs a nested executable literally + # named `claude` with its own node child, and that fallback's *claude* + # args glob would otherwise claim it if that subtree were ever walked. + omp) echo omp; return ;; node*|python*) # Bare interpreter: match the harness name in its script path. args=$(ps -o args= -p "$pid" 2>/dev/null) @@ -160,6 +194,20 @@ detect_own() { echo unknown } +# True when an exact `omp` process sits within eight parents of this one. The +# same anchored match as the ancestry walk in detect_own, kept separate so the +# marker precedence above can demand real process evidence. +ancestry_names_omp() { + local pid=$$ comm + for _ in 1 2 3 4 5 6 7 8; do + comm=$(ps -o comm= -p "$pid" 2>/dev/null) || return 1 + [ "$(basename -- "$comm")" = omp ] && return 0 + pid=$(ps -o ppid= -p "$pid" 2>/dev/null | tr -d ' ') + [ -n "$pid" ] && [ "$pid" -gt 1 ] || return 1 + done + return 1 +} + # Resolve the effective crewmate harness: config/crew-harness (a bare adapter # name) wins; absent or "default" mirrors firstmate's own harness. resolve_crew() { diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 52064545836..9c30a9074be 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -121,7 +121,7 @@ if [ "$FM_INACTIVE_RECONCILE_BUDGET_SECS" -gt 30 ]; then fi if [ "$(uname)" = Darwin ]; then - file_mtime() { stat -f %m "$1" 2>/dev/null; } + file_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else file_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-lint.sh b/bin/fm-lint.sh index 3eb5b53609d..3ca44f802d2 100755 --- a/bin/fm-lint.sh +++ b/bin/fm-lint.sh @@ -3,27 +3,37 @@ # # Runs its file set with ShellCheck's default severity, extended analysis, # ambient configuration disabled, and one exact ShellCheck version. CI and -# no-mistakes both invoke this script with no arguments, so the rule set, -# version, bounded execution, and diagnostics ordering cannot drift. +# no-mistakes both invoke this script with no arguments, so this owner selects +# the context-appropriate rule set without duplicating lint configuration. # The explicit --fast mode is local-only and disables ShellCheck's extended -# dataflow analysis while preserving ordinary shell lint checks. CI and -# no-mistakes keep the full-analysis no-argument default. -# Tests stop source analysis at imported production modules because every -# production shell is already a canonical, source-aware root of this same run. +# dataflow analysis while preserving ordinary shell lint checks and source +# following. CI, main, and merge-base-less runs keep --norc --external-sources +# with full dataflow over the whole canonical set. An ordinary local branch +# (changed-file mode, including the no-mistakes lint step) drops +# --external-sources, keeps dataflow, and excludes SC1091, SC2034, SC2153, +# and SC2329, the codes that need library context. Those codes still run in +# CI over the whole set. Explicit paths keep --external-sources with the +# selected dataflow mode. +# Tests stop source analysis at imported production modules because CI analyzes +# every production shell separately as a canonical, source-aware root. # The default (no explicit-path) path also runs bin/fm-lint-workflows.sh so a # malformed GitHub workflow, including a self-broken ci.yml, fails locally # before merge instead of only failing to run as CI. # -# With no explicit paths, the file set depends on context: +# With no explicit paths, the file set and source-following posture depend +# on context: # - In CI (GITHUB_ACTIONS=true or CI=true), on the main branch, or when no # merge-base against origin/main (or local main) can be found, it lints -# the full canonical set: bin/*.sh bin/backends/*.sh tests/*.sh. This is -# what CI always runs, so CI coverage never depends on a local diff. +# the full canonical set: bin/*.sh bin/backends/*.sh tests/*.sh, with +# --external-sources and full dataflow. This is what CI always runs, so +# CI coverage never depends on a local diff. # - Otherwise (an ordinary local branch with a real merge-base) it lints # only the canonical-set files changed since that merge-base, including # uncommitted local edits, via plain local `git diff` (no network, no -# `gh`). A branch with zero matching changed files skips ShellCheck and -# prints a "no changed lint targets" note, then still validates workflows. +# `gh`). That local pass drops --external-sources and excludes SC1091, +# SC2034, SC2153, and SC2329. A branch with zero matching changed files +# skips ShellCheck and prints a "no changed lint targets" note, then +# still validates workflows. # Explicit paths always bypass this file-set selection and lint exactly the # given paths, matching the same config, without the workflow YAML check. # @@ -47,6 +57,9 @@ set -u REQUIRED_SHELLCHECK=0.11.0 +# Cross-file codes that need --external-sources. Local changed-file mode +# cannot judge them, so they stay CI-only. +LOCAL_NOX_EXCLUDE=SC1091,SC2034,SC2153,SC2329 SELF_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SELF="$SELF_DIR/fm-lint.sh" ROOT="$(cd "$SELF_DIR/.." && pwd)" @@ -62,7 +75,7 @@ fm_lint_worker_stop() { } fm_lint_worker() { # <manifest> <output-dir> <shard-index> - local manifest=$1 output_dir=$2 shard_index=$3 tab index path output rc=0 + local manifest=$1 output_dir=$2 shard_index=$3 tab index path output invocation_rc rc=0 local -a roots shellcheck_args roots=() tab=$(printf '\t') @@ -75,14 +88,34 @@ fm_lint_worker() { # <manifest> <output-dir> <shard-index> trap 'fm_lint_worker_stop; exit 129' HUP trap 'fm_lint_worker_stop; exit 130' INT trap 'fm_lint_worker_stop; exit 143' TERM - shellcheck_args=(--norc --external-sources) + shellcheck_args=(--norc) + if [ "${FM_LINT_INTERNAL_FOLLOW_SOURCES:-1}" -eq 1 ]; then + shellcheck_args+=(--external-sources) + fi + if [ -n "${FM_LINT_INTERNAL_EXCLUDE:-}" ]; then + shellcheck_args+=(--exclude="$FM_LINT_INTERNAL_EXCLUDE") + fi if [ "${FM_LINT_INTERNAL_FAST:-0}" -eq 1 ]; then shellcheck_args+=(--extended-analysis=false) fi - "$FM_LINT_SHELLCHECK" "${shellcheck_args[@]}" -- "${roots[@]}" > "$output.out" 2>&1 & - FM_LINT_WORKER_SHELLCHECK_PID=$! - wait "$FM_LINT_WORKER_SHELLCHECK_PID" || rc=$? - FM_LINT_WORKER_SHELLCHECK_PID= + : > "$output.out" + if [ "${FM_LINT_INTERNAL_FOLLOW_SOURCES:-1}" -eq 1 ]; then + "$FM_LINT_SHELLCHECK" "${shellcheck_args[@]}" -- "${roots[@]}" >> "$output.out" 2>&1 & + FM_LINT_WORKER_SHELLCHECK_PID=$! + wait "$FM_LINT_WORKER_SHELLCHECK_PID" || rc=$? + FM_LINT_WORKER_SHELLCHECK_PID= + else + for path in "${roots[@]}"; do + invocation_rc=0 + "$FM_LINT_SHELLCHECK" "${shellcheck_args[@]}" -- "$path" >> "$output.out" 2>&1 & + FM_LINT_WORKER_SHELLCHECK_PID=$! + wait "$FM_LINT_WORKER_SHELLCHECK_PID" || invocation_rc=$? + FM_LINT_WORKER_SHELLCHECK_PID= + if [ "$rc" -eq 0 ] && [ "$invocation_rc" -ne 0 ]; then + rc=$invocation_rc + fi + done + fi trap - HUP INT TERM else : > "$output.out" @@ -215,6 +248,8 @@ fm_lint_is_canonical_root() { CHANGED_MODE=0 EXPLICIT_PATHS=0 +FOLLOW_SOURCES=1 +EXCLUDE_CODES= if [ "$#" -gt 0 ]; then EXPLICIT_PATHS=1 ROOTS=("$@") @@ -242,6 +277,11 @@ else done < <(git diff --name-only --diff-filter=ACMR -z "$merge_base" -- 2>/dev/null | LC_ALL=C sort -z) fi fi +if [ "$CHANGED_MODE" -eq 1 ] && [ "$FAST" -eq 0 ]; then + FOLLOW_SOURCES=0 + EXCLUDE_CODES=$LOCAL_NOX_EXCLUDE + ANALYSIS_MODE=local +fi ROOT_COUNT=${#ROOTS[@]} if [ "$LIST_FILES" -eq 1 ]; then @@ -273,6 +313,8 @@ if [ "$resolved" != "$REQUIRED_SHELLCHECK" ]; then fi if [ "$FAST" -eq 1 ]; then printf 'fm-lint.sh: fast local mode; ShellCheck extended analysis disabled\n' >&2 +elif [ "$FOLLOW_SOURCES" -eq 0 ]; then + printf 'fm-lint.sh: local changed-file mode; ShellCheck source following disabled\n' >&2 else printf 'fm-lint.sh: full ShellCheck extended analysis enabled\n' >&2 fi @@ -408,18 +450,24 @@ fm_lint_run_worker() { # <worker-index> if [ "$(uname)" = Darwin ]; then exec "$PERL_BIN" -e 'setpgrp(0, 0) or die "setpgrp: $!"; exec @ARGV or die "exec: $!"' \ /usr/bin/time -lp -o "$timing" \ - env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ + env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" \ + FM_LINT_INTERNAL_FOLLOW_SOURCES="$FOLLOW_SOURCES" FM_LINT_INTERNAL_EXCLUDE="$EXCLUDE_CODES" \ + FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ "${BASH:-bash}" "$SELF" --internal-worker "$manifest" "$OUTPUT_DIR" "$worker_index" else exec "$PERL_BIN" -e 'setpgrp(0, 0) or die "setpgrp: $!"; exec @ARGV or die "exec: $!"' \ /usr/bin/time -f 'wall_seconds=%e\nuser_seconds=%U\nsystem_seconds=%S\nmax_rss_kib=%M' -o "$timing" \ - env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ + env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" \ + FM_LINT_INTERNAL_FOLLOW_SOURCES="$FOLLOW_SOURCES" FM_LINT_INTERNAL_EXCLUDE="$EXCLUDE_CODES" \ + FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ "${BASH:-bash}" "$SELF" --internal-worker "$manifest" "$OUTPUT_DIR" "$worker_index" fi else [ -z "$TELEMETRY" ] || printf 'timing_unavailable=1\n' > "$timing" exec "$PERL_BIN" -e 'setpgrp(0, 0) or die "setpgrp: $!"; exec @ARGV or die "exec: $!"' \ - env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ + env FM_LINT_INTERNAL=1 FM_LINT_INTERNAL_FAST="$FAST" \ + FM_LINT_INTERNAL_FOLLOW_SOURCES="$FOLLOW_SOURCES" FM_LINT_INTERNAL_EXCLUDE="$EXCLUDE_CODES" \ + FM_LINT_SHELLCHECK="$SHELLCHECK_BIN" \ "${BASH:-bash}" "$SELF" --internal-worker "$manifest" "$OUTPUT_DIR" "$worker_index" fi } @@ -505,7 +553,11 @@ if [ -n "$TELEMETRY" ]; then source_directives=$(wc -l < "$TMP_ROOT/source-targets" | tr -d '[:space:]') source_boundaries=$(grep -c '^/dev/null$' "$TMP_ROOT/source-targets" 2>/dev/null || true) case "$source_boundaries" in ''|*[!0-9]*) source_boundaries=0 ;; esac - source_followed=$((source_directives - source_boundaries)) + if [ "$FOLLOW_SOURCES" -eq 1 ]; then + source_followed=$((source_directives - source_boundaries)) + else + source_followed=0 + fi source_targets=$(LC_ALL=C sort -u "$TMP_ROOT/source-targets" | wc -l | tr -d '[:space:]') content_cksum=$(cksum "$TMP_ROOT/content-cksums" | awk '{print $1 "-" $2}') git_head=$(git rev-parse HEAD 2>/dev/null || printf 'unavailable') diff --git a/bin/fm-lock-lib.sh b/bin/fm-lock-lib.sh index f3b070ec8cf..7303ac571ad 100644 --- a/bin/fm-lock-lib.sh +++ b/bin/fm-lock-lib.sh @@ -24,7 +24,7 @@ fm_lock_log() { # no wake-queue machinery when a caller only needs the staleness proof. fm_lock_path_mtime() { if [ "$(uname)" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi diff --git a/bin/fm-pending-reply-lib.sh b/bin/fm-pending-reply-lib.sh index 7e4b00e4948..27c9bc9e90d 100755 --- a/bin/fm-pending-reply-lib.sh +++ b/bin/fm-pending-reply-lib.sh @@ -571,7 +571,7 @@ fm_pending_reply_file_signature() { # <path> local path=$1 [ -f "$path" ] || { printf 'missing'; return 0; } if [ "$(uname -s 2>/dev/null)" = Darwin ]; then - LC_ALL=C stat -f '%d:%i:%z:%m:%c' "$path" 2>/dev/null || printf 'unreadable' + LC_ALL=C /usr/bin/stat -f '%d:%i:%z:%m:%c' "$path" 2>/dev/null || printf 'unreadable' else LC_ALL=C stat -c '%d:%i:%s:%Y:%Z' "$path" 2>/dev/null || printf 'unreadable' fi diff --git a/bin/fm-pr-lib.sh b/bin/fm-pr-lib.sh index e934e96db8f..48f61ee8e12 100755 --- a/bin/fm-pr-lib.sh +++ b/bin/fm-pr-lib.sh @@ -215,7 +215,7 @@ fm_pr_head_valid() { fm_pr_file_mode() { if [ "$(uname)" = Darwin ]; then - stat -f %Lp "$1" 2>/dev/null + /usr/bin/stat -f %Lp "$1" 2>/dev/null else stat -c %a "$1" 2>/dev/null fi @@ -223,7 +223,7 @@ fm_pr_file_mode() { fm_pr_file_device() { if [ "$(uname)" = Darwin ]; then - stat -f %d "$1" 2>/dev/null + /usr/bin/stat -f %d "$1" 2>/dev/null else stat -c %d "$1" 2>/dev/null fi @@ -231,7 +231,7 @@ fm_pr_file_device() { fm_pr_file_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi @@ -239,7 +239,7 @@ fm_pr_file_link_count() { fm_pr_file_inode() { if [ "$(uname)" = Darwin ]; then - stat -f %i "$1" 2>/dev/null + /usr/bin/stat -f %i "$1" 2>/dev/null else stat -c %i "$1" 2>/dev/null fi diff --git a/bin/fm-procevent-extension-capture.pl b/bin/fm-procevent-extension-capture.pl index 3e877ae6c43..485c1a0343d 100644 --- a/bin/fm-procevent-extension-capture.pl +++ b/bin/fm-procevent-extension-capture.pl @@ -1,7 +1,7 @@ use strict; use warnings; use Cwd qw(getcwd); -use Fcntl qw(O_CREAT O_EXCL O_NOFOLLOW O_RDONLY O_RDWR); +use Fcntl qw(O_CREAT O_EXCL O_NOFOLLOW O_RDONLY O_RDWR O_WRONLY); use JSON::PP qw(encode_json); use POSIX qw(dup2); @@ -92,8 +92,12 @@ my ($registry_fd, $inbox_fd, $reservation_fd, $id, $adapter, $extension_id, $extension_version, $capability_version, $package_digest, $binding_digest, $claim_token, $runner_name, $output_name, $runner_pid, $claim_identity, $limit, @command) = @ARGV; +my $launch_ready_name; +$launch_ready_name = shift @command if @command && $command[0] ne "--"; die "missing command\n" unless @command && shift(@command) eq "--"; die "invalid limit\n" unless defined $limit && $limit =~ /\A\d+\z/; +die "invalid launch boundary\n" if defined($launch_ready_name) + && $launch_ready_name !~ /\A\.[A-Za-z0-9._-]{1,384}\.launch-ready\z/; our ($registry_dir, $registry, $reservation_dir, $reservation_root, $sequence); sub fail { die "capture failed: $_[0]\n"; } @@ -180,6 +184,14 @@ sub write_reservation { write_all($runner, "$runner_pid\n"); close($runner) or fail("cannot close runner record"); my $stage = open_new($output_name); +my $launch_ready; +if (defined $launch_ready_name) { + sysopen($launch_ready, $launch_ready_name, O_WRONLY | O_NOFOLLOW) + or fail("cannot open launch boundary"); + my @launch_ready_stat = stat($launch_ready); + fail("unsafe launch boundary") unless @launch_ready_stat && -f _ && $launch_ready_stat[4] == $< + && ($launch_ready_stat[2] & 07777) == 0600 && $launch_ready_stat[3] == 1; +} pipe(my $reader, my $writer) or fail("cannot create output pipe"); my $child = fork(); defined $child or fail("cannot fork adapter"); @@ -191,6 +203,10 @@ sub write_reservation { exit 127; } close($writer); +if (defined $launch_ready) { + write_all($launch_ready, "ready\n"); + close($launch_ready) or fail("cannot close launch boundary"); +} my ($written, $truncated) = (0, 0); while (1) { my $read = sysread($reader, my $buffer, 65536); diff --git a/bin/fm-procevent-lavish.sh b/bin/fm-procevent-lavish.sh index 91b2ac5e3b4..ece166e7d37 100755 --- a/bin/fm-procevent-lavish.sh +++ b/bin/fm-procevent-lavish.sh @@ -7,6 +7,7 @@ # fm-procevent-lavish.sh terminal <result-file> # fm-procevent-lavish.sh silent <result-file> # fm-procevent-lavish.sh answers <result-file> +# fm-procevent-lavish.sh reconciles <result-file> # fm-procevent-lavish.sh read <result-file> # fm-procevent-lavish.sh source-id <artifact.html> # fm-procevent-lavish.sh retire <artifact.html> @@ -96,7 +97,7 @@ # That is an internal retry, not news, so registering the raw poll made the # generic runner capture it and wake the whole fleet. `poll` therefore re-runs # the published poll up to POLL_RETRY_LIMIT times for that exact response, with -# POLL_RETRY_DELAY_DEFAULT seconds between attempts. The match is exact and +# attempt starts at least POLL_RETRY_DELAY_DEFAULT seconds apart. The match is exact and # deliberately narrow: real feedback, ended and missing sessions, any other # SERVER_ERROR, and the same interruption still standing after the bound is # spent are all printed straight through and captured normally. The retry is a @@ -174,6 +175,7 @@ cmd_retire() { # without waiting it out. POLL_RETRY_LIMIT=12 POLL_RETRY_DELAY_DEFAULT=5 +POLL_RETRY_DELAY_MIN=1 POLL_RETRY_DELAY_MAX=60 # Exit 0 only for the exact two-line interruption, and nothing else. The whole @@ -226,8 +228,8 @@ poll_response_filter() { # <response-file> ' "$1" } -# Seconds between retries. FM_LAVISH_POLL_RETRY_DELAY is a bounded test -# override; a malformed or out-of-range value is refused rather than quietly +# Minimum seconds between retry attempt starts. FM_LAVISH_POLL_RETRY_DELAY is a +# bounded test override; a malformed or out-of-range value is refused rather than quietly # rounded, because silently changing a retry cadence is how a bound stops # meaning anything. poll_retry_delay() { @@ -237,15 +239,28 @@ poll_retry_delay() { return 0 fi case "$delay" in - *[!0-9]*) die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" ;; + *[!0-9]*) die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from $POLL_RETRY_DELAY_MIN to $POLL_RETRY_DELAY_MAX: $delay" ;; esac - [ "$delay" -le "$POLL_RETRY_DELAY_MAX" ] \ - || die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from 0 to $POLL_RETRY_DELAY_MAX: $delay" + [ "$delay" -ge "$POLL_RETRY_DELAY_MIN" ] && [ "$delay" -le "$POLL_RETRY_DELAY_MAX" ] \ + || die "FM_LAVISH_POLL_RETRY_DELAY must be whole seconds from $POLL_RETRY_DELAY_MIN to $POLL_RETRY_DELAY_MAX: $delay" printf '%s\n' "$delay" } +poll_iteration_started() { + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e \ + 'printf "%.6f\\n", clock_gettime(CLOCK_MONOTONIC)' +} + +poll_iteration_floor_wait() { + perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' + my ($started, $floor) = @ARGV; + my $remaining = $floor - (clock_gettime(CLOCK_MONOTONIC) - $started); + sleep($remaining) if $remaining > 0; + ' "$1" "$2" +} + cmd_poll() { - local artifact=${1-} delay attempt=0 response cleanup_command rc filter_rc + local artifact=${1-} delay attempt=0 response cleanup_command rc filter_rc iteration_started local pipeline_status [ -n "$artifact" ] || usage [ "$#" -eq 1 ] || usage @@ -265,6 +280,7 @@ cmd_poll() { trap "$cleanup_command; trap - $signal; kill -$signal $$" "$signal" done while :; do + iteration_started=$(poll_iteration_started) || die "cannot start the poll rate governor" lavish-axi poll "$artifact" | poll_response_filter "$response" pipeline_status=("${PIPESTATUS[@]}") rc=${pipeline_status[0]} @@ -274,7 +290,8 @@ cmd_poll() { 10) if [ "$attempt" -lt "$POLL_RETRY_LIMIT" ]; then attempt=$((attempt + 1)) - sleep "$delay" + poll_iteration_floor_wait "$iteration_started" "$delay" \ + || die "cannot enforce the poll rate governor" else cat -- "$response" break @@ -393,7 +410,7 @@ cmd_silent() { [ "$content_rc" -eq 1 ] } -# Print `key<TAB>answer<TAB>label[<TAB>mode]` for every structured choice the +# Print `key<TAB>answer<TAB>label[<TAB>mode]` for each non-reconcile structured choice the # captain submitted in a captured result; the optional mode column relays the # card's declared close mode (`done` or `release`) to the keyed-answer intake. The published response frames queued feedback as # a `prompts[N]{field,...}:` header followed by exactly N indented CSV rows whose @@ -401,18 +418,20 @@ cmd_silent() { # rather than assuming a fixed column, and takes only rows whose `tag` field is # `choice`. A freeform `message` row is captain prose and is deliberately never a # source of decision keys. A row that does not carry both a slug-shaped `question` -# and an `answer` inside its `Context data:` block is skipped, so a deck that does -# not key its forms by decision key simply yields nothing. +# and the versioned `selection` and `note` fields inside its `Context data:` block +# is skipped. A time-limited rollout branch accepts the old question/answer +# shape only for ordinary answers and rejects its bare or annotated reconcile +# values because old rows do not separate the selected option from its note. # The question cap is 128 so any task id fits, including the long legacy # `<origin>-decision-<key>` identities pre-collapse decks still carry; the # security property is the slug SHAPE, which is unchanged. -cmd_answers() { - local file=${1-} +cmd_choice_rows() { + local selection=$1 file=${2-} [ -n "$file" ] || usage [ -f "$file" ] && [ ! -L "$file" ] || die "result file does not exist: $file" perl -MJSON::PP -e ' use strict; use warnings; - my ($path) = @ARGV; + my ($selection, $path) = @ARGV; open my $fh, "<", $path or exit 1; my (@fields, $want, @rows); while (my $line = <$fh>) { @@ -428,7 +447,7 @@ cmd_answers() { } close $fh; my %seen; - my @out; + my @choices; for my $row (@rows) { $row =~ s/^\s+//; my @vals; @@ -451,29 +470,72 @@ cmd_answers() { my $ctx = $1; my $data = eval { decode_json($ctx) }; next unless ref($data) eq "HASH"; - my $key = $data->{question}; - my $answer = $data->{answer}; - next if !defined($key) || ref($key) || !defined($answer) || ref($answer); + my ($key, $selected, $note, $answer, $legacy); + if (defined($data->{schema}) && !ref($data->{schema}) + && $data->{schema} eq "fm-bearings-answer.v1") { + $key = $data->{question}; + $selected = $data->{selection}; + $note = $data->{note}; + next if !defined($key) || ref($key) || !defined($selected) || ref($selected) + || !defined($note) || ref($note); + next unless $selected eq "" || $selected =~ /\A[A-Za-z0-9._-]{1,128}\z/; + next unless length($note) <= 512; + next unless length($selected) || length($note); + $answer = length($selected) ? $selected : $note; + $legacy = 0; + # Time-limited compatibility for captures from pre-change boards; remove + # once no board carrying the old question/answer context can remain armed. + } elsif (!exists($data->{schema}) && !exists($data->{selection}) + && !exists($data->{note})) { + $key = $data->{question}; + $answer = $data->{answer}; + next if !defined($key) || ref($key) || !defined($answer) || ref($answer); + next unless length($answer) && length($answer) <= 512; + next if $answer eq "reconcile" || index($answer, "reconcile - ") == 0; + $selected = ""; + $note = ""; + $legacy = 1; + } else { + next; + } + next unless $key =~ /\A[A-Za-z0-9._-]{1,128}\z/; my $mode = ""; if (exists $data->{close}) { next if !defined($data->{close}) || ref($data->{close}) || ($data->{close} ne "done" && $data->{close} ne "release"); $mode = $data->{close}; } - next unless $key =~ /\A[A-Za-z0-9._-]{1,128}\z/; - next unless length $answer && length($answer) <= 512; my $label = defined $f{text} ? $f{text} : ""; - s/[\x00-\x1f\x7f]/ /g for ($answer, $label); + s/[\x00-\x1f\x7f]/ /g for ($answer, $note, $label); $label = substr($label, 0, 512); - # A re-answered form appears again later in the queue; the last submission wins. - if (defined $seen{$key}) { $out[$seen{$key}] = undef } - $seen{$key} = scalar @out; - push @out, length $mode ? "$key\t$answer\t$label\t$mode" : "$key\t$answer\t$label"; + if (defined $seen{$key}) { $choices[$seen{$key}] = undef } + $seen{$key} = scalar @choices; + push @choices, { + key => $key, selection => $selected, note => $note, legacy => $legacy, + answer => $answer, label => $label, mode => $mode + }; } - print "$_\n" for grep { defined } @out; - ' "$file" + for my $choice (grep { defined } @choices) { + if ($selection eq "reconciles") { + next if $choice->{legacy}; + if ($choice->{selection} eq "reconcile") { + print length($choice->{note}) + ? "$choice->{key}\t$choice->{note}\n" + : "$choice->{key}\n"; + } + next; + } + next if $choice->{selection} eq "reconcile"; + print length $choice->{mode} + ? "$choice->{key}\t$choice->{answer}\t$choice->{label}\t$choice->{mode}\n" + : "$choice->{key}\t$choice->{answer}\t$choice->{label}\n"; + } + ' "$selection" "$file" } +cmd_answers() { cmd_choice_rows answers "$@"; } +cmd_reconciles() { cmd_choice_rows reconciles "$@"; } + # Present one already-captured result for a handler. Body lines are prefixed # so a captain-supplied string cannot forge a section label. The session-ending # message is printed before the count line and before any annotation, because @@ -622,6 +684,7 @@ case "${1-}" in terminal) shift; cmd_terminal "$@" ;; silent) shift; cmd_silent "$@" ;; answers) shift; cmd_answers "$@" ;; + reconciles) shift; cmd_reconciles "$@" ;; read) shift; cmd_read "$@" ;; ''|-h|--help|help) usage ;; *) die "unknown command: $1" ;; diff --git a/bin/fm-procevent-lib.sh b/bin/fm-procevent-lib.sh index 5bfaca8bd74..c1e4a9da13b 100644 --- a/bin/fm-procevent-lib.sh +++ b/bin/fm-procevent-lib.sh @@ -103,6 +103,202 @@ fm_procevent_any_registered() { return 1 } +# --- owning-session lease --------------------------------------------------- +# A runner is detached into its own process group so it survives the turn that +# started it. That is what makes a persistent source work, and on its own it is +# also what lets a runner outlive its whole home: once reparented to init, +# nothing bounds its lifetime, so its blocking child - and everything that child +# spawns - can keep running indefinitely. +# +# The bound is a lease on the OWNING STATE ROOT. Owner-presence operations +# refresh it, an attached public start keeps it fresh while its caller remains +# attached, and the watcher's reconcile cycle keeps it fresh in a live home. +# A guard proves the runner's owner is still there by reading that lease from +# the physical state root recorded in the claim. After two consecutive checks +# cannot prove both the root identity and a fresh lease, it stops the runner's +# process group. The lease is keyed by state root, so another home's live runner +# is untouched: that home refreshes its own lease. Nothing here keys on a script +# name, a command line, or a process name, all of which are shared across homes. + +fm_procevent_owner_lease_path() { # <state-root> + printf '%s/.owner-lease\n' "$(fm_procevent_registry_dir "$1")" +} + +# Record owner-presence activity in this home's process-event state. Best +# effort by design: a home with no registry directory yet owns no runner. +fm_procevent_owner_lease_touch() { # <state-root> + local reg lease tmp now + reg=$(fm_procevent_registry_dir "$1") + [ -d "$reg" ] && [ ! -L "$reg" ] || return 1 + lease=$(fm_procevent_owner_lease_path "$1") + now=$(perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e \ + 'printf "%.6f\n", clock_gettime(CLOCK_MONOTONIC)') || return 1 + tmp=$(umask 077; mktemp "$reg/.owner-lease.XXXXXX") || return 1 + if ! printf '%s\n' "$now" > "$tmp" || ! mv -f -- "$tmp" "$lease"; then + rm -f -- "$tmp" + return 1 + fi +} + +# Seconds since the last refresh. Fails when the lease is absent or unreadable, +# which is what a removed home looks like from inside a surviving runner. +fm_procevent_owner_lease_age() { # <state-root> + local lease value + lease=$(fm_procevent_owner_lease_path "$1") + [ -f "$lease" ] && [ ! -L "$lease" ] || return 1 + IFS= read -r value < "$lease" || return 1 + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -e ' + use strict; + use warnings; + my $value = shift; + $value =~ /\A[0-9]+(?:\.[0-9]+)?\z/ or exit 1; + my $now = clock_gettime(CLOCK_MONOTONIC); + $now >= $value or exit 1; + printf "%d\n", int($now - $value); + ' "$value" +} + +# How long a runner keeps going with no activity in its owning home. The default +# is forty watcher cycles at the default poll interval, so an ordinary busy or +# briefly wedged home never trips it, while a home that is simply gone stops +# owning processes within the hour rather than within a day. +FM_PROCEVENT_OWNER_LEASE_DEFAULT_SECONDS=600 +FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS=1 +FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS=86400 + +fm_procevent_owner_lease_seconds() { + local value=${FM_PROCEVENT_OWNER_LEASE_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_OWNER_LEASE_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +# How often a runner's guard re-reads that lease. One watcher cycle at the +# default poll interval, so the guard costs about as much as the cycle that +# refreshes what it reads. +FM_PROCEVENT_OWNER_CHECK_DEFAULT_SECONDS=15 +FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS=1 +FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS=3600 + +fm_procevent_owner_check_seconds() { + local value=${FM_PROCEVENT_OWNER_CHECK_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_OWNER_CHECK_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +FM_PROCEVENT_LAUNCH_FLOOR_DEFAULT_SECONDS=1 +FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS=1 +FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS=3600 + +fm_procevent_launch_floor_seconds() { + local value=${FM_PROCEVENT_LAUNCH_FLOOR_SECONDS-} + if [ -z "$value" ]; then + printf '%s\n' "$FM_PROCEVENT_LAUNCH_FLOOR_DEFAULT_SECONDS" + return 0 + fi + case "$value" in ''|*[!0-9]*) return 1 ;; esac + [ "$value" -ge "$FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS" ] || return 1 + [ "$value" -le "$FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS" ] || return 1 + printf '%s\n' "$value" +} + +fm_procevent_launch_floor_reset_locked() { # <state-root> <source-id> <registration-identity> + local reg identity + case "$3" in *:*) ;; *) return 1 ;; esac + case "$3" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$1") || return 1 + identity=${3//:/-} + rm -f -- "$reg/$2.$identity.last-launch" +} + +fm_procevent_launch_floor_prune_locked() { # <state-root> <source-id> <registration-identity> + local reg identity keep stamp + case "$3" in *:*) ;; *) return 1 ;; esac + case "$3" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$1") || return 1 + identity=${3//:/-} + keep="$reg/$2.$identity.last-launch" + for stamp in "$reg/$2".*.last-launch "$reg/$2.last-launch"; do + [ "$stamp" = "$keep" ] && continue + [ -e "$stamp" ] || [ -L "$stamp" ] || continue + rm -f -- "$stamp" || return 1 + done +} + +fm_procevent_launch_floor_wait() { # <state-root> <source-id> <registration-identity> <seconds> + local state=$1 id=$2 expected=$3 floor=$4 reg stamp identity registration current_identity status=0 + case "$expected" in *:*) ;; *) return 1 ;; esac + case "$expected" in ''|*[!0-9:]*) return 1 ;; esac + reg=$(fm_procevent_registry_dir "$state") || return 1 + identity=${expected//:/-} + stamp="$reg/$id.$identity.last-launch" + [ ! -L "$stamp" ] || return 1 + [ ! -e "$stamp" ] || [ -f "$stamp" ] || return 1 + perl -MTime::HiRes=clock_gettime,sleep,CLOCK_MONOTONIC -e ' + use strict; + use warnings; + my ($path, $floor) = @ARGV; + my $previous; + if (-e $path) { + open my $in, "<", $path or exit 1; + my $value = <$in>; + close $in or exit 1; + defined($value) && $value =~ /\A([0-9]+(?:\.[0-9]+)?)\n?\z/ or exit 1; + $previous = 0 + $1; + } + my $now = clock_gettime(CLOCK_MONOTONIC); + my $elapsed = defined($previous) && $now >= $previous ? $now - $previous : undef; + sleep($floor - $elapsed) if defined($elapsed) && $elapsed < $floor; + ' "$stamp" "$floor" || return 1 + + # Registration publication holds this same source lock while replacing and + # pruning pacing state, so a superseded sleeper cannot recreate its stamp. + fm_procevent_source_lock_acquire "$id" || return 1 + registration="$reg/$id.source" + current_identity=$(fm_pr_file_identity "$registration" 2>/dev/null) || current_identity= + if [ "$current_identity" != "$expected" ]; then + fm_procevent_source_lock_release "$id" || return 1 + return 2 + fi + [ ! -L "$stamp" ] && { [ ! -e "$stamp" ] || [ -f "$stamp" ]; } || status=1 + if [ "$status" -eq 0 ]; then + perl -MTime::HiRes=clock_gettime,CLOCK_MONOTONIC -MFcntl=:DEFAULT -e ' + use strict; + use warnings; + my $path = shift; + my $now = clock_gettime(CLOCK_MONOTONIC); + my $tmp = "$path.$$"; + sysopen(my $out, $tmp, O_WRONLY | O_CREAT | O_EXCL, 0600) or exit 1; + print {$out} "$now\n" or exit 1; + close $out or exit 1; + rename $tmp, $path or exit 1; + ' "$stamp" || status=1 + fi + if [ "$status" -ne 0 ]; then + fm_procevent_source_lock_release "$id" || : + return "$status" + fi + return 0 +} + +# True while the owning home is provably still active. +fm_procevent_owner_alive() { # <state-root> <lease-seconds> + local age + age=$(fm_procevent_owner_lease_age "$1") || return 1 + [ "$age" -le "$2" ] +} + # --- ownership -------------------------------------------------------------- # A claim is a private file recording the home, runner pid, claim generation, # and process identity. Registration and every ownership transition are @@ -130,7 +326,7 @@ fm_procevent_source_lock_release() { } fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <argv...> - local state=$1 adapter=$2 id=$3 reg dest tmp arg + local state=$1 adapter=$2 id=$3 reg dest tmp arg identity shift 3 fm_procevent_adapter_valid "$adapter" || return 1 fm_procevent_source_id_valid "$id" || return 1 @@ -148,7 +344,11 @@ fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <a printf 'argc=%s\n' "$#" printf 'argv:\n' printf '%s\n' "$@" - } > "$tmp" && fm_pr_private_file_secure "$tmp" 600 && mv -f -- "$tmp" "$dest"; then + } > "$tmp" && fm_pr_private_file_secure "$tmp" 600 \ + && identity=$(fm_pr_file_identity "$tmp") \ + && fm_procevent_launch_floor_reset_locked "$state" "$id" "$identity" \ + && mv -f -- "$tmp" "$dest"; then + fm_procevent_launch_floor_prune_locked "$state" "$id" "$identity" 2>/dev/null || : return 0 fi rm -f -- "$tmp" @@ -160,7 +360,7 @@ fm_procevent_registration_publish_locked() { # <state> <adapter> <source-id> <a # stored because the tracked host constructs that command at run time. fm_procevent_extension_registration_publish_locked() { # <state> <adapter> <source-id> <extension-id> <extension-version> <capability-version> <package-digest> <binding-digest> <config-ref> <registration-token> local state=$1 adapter=$2 id=$3 extension_id=$4 extension_version=$5 capability_version=$6 - local package_digest=$7 binding_digest=$8 config_ref=$9 registration_token=${10} reg dest tmp + local package_digest=$7 binding_digest=$8 config_ref=$9 registration_token=${10} reg dest tmp identity fm_procevent_adapter_valid "$adapter" || return 1 fm_procevent_source_id_valid "$id" || return 1 fm_procevent_extension_id_valid "$extension_id" || return 1 @@ -188,7 +388,11 @@ fm_procevent_extension_registration_publish_locked() { # <state> <adapter> <sou printf 'registration_token=%s\n' "$registration_token" printf 'argc=0\n' printf 'argv:\n' - } > "$tmp" && fm_pr_private_file_secure "$tmp" 600 && mv -f -- "$tmp" "$dest"; then + } > "$tmp" && fm_pr_private_file_secure "$tmp" 600 \ + && identity=$(fm_pr_file_identity "$tmp") \ + && fm_procevent_launch_floor_reset_locked "$state" "$id" "$identity" \ + && mv -f -- "$tmp" "$dest"; then + fm_procevent_launch_floor_prune_locked "$state" "$id" "$identity" 2>/dev/null || : return 0 fi rm -f -- "$tmp" @@ -380,27 +584,49 @@ fm_procevent_claim_capture_reservation_remove_locked() { fm_procevent_capture_reservation_remove_claim "$FM_PROCEVENT_CLAIM_STATE_ROOT" "$FM_PROCEVENT_CLAIM_TOKEN" } +# fm_procevent_claim_generation_gone_locked +# True only when the loaded claim's owner is stale and the process group it led +# independently has no members left. The separate group check also covers a +# reused live pid whose identity differs while the old generation survives. +# A live matched owner (state 0), an unreadable identity (state 2), and a +# crashed leader with a still-live ambiguous group (state 3) all return false. +fm_procevent_claim_generation_gone_locked() { + local state=0 + fm_procevent_pid_state "${FM_PROCEVENT_CLAIM_PID:-}" "${FM_PROCEVENT_CLAIM_IDENTITY:-}" || state=$? + [ "$state" -eq 1 ] \ + && ! fm_procevent_group_alive "${FM_PROCEVENT_CLAIM_PID:-}" +} + +# Capture-reservation cleanup for a claim being reclaimed. +# +# Reservation records are keyed by CLAIM TOKEN, and every replacement claims a +# fresh token, so a dead generation's leftovers can never collide with the +# generation that replaces it. They are hygiene, not an ownership invariant - +# the runner's own successful-capture path already tidies them best-effort. +# The cleanup is still attempted and remains authoritative for a generation +# that is not provably gone; it stops being a veto only after the stale owner +# and independent group check prove the whole generation gone. +fm_procevent_claim_capture_reservation_reclaim_locked() { + fm_procevent_claim_capture_reservation_remove_locked && return 0 + fm_procevent_claim_generation_gone_locked +} + # fm_procevent_group_alive <pid> -# True while any process remains in the process group a runner leads. A runner -# started by reconcile is its own group leader, so this is what distinguishes a -# generation that is really gone from one whose leader died while its blocking -# source child kept running. +# True while any process remains in the runner's numeric process group. A runner +# starts as its own group leader, but after that leader exits a same-numbered +# group may be reused, so group presence prevents proving the generation gone. fm_procevent_group_alive() { case "$1" in ''|*[!0-9]*) return 1 ;; esac kill -0 -"$1" 2>/dev/null } # fm_procevent_pid_state <pid> <identity> -# 0 live match, 1 stale, 2 uncertain, 3 orphaned group. +# 0 live match, 1 stale, 2 uncertain, 3 ambiguous leaderless group. # -# State 3 is the crash cut: the runner leader is gone, but its owned process -# group still has members, so the old generation can still be consuming the -# source. Treating that as stale would release ownership and let a second -# poller start against one canonical source. Only the leader being absent -# reaches state 3, which is also what makes signalling that group safe: if this -# pid had been reused by an unrelated process the leader would be alive, so the -# identity comparison below would classify it stale or uncertain and no group -# signal would ever follow. +# State 3 is the crash cut: the runner leader is gone, but a process group with +# its numeric id still has members. That group may be the old generation or a +# leaderless group created after PID/PGID reuse, so cleanup preserves the claim +# without signalling the group or starting a replacement. fm_procevent_pid_state() { local pid=$1 expected=$2 actual if ! fm_pid_alive "$pid"; then @@ -469,7 +695,7 @@ fm_procevent_claim_acquire_locked() { fi fi if [ "$status" -eq 0 ]; then - fm_procevent_claim_capture_reservation_remove_locked || status=1 + fm_procevent_claim_capture_reservation_reclaim_locked || status=1 fi [ "$status" -ne 0 ] || rm -f -- "$claim" || status=1 else @@ -499,6 +725,11 @@ fm_procevent_claim_acquire_locked() { if [ "$status" -eq 0 ]; then FM_PROCEVENT_CLAIM_TOKEN=$token FM_PROCEVENT_CLAIM_REG_IDENTITY=$reg_identity + FM_PROCEVENT_CLAIM_STATE_ROOT=$state_root + FM_PROCEVENT_CLAIM_STATE_DEVICE=$state_device + FM_PROCEVENT_CLAIM_STATE_INODE=$state_inode + FM_PROCEVENT_CLAIM_STATE_OWNER=$state_owner + FM_PROCEVENT_CLAIM_STATE_MODE=$state_mode fi fi [ "$status" -eq 0 ] || { [ -z "${tmp:-}" ] || rm -f -- "$tmp"; } @@ -544,8 +775,28 @@ fm_procevent_claim_mark_terminal_locked() { } # fm_procevent_claim_release_locked <source-id> <home> <pid> <token> +# The live owner uses this path for its own release. Reservation cleanup must +# succeed normally; stale-generation relaxation is never consulted. fm_procevent_claim_release_locked() { - local id=$1 home=$2 pid=$3 token=$4 claim + fm_procevent_claim_release_mode_locked release "$@" +} + +# fm_procevent_claim_release_terminal_self_locked <source-id> <home> <pid> <token> +# A live runner uses this only while retiring its own terminal source mid-capture. +# Its in-flight reservation is transient, so attempt cleanup without making that +# cleanup a veto; exact ownership still must match before releasing the claim. +fm_procevent_claim_release_terminal_self_locked() { + fm_procevent_claim_release_mode_locked terminal-self "$@" +} + +# fm_procevent_claim_reclaim_locked <source-id> <home> <pid> <token> +# Lifecycle commands use this only after proving or stopping a dead generation. +fm_procevent_claim_reclaim_locked() { + fm_procevent_claim_release_mode_locked reclaim "$@" +} + +fm_procevent_claim_release_mode_locked() { + local mode=$1 id=$2 home=$3 pid=$4 token=$5 claim fm_procevent_source_id_valid "$id" || return 1 claim=$(fm_procevent_claim_path "$id") [ -e "$claim" ] || return 0 @@ -553,7 +804,18 @@ fm_procevent_claim_release_locked() { && [ "$FM_PROCEVENT_CLAIM_HOME" = "$home" ] \ && [ "$FM_PROCEVENT_CLAIM_PID" = "$pid" ] \ && [ "$FM_PROCEVENT_CLAIM_TOKEN" = "$token" ]; then - fm_procevent_claim_capture_reservation_remove_locked || return 1 + case "$mode" in + reclaim) + fm_procevent_claim_capture_reservation_reclaim_locked || return 1 + ;; + terminal-self) + fm_procevent_claim_capture_reservation_remove_locked || true + ;; + release) + fm_procevent_claim_capture_reservation_remove_locked || return 1 + ;; + *) return 1 ;; + esac rm -f -- "$claim" return $? fi @@ -584,7 +846,7 @@ fm_procevent_path_normalize() { fm_procevent_directory_owned_by_current_user() { local owner if [ "$(uname)" = Darwin ]; then - owner=$(stat -f %u "$1" 2>/dev/null) + owner=$(/usr/bin/stat -f %u "$1" 2>/dev/null) else owner=$(stat -c %u "$1" 2>/dev/null) fi diff --git a/bin/fm-procevent.sh b/bin/fm-procevent.sh index 64e2f6d2065..c8ed7eb2d11 100755 --- a/bin/fm-procevent.sh +++ b/bin/fm-procevent.sh @@ -137,19 +137,44 @@ # captain chose; the intake owns every rule about what happens next. This runner # names no adapter, parses no result, and knows no decision rule, so a future # built-in source needs nothing here beyond an `answers` command and a binding. -# External binding responses never enter this authority-bearing intake. +# Reconcile selections use the parallel `reconciles` adapter command and the +# binding-verified `reconcile-requests` intake, never the keyed-answer value. +# External binding responses never enter either authority-bearing intake. # # Feeding is deliberately independent of handling: it never acknowledges a result # and never suppresses a wake. Recording the captain's answer is transcription, # while ACTING on it is firstmate's judgement, so the capture stays unacknowledged # and its `check` wake reaches the handler exactly as it would have anyway. # +# A runner is bound to the HOME that owns it, not to the one session that armed +# it: a persistent source is meant to outlive that session, so reconcile stops a +# runner whose source is retired in a live home, and this lease is the backstop +# for a home that is GONE. Detaching a runner into its own +# process group is what lets a persistent source outlive the turn that armed it, +# and with nothing else it is also what lets a runner outlive its whole home: +# reparented to init, it keeps its blocking child - and everything that child +# spawns - running with nobody left to reap it. So every runner starts a small +# guard beside it, in its own separate process group, which re-reads the owning +# state root's lease on a bounded cadence and stops the runner's whole process +# group once that lease can no longer be proved fresh. Owner-presence operations +# refresh the lease, an attached public start keeps it fresh while its caller +# remains attached, and the watcher's reconcile cycle keeps it fresh in a live +# home. A runner exports the inherited FM_PROCEVENT_IN_RUNNER marker and every +# refresh is skipped under it, so a runner and its ordinary children do not +# certify their own owner. That rule is CONFUSED-AGENT-GRADE, the grade +# bin/fm-lease-lib.sh documents: a source that DELIBERATELY strips the marker +# can still refresh, and adversarial-grade unforgeability is out of scope (see +# docs/configuration.md). Scope is the owning state root and one runner +# generation, never a script or process name, so a live source in +# another home is untouched. See bin/fm-procevent-lib.sh for the lease itself. +# # Ownership is machine-wide per canonical source, because separate Firstmate # homes can share one underlying source store. A live owner is never displaced; -# only a claim whose whole generation is gone is reclaimed. A runner leads its -# own process group, so a crashed leader whose group still has members is not -# stale: reconcile stops that surviving group and releases its generation before -# any replacement starts, and keeps the claim for a later retry when it cannot. +# only a claim whose stale owner and independently absent process group prove +# its whole generation gone is reclaimed. A crashed leader or reused pid whose +# process group still has members cannot relax ownership cleanup. Reconcile +# signals only a live identity-matched runner group and otherwise keeps the +# claim without starting a replacement. # # Durability boundary: see bin/fm-procevent-lib.sh. This runner proves capture # before publication and bounded re-announcement until handled, and nothing @@ -360,6 +385,19 @@ feed_keyed_answers() { # <adapter> <source-id> <result-file> --source "the captured result $id sequence $seq" >/dev/null 2>&1 } +feed_reconcile_requests() { # <adapter> <source-id> <result-file> + local adapter=$1 id=$2 result=$3 script origin seq rows + script=$(adapter_script "$adapter") + [ -f "$script" ] && [ ! -L "$script" ] || return 1 + origin=$("$SCRIPT_DIR/fm-captain-hold.sh" binding "$id" 2>/dev/null) || return 1 + [ -n "$origin" ] || return 1 + seq=$(fm_procevent_result_sequence "$result") || return 1 + rows=$("$script" reconciles "$result" 2>/dev/null) || return 1 + printf '%s\n' "$rows" \ + | "$SCRIPT_DIR/fm-captain-hold.sh" reconcile-requests \ + --source-id "$id" --source "the captured result $id sequence $seq" >/dev/null 2>&1 +} + read_adapter() { # <source-id> local f; f=$(source_file "$1") [ -f "$f" ] && [ ! -L "$f" ] || return 1 @@ -421,6 +459,7 @@ cmd_register() { die "cannot publish the registration" fi fm_procevent_source_lock_release "$id" + owner_lease_refresh printf 'registered: %s (%s)\n' "$id" "$adapter" } @@ -510,6 +549,7 @@ cmd_register_extension() { fi fm_procevent_source_lock_release "$id" extension_lifecycle_lock_release + owner_lease_refresh printf 'registered: %s (%s from %s@%s)\n' "$id" "$adapter" "$extension_id" "$extension_version" printf 'owner-token: %s\n' "$registration_token" printf 'retire: bin/fm-procevent.sh retire %s --if-owner %s\n' "$id" "$registration_token" @@ -569,8 +609,15 @@ publish_pending() { # [result-file-to-skip] printf '%s\n' "$published" } -isolate_runner() { # <wait|detach> <source-id> - local mode=$1 id=$2 program +# Start one command as the leader of a fresh process group, either waiting for +# it (the public `start` boundary) or detaching from it (reconcile's restart and +# the runner's own owner guard). The guard deliberately gets its OWN group +# rather than joining the runner's: it has to survive the group signal it sends, +# and a member of the runner's group would also make that group read as alive +# after the runner itself is gone. +isolate_process() { # <wait|detach> <command> [argv...] + local mode=$1 program + shift # shellcheck disable=SC2016 # Perl owns every $ expression in this literal program. program='my $mode = shift @ARGV; defined(my $pid = fork) or exit 125; @@ -586,9 +633,14 @@ isolate_runner() { # <wait|detach> <source-id> exit(128 + ($status & 127)) if $status & 127; exit($status >> 8);' if [ "$mode" = wait ]; then - exec perl -e "$program" "$mode" "$SCRIPT_DIR/fm-procevent.sh" _start "$id" + perl -e "$program" "$mode" "$@" + return $? fi - perl -e "$program" "$mode" "$SCRIPT_DIR/fm-procevent.sh" _start "$id" >/dev/null 2>&1 & + perl -e "$program" "$mode" "$@" >/dev/null 2>&1 & +} + +isolate_runner() { # <wait|detach> <source-id> + isolate_process "$1" "$SCRIPT_DIR/fm-procevent.sh" _start "$2" } fm_procevent_process_group_id() { # <pid> @@ -600,22 +652,56 @@ fm_procevent_process_group_id() { # <pid> ' "$1" 2>/dev/null } -require_runner_group() { - local pgid +require_isolated_group() { # <role> + local role=$1 pgid [ "${FM_PROCEVENT_RUNNER_GROUP:-}" = "$$" ] \ - || die "runner process group was not isolated" + || die "$role process group was not isolated" pgid=$(fm_procevent_process_group_id "$$") \ - || die "cannot inspect runner process group" - [ -n "$pgid" ] || die "cannot inspect runner process group" - [ "$pgid" = "$$" ] || die "runner does not lead its process group" + || die "cannot inspect $role process group" + [ -n "$pgid" ] || die "cannot inspect $role process group" + [ "$pgid" = "$$" ] || die "$role does not lead its process group" unset FM_PROCEVENT_RUNNER_GROUP } +require_runner_group() { require_isolated_group runner; } + +# Record owner-presence activity for this home. Skipped under the inherited +# FM_PROCEVENT_IN_RUNNER marker, so a runner and its ordinary children do not +# keep refreshing their own lease after the home goes away. +# Confused-agent-grade: a source that deliberately unsets the marker can still +# refresh, and that is out of scope (see docs/configuration.md). +owner_lease_refresh() { + [ "${FM_PROCEVENT_IN_RUNNER:-0}" = 1 ] && return 0 + fm_procevent_owner_lease_touch "$STATE" 2>/dev/null || true +} + +owner_lease_keepalive() { # <parent-pid> <parent-identity> + local parent=$1 identity=$2 state + while :; do + sleep 1 + fm_procevent_pid_state "$parent" "$identity" + state=$? + case "$state" in + 0) owner_lease_refresh ;; + 2) ;; + *) return 0 ;; + esac + done +} + cmd_start_public() { - local id=${1-} + local id=${1-} identity keeper status [ "$#" -eq 1 ] || usage fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" + owner_lease_refresh + identity=$(fm_pid_identity "$$" 2>/dev/null) || die "cannot identify the attached owner" + owner_lease_keepalive "$$" "$identity" & + keeper=$! isolate_runner wait "$id" + status=$? + kill "$keeper" 2>/dev/null || true + wait "$keeper" 2>/dev/null || true + return "$status" } cmd_start() { @@ -674,6 +760,10 @@ cmd_start() { die "extension registration owner is unreadable: $id" ;; esac + exec 7<"$(source_file "$id")" || { + fm_procevent_source_lock_release "$id" + die "cannot retain registration identity: $id" + } fm_procevent_claim_acquire_locked "$id" "$FM_HOME" "$$" "$(source_file "$id")" "$STATE" claimed=$? fm_procevent_source_lock_release "$id" @@ -687,6 +777,8 @@ cmd_start() { CLAIM_PID=$$ CLAIM_TOKEN=$FM_PROCEVENT_CLAIM_TOKEN CLAIM_REG_IDENTITY=$FM_PROCEVENT_CLAIM_REG_IDENTITY + CLAIM_STATE_DEVICE=$FM_PROCEVENT_CLAIM_STATE_DEVICE + CLAIM_STATE_INODE=$FM_PROCEVENT_CLAIM_STATE_INODE STAGED_OUTPUT= release_start_claim() { extension_lifecycle_lock_release 2>/dev/null || true @@ -704,7 +796,14 @@ cmd_start() { fm_procevent_source_lock_release "$CLAIM_ID" 2>/dev/null || true } trap release_start_claim EXIT - local runner inbox reservation_dir staging + # The inherited marker keeps the runner and its ordinary children from + # accidentally refreshing the owner lease. A source that deliberately strips + # it is outside this confused-agent-grade boundary. + export FM_PROCEVENT_IN_RUNNER=1 + start_owner_guard "$id" || die "cannot start the runner's owner guard: $id" + local launch_floor runner inbox reservation_dir staging launch_ready launch_reply launch_pid + launch_floor=$(fm_procevent_launch_floor_seconds) \ + || die "FM_PROCEVENT_LAUNCH_FLOOR_SECONDS must be whole seconds from $FM_PROCEVENT_LAUNCH_FLOOR_MIN_SECONDS to $FM_PROCEVENT_LAUNCH_FLOOR_MAX_SECONDS" if [ "$extension_owner" -eq 1 ]; then staging=$(fm_procevent_extension_staging_prepare "$STATE") \ || die "cannot safely prepare the external registry staging boundary" @@ -741,13 +840,45 @@ cmd_start() { # Built-in adapters do not run the extension capture helper, so keep this # sentinel defined while sharing the no-result branch below under `set -u`. local truncated=0 capture_state='' durable='' reservation_terminal='' reservation_silent='' + fm_procevent_launch_floor_wait "$STATE" "$id" "$CLAIM_REG_IDENTITY" "$launch_floor" + case "$?" in + 0) ;; + # A superseded generation leaves nothing behind. The runner marker is + # written before this wait, and a home sweep counts a marker with no owned + # claim as a preflight failure, so exiting without removing it would make + # that home refuse to sweep. + 2) [ "$extension_owner" -eq 1 ] || rm -f -- "$runner"; exit 0 ;; + *) die "cannot enforce the source launch floor: $id" ;; + esac + exec 7<&- if [ "$extension_owner" -eq 1 ]; then - capture_state=$(perl "$SCRIPT_DIR/fm-procevent-extension-capture.pl" \ + launch_ready=".$id.$CLAIM_TOKEN.launch-ready" + launch_reply="$REG/.$id.$CLAIM_TOKEN.launch-reply" + (umask 077; : > "$REG/$launch_ready" && : > "$launch_reply") || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + fm_procevent_source_lock_release "$id" + die "cannot prepare the source launch boundary: $id" + } + perl "$SCRIPT_DIR/fm-procevent-extension-capture.pl" \ 9 8 6 "$id" "$adapter" "$FM_PROCEVENT_EXTENSION_ID" \ "$FM_PROCEVENT_EXTENSION_VERSION" "$FM_PROCEVENT_EXTENSION_CAPABILITY_VERSION" \ "$FM_PROCEVENT_EXTENSION_PACKAGE_DIGEST" "$FM_PROCEVENT_EXTENSION_BINDING_DIGEST" \ - "$CLAIM_TOKEN" "$runner" "$out" "$$" "$(fm_pid_identity "$$")" "$MAX_OUTPUT_BYTES" -- "${ARGV[@]}") \ - || die "cannot safely stage the extension result" + "$CLAIM_TOKEN" "$runner" "$out" "$$" "$(fm_pid_identity "$$")" "$MAX_OUTPUT_BYTES" \ + "$launch_ready" -- "${ARGV[@]}" > "$launch_reply" & + launch_pid=$! + while [ ! -s "$REG/$launch_ready" ] && kill -0 "$launch_pid" 2>/dev/null; do sleep 0.01; done + fm_procevent_source_lock_release "$id" \ + || die "cannot release the source launch boundary: $id" + wait "$launch_pid" || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + die "cannot safely stage the extension result" + } + [ -s "$REG/$launch_ready" ] || { + rm -f -- "$REG/$launch_ready" "$launch_reply" + die "cannot establish the source launch boundary: $id" + } + IFS= read -r capture_state < "$launch_reply" || capture_state= + rm -f -- "$REG/$launch_ready" "$launch_reply" IFS=$'\t' read -r capture_state durable rc truncated reservation_terminal reservation_silent <<EOF $capture_state EOF @@ -762,10 +893,38 @@ EOF FM_PROCEVENT_CAPTURE_RESERVATION_SILENT=$reservation_silent fi else - [ ! -e "$out" ] && [ ! -L "$out" ] || die "cannot safely stage output" - (umask 077; : > "$out") || die "cannot stage output" + [ ! -e "$out" ] && [ ! -L "$out" ] || { + fm_procevent_source_lock_release "$id" + die "cannot safely stage output" + } + (umask 077; : > "$out") || { + fm_procevent_source_lock_release "$id" + die "cannot stage output" + } STAGED_OUTPUT=$out - "${ARGV[@]}" 2>/dev/null | perl -e ' + launch_ready="$REG/.$id.$CLAIM_TOKEN.launch-pipe" + mkfifo -m 600 "$launch_ready" || { + fm_procevent_source_lock_release "$id" + die "cannot prepare the source launch boundary: $id" + } + exec 5<> "$launch_ready" || { + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" + die "cannot retain the source launch boundary: $id" + } + exec 4< "$launch_ready" || { + exec 5>&- + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" + die "cannot retain the source output boundary: $id" + } + "${ARGV[@]}" >&5 5>&- 4<&- 2>/dev/null & + launch_pid=$! + exec 5>&- + rm -f -- "$launch_ready" + fm_procevent_source_lock_release "$id" \ + || die "cannot release the source launch boundary: $id" + perl -e ' use strict; use warnings; my $limit = shift; @@ -788,10 +947,11 @@ EOF $truncated = 1 if $take < $count; } exit($truncated ? 3 : 0); - ' "$MAX_OUTPUT_BYTES" > "$out" - local pipe_status=("${PIPESTATUS[@]}") - rc=${pipe_status[0]} - bound_rc=${pipe_status[1]} + ' "$MAX_OUTPUT_BYTES" <&4 > "$out" + bound_rc=$? + exec 4<&- + wait "$launch_pid" + rc=$? case "$bound_rc" in 0) ;; 3) truncated=1 ;; @@ -825,6 +985,10 @@ EOF # Independent of publication and acknowledgement, so it runs once per capture # for every adapter and cannot change what the handler receives. + if [ "$extension_owner" -eq 0 ] \ + && feed_reconcile_requests "$adapter" "$id" "$durable"; then + printf 'reconciles-fed: %s\n' "$id" + fi if [ "$extension_owner" -eq 0 ] \ && feed_keyed_answers "$adapter" "$id" "$durable"; then printf 'answers-fed: %s\n' "$id" @@ -901,7 +1065,7 @@ retire_owned_terminal_source() { # <source-id> && [ "$current_identity" = "$CLAIM_REG_IDENTITY" ] \ && fm_procevent_claim_mark_terminal_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN"; then if rm -f -- "$registration" && [ ! -e "$registration" ] && [ ! -L "$registration" ]; then - fm_procevent_claim_release_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN" || status=1 + fm_procevent_claim_release_terminal_self_locked "$id" "$CLAIM_HOME" "$CLAIM_PID" "$CLAIM_TOKEN" || status=1 else status=1 fi @@ -912,6 +1076,106 @@ retire_owned_terminal_source() { # <source-id> return "$status" } +# Bind this runner's lifetime to the home that owns it. Started once the +# claim is held, so the guard names the exact generation it protects, and +# detached into its OWN process group so the group signal it may later send +# reaches the runner and every descendant without killing the guard first. +# If signalling cannot be proved safe or does not finish, the guard remains +# alive and retries on its normal check cadence rather than abandoning cleanup. +start_owner_guard() { # <source-id> + local identity ready value + identity=$(fm_pid_identity "$$" 2>/dev/null) || return 1 + ready=$(umask 077; mktemp "$REG/.owner-guard-ready.XXXXXX") || return 1 + if ! isolate_process detach "$SCRIPT_DIR/fm-procevent.sh" _owner-watchdog \ + "$1" "$$" "$identity" "$ready" "$CLAIM_STATE_DEVICE" "$CLAIM_STATE_INODE"; then + rm -f -- "$ready" + return 1 + fi + for _ in $(seq 1 50); do + if [ -s "$ready" ]; then + IFS= read -r value < "$ready" || value= + rm -f -- "$ready" + [ "$value" = ready ] + return $? + fi + sleep 0.1 + done + rm -f -- "$ready" + return 1 +} + +# The runner's owner guard, which bounds an accidentally orphaned detached +# runner after its home ends. It revalidates the recorded physical state root +# and its lease on a bounded cadence and, after two consecutive checks cannot prove +# both, invokes the identity-gated stop for the runner's whole process group - +# which is what reaches the blocking child and everything that child spawned, +# exactly as retirement does. A failed verified stop stays on the retry cadence; +# an absent leader ends the guard without signalling an ambiguous group. +# +# Scope is the owning state root and this one runner generation. It never +# matches on a script name, a command line, or a process name: those are shared +# by every home running the same adapter, and a live source in another home +# proves its own owner through that home's own lease. +cmd_owner_watchdog() { # <source-id> <runner-pid> <runner-identity> <ready-file> <state-device> <state-inode> + local id=${1-} pid=${2-} identity=${3-} ready=${4-} state_device=${5-} state_inode=${6-} + local lease tick misses=0 pid_state state_identity current_device current_inode + [ "$#" -eq 6 ] || usage + fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" + case "$pid" in ''|*[!0-9]*) die "runner pid must be a positive integer: $pid" ;; esac + [ -n "$identity" ] || die "runner identity is required" + case "$state_device" in ''|*[!0-9]*) die "state device must be an integer" ;; esac + case "$state_inode" in ''|*[!0-9]*) die "state inode must be an integer" ;; esac + [ "${ready%/*}" = "$REG" ] && [ -f "$ready" ] && [ ! -L "$ready" ] \ + || die "owner guard readiness boundary is invalid" + trap 'printf "failed\n" > "$ready" 2>/dev/null || true' EXIT + require_isolated_group guard + lease=$(fm_procevent_owner_lease_seconds) \ + || die "FM_PROCEVENT_OWNER_LEASE_SECONDS must be whole seconds from $FM_PROCEVENT_OWNER_LEASE_MIN_SECONDS to $FM_PROCEVENT_OWNER_LEASE_MAX_SECONDS" + tick=$(fm_procevent_owner_check_seconds) \ + || die "FM_PROCEVENT_OWNER_CHECK_SECONDS must be whole seconds from $FM_PROCEVENT_OWNER_CHECK_MIN_SECONDS to $FM_PROCEVENT_OWNER_CHECK_MAX_SECONDS" + fm_procevent_pid_state "$pid" "$identity" + pid_state=$? + [ "$pid_state" -eq 0 ] || die "runner identity changed before owner guard initialization" + state_identity=$(fm_procevent_claim_state_root_identity "$STATE") \ + || die "owning state root identity is unreadable at owner guard initialization" + IFS=$'\t' read -r _ current_device current_inode _ _ <<< "$state_identity" + [ "$current_device" = "$state_device" ] && [ "$current_inode" = "$state_inode" ] \ + || die "owning state root identity changed before owner guard initialization" + fm_procevent_owner_alive "$STATE" "$lease" \ + || die "owning home lease is not fresh at owner guard initialization" + printf 'ready\n' > "$ready" || die "cannot confirm owner guard initialization" + trap - EXIT + while :; do + sleep "$tick" + fm_procevent_pid_state "$pid" "$identity" + pid_state=$? + case "$pid_state" in + 1|3) exit 0 ;; + 0) ;; + *) continue ;; + esac + state_identity=$(fm_procevent_claim_state_root_identity "$STATE" 2>/dev/null || true) + current_device= + current_inode= + [ -z "$state_identity" ] \ + || IFS=$'\t' read -r _ current_device current_inode _ _ <<< "$state_identity" + if [ "$current_device" = "$state_device" ] \ + && [ "$current_inode" = "$state_inode" ] \ + && fm_procevent_owner_alive "$STATE" "$lease"; then + misses=0 + continue + fi + # Two consecutive misses, so one unreadable read cannot end a live runner. + misses=$((misses + 1)) + [ "$misses" -ge 2 ] || continue + if stop_runner_pid "$pid" "$identity"; then + exit 0 + fi + # Identity/group inspection and signalling can fail transiently. Keep the + # guard alive so the next normal tick retries the same generation cleanup. + done +} + # Start a runner outside the watcher cycle that noticed it was missing. The # public start boundary establishes its own process group before claiming. detach_runner() { # <source-id> @@ -920,6 +1184,7 @@ detach_runner() { # <source-id> cmd_reconcile() { local rec id published started=0 stopped=0 uncertain=0 claim owner pid token identity claim_state stop_state + owner_lease_refresh published=$(publish_pending) # Stop a runner this home owns whose source is no longer registered. Without @@ -951,7 +1216,7 @@ cmd_reconcile() { stop_state=$? case "$stop_state" in 0|1) - if fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then + if fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then rm -f -- "$(staging_file "$id" "$token")" rm -f -- "$(runner_file "$id")" stopped=$((stopped + 1)) @@ -991,36 +1256,14 @@ cmd_reconcile() { && rm -f -- "$(source_file "$id")" \ && [ ! -e "$(source_file "$id")" ] \ && [ ! -L "$(source_file "$id")" ] \ - && fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then + && fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then stopped=$((stopped + 1)) else uncertain=$((uncertain + 1)) fi elif [ "$claim_state" -eq 3 ]; then - # The leader crashed but its owned group is still consuming the - # source. Never start a replacement alongside it: stop that group and - # release its generation first, and if either cannot be proved, keep - # the claim and retry on a later cycle rather than adding a second - # poller. Only the owning home may signal its own group. - owner=$FM_PROCEVENT_CLAIM_HOME - pid=$FM_PROCEVENT_CLAIM_PID - token=$FM_PROCEVENT_CLAIM_TOKEN - identity=$FM_PROCEVENT_CLAIM_IDENTITY - stop_state=2 - if fm_procevent_claim_owned_by_state "$STATE" "$FM_HOME"; then - stop_runner_pid "$pid" "$identity" - stop_state=$? - fi - if [ "$stop_state" -eq 0 ] \ - && cleanup_extension_registration_invocations_locked "$id" \ - && fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token" 2>/dev/null; then - rm -f -- "$(staging_file "$id" "$token")" - rm -f -- "$(runner_file "$id")" - fm_procevent_source_lock_release "$id" - detach_runner "$id" - started=$((started + 1)) - continue - fi + # A leaderless group's generation is ambiguous under PID/PGID reuse, + # so preserve its claim without signalling or starting a replacement. uncertain=$((uncertain + 1)) elif [ "$claim_state" -eq 2 ]; then uncertain=$((uncertain + 1)) @@ -1036,38 +1279,40 @@ cmd_reconcile() { # its own process group leader, so the group signal is what actually reaches the # blocking child - signalling only the runner would leave that child alive and # reparented, which is exactly how a source that never completes leaks. -stop_runner_pid() { # <pid> <identity> - local pid=${1-} identity=${2-} state pgid i=0 - case "$pid" in ''|*[!0-9]*) return 2 ;; esac - [ -n "$identity" ] || return 2 +runner_group_signal() { # <signal> <pid> <identity> + local signal=$1 pid=$2 identity=$3 state pgid + # KNOWN LIMIT: only an alive identity-matched leader proves group ownership. + # Detected reused PIDs and absent leaders are refused before signalling; + # launch pacing, leases, and reconcile cleanup are the backstop. fm_procevent_pid_state "$pid" "$identity" state=$? case "$state" in - 0) - # A live identity-matched leader still owns its group, so prove the group - # really is the one this pid leads before signalling it. - pgid=$(fm_procevent_process_group_id "$pid") || return 2 - [ "$pgid" = "$pid" ] || return 2 - ;; - 3) - # The leader crashed but its owned group is still running. Its pgid cannot - # be read from the dead leader, and it does not need to be: only an absent - # leader reaches this state, so the group cannot belong to a reused pid. - ;; - *) return "$state" ;; + 0) ;; + 1) fm_procevent_group_alive "$pid" && return 2; return 1 ;; + *) return 2 ;; esac - kill -TERM -"$pid" 2>/dev/null || return 2 + pgid=$(fm_procevent_process_group_id "$pid") || return 2 + [ "$pgid" = "$pid" ] || return 2 + # KNOWN LIMIT: portable shell cannot make this verification and signal atomic, + # so the PID and group could be reused in the interval between them. + kill -"$signal" -"$pid" 2>/dev/null || return 2 +} + +stop_runner_pid() { # <pid> <identity> + local pid=${1-} identity=${2-} signal_state i=0 + case "$pid" in ''|*[!0-9]*) return 2 ;; esac + [ -n "$identity" ] || return 2 + runner_group_signal TERM "$pid" "$identity" + signal_state=$? + [ "$signal_state" -eq 0 ] || return "$signal_state" while [ "$i" -lt 20 ]; do kill -0 -"$pid" 2>/dev/null || return 0 - if kill -0 "$pid" 2>/dev/null; then - fm_procevent_pid_state "$pid" "$identity" - state=$? - [ "$state" -eq 2 ] && return 2 - fi sleep 0.1 i=$((i + 1)) done - kill -KILL -"$pid" 2>/dev/null || return 2 + runner_group_signal KILL "$pid" "$identity" + signal_state=$? + [ "$signal_state" -eq 0 ] || return "$signal_state" i=0 while [ "$i" -lt 20 ]; do kill -0 -"$pid" 2>/dev/null || return 0 @@ -1104,6 +1349,7 @@ cmd_handled() { local id=${1-} seq=${2-} status fm_procevent_source_id_valid "$id" || die "source id must be path-safe: $id" case "$seq" in ''|*[!0-9]*) die "sequence must be a nonnegative integer: $seq" ;; esac + owner_lease_refresh fm_procevent_source_lock_acquire "$id" || die "cannot lock source: $id" fm_procevent_mark_handled "$STATE" "$id" "$seq" status=$? @@ -1202,7 +1448,7 @@ cmd_retire() { fm_procevent_source_lock_release "$id" die "cannot prove external adapter cleanup; source remains registered: $id" fi - if ! fm_procevent_claim_release_locked "$id" "$owner" "$pid" "$token"; then + if ! fm_procevent_claim_reclaim_locked "$id" "$owner" "$pid" "$token"; then fm_procevent_source_lock_release "$id" die "cannot release source ownership: $id" fi @@ -1380,6 +1626,7 @@ cmd_sweep_home() { cmd_list() { local rec id adapter owner pending + owner_lease_refresh if ! fm_procevent_any_registered "$STATE"; then printf 'no sources registered\n' return 0 @@ -1500,6 +1747,7 @@ case "${1-}" in register-extension) shift; cmd_register_extension "$@" ;; start) shift; cmd_start_public "$@" ;; _start) shift; cmd_start "$@" ;; + _owner-watchdog) shift; cmd_owner_watchdog "$@" ;; reconcile) shift; cmd_reconcile "$@" ;; classify) shift; cmd_classify "$@" ;; handled) shift; cmd_handled "$@" ;; diff --git a/bin/fm-quota-choose.sh b/bin/fm-quota-choose.sh index 43ff8c7c4b9..3c7fa891c56 100755 --- a/bin/fm-quota-choose.sh +++ b/bin/fm-quota-choose.sh @@ -34,6 +34,16 @@ # provider - is owned by AGENTS.md section 4 and the quota-array-dispatch skill, # not by this helper. Use this helper only when the brief already fixed the # candidate order and every candidate's provider is the harness's primary family. +# +# omp (Oh My Pi) has no single primary family, so its candidate model prefix +# selects the family: openai-codex/<id> checks the codex row and +# claude-bridge/<id> checks the claude row, each against the bare <id> for +# model: and product: scopes. Any other or absent prefix is refused up front, +# the same shape as an unknown harness, because no quota-axi row measures it. +# quota-axi reports Codex quota unavailable on this host because omp carries +# its own Codex login, so an openai-codex candidate reads as unknown quota here +# and is never selected on this host; its runway is disclosed uncertainty for +# the agent-side gates, not measured headroom. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -298,13 +308,22 @@ fi printf '%s\n' "$QUOTA_JSON" | fm_quota_json_valid || die "invalid quota-axi provider data" -# provider_for_harness <harness> +# provider_for_harness <harness> [<model>] # Map a firstmate harness name to its primary quota-axi provider family. # Multi-provider harnesses (Pi, OpenCode) map to their primary family only; see -# the header limitation note. Authoritative multi-provider routing is owned by -# AGENTS.md section 4 and the quota-array-dispatch skill, not this helper. +# the header limitation note. omp is keyed on the candidate model prefix instead +# and has no family for any other prefix (see the header). Authoritative +# multi-provider routing is owned by AGENTS.md section 4 and the +# quota-array-dispatch skill, not this helper. provider_for_harness() { case "$1" in + omp) + case "${2:-}" in + openai-codex/*) printf 'codex\n' ;; + claude-bridge/*) printf 'claude\n' ;; + *) return 1 ;; + esac + ;; claude) printf 'claude\n' ;; codex) printf 'codex\n' ;; opencode) printf 'codex\n' ;; @@ -352,7 +371,10 @@ for c in "${CANDIDATES[@]}"; do [ "$model" = "$c" ] && model="default" [ -n "$model" ] || die "invalid candidate: $c" fm_control_harness_supported "$harness" || die "unknown harness: $harness" - provider_for_harness "$harness" >/dev/null || die "unknown harness: $harness" + provider_for_harness "$harness" "$model" >/dev/null || case "$harness" in + omp) die "omp quota mapping covers only the openai-codex and claude-bridge prefixes: $model" ;; + *) die "unknown harness: $harness" ;; + esac done chosen="none" @@ -360,8 +382,10 @@ for c in "${CANDIDATES[@]}"; do harness=${c%%:*} model=${c#*:} [ "$model" = "$c" ] && model="default" - provider=$(provider_for_harness "$harness") - effective=$(effective_for_provider_model "$provider" "$model") + provider=$(provider_for_harness "$harness" "$model") + scope_model=$model + [ "$harness" != omp ] || scope_model=${model#*/} + effective=$(effective_for_provider_model "$provider" "$scope_model") if [ -z "$effective" ] || [ "$effective" = "null" ]; then continue fi diff --git a/bin/fm-remote-file.sh b/bin/fm-remote-file.sh index 34a993db5b4..31887ac27ba 100755 --- a/bin/fm-remote-file.sh +++ b/bin/fm-remote-file.sh @@ -77,7 +77,7 @@ snapshot_bounded_file() { # <file> <max-bytes> <destination> <size-file> directory_identity() { if [ "$(uname)" = Darwin ]; then - stat -f '%d:%i' . 2>/dev/null + /usr/bin/stat -f '%d:%i' . 2>/dev/null else stat -c '%d:%i' . 2>/dev/null fi diff --git a/bin/fm-remote-home-seed.sh b/bin/fm-remote-home-seed.sh index 7deafc40dcf..d1a434f1f24 100755 --- a/bin/fm-remote-home-seed.sh +++ b/bin/fm-remote-home-seed.sh @@ -157,7 +157,7 @@ done < "$BRIEF" > "$TMP/charter.remote" PROJECTS_CSV= : > "$TMP/project.records" PROJECT_INDEX=0 -for project in "${PROJECT_NAMES[@]}"; do +for project in "${PROJECT_NAMES[@]+"${PROJECT_NAMES[@]}"}"; do ORIGIN=${PROJECT_ORIGINS[$PROJECT_INDEX]} PROJECT_INDEX=$((PROJECT_INDEX + 1)) MODE_LINE=$(FM_HOME="$FM_HOME" FM_DATA_OVERRIDE="$DATA" "$SCRIPT_DIR/fm-project-mode.sh" "$project") diff --git a/bin/fm-remote-inherit-push.sh b/bin/fm-remote-inherit-push.sh index ed068622986..518e849b762 100755 --- a/bin/fm-remote-inherit-push.sh +++ b/bin/fm-remote-inherit-push.sh @@ -27,7 +27,7 @@ sha256_file() { if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi } file_link_count() { - if [ "$(uname)" = Darwin ]; then stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi + if [ "$(uname)" = Darwin ]; then /usr/bin/stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi } shared_captain_header_valid() { local head @@ -69,7 +69,8 @@ while IFS= read -r rel; do config/*) source="$CONFIG/${rel#config/}" ;; data/*) source="$DATA/${rel#data/}" ;; esac - if [ -e "$source" ] || [ -L "$source" ]; then + source_present=$(fm_config_source_present "$source") || exit 1 + if [ "$source_present" = 1 ]; then [ -f "$source" ] && [ ! -L "$source" ] || die "inherited source is unsafe: $source" [ "$(file_link_count "$source")" = 1 ] || die "inherited source is hardlinked: $source" if [ "$rel" = data/captain-shared.md ]; then diff --git a/bin/fm-remote-inherit.sh b/bin/fm-remote-inherit.sh index be995d75c70..15bb0d4cb1c 100755 --- a/bin/fm-remote-inherit.sh +++ b/bin/fm-remote-inherit.sh @@ -22,7 +22,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" die() { printf 'error: %s\n' "$1" >&2; exit 1; } usage() { sed -n '2,10p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } file_link_count() { - if [ "$(uname)" = Darwin ]; then stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi + if [ "$(uname)" = Darwin ]; then /usr/bin/stat -f %l "$1" 2>/dev/null; else stat -c %h "$1" 2>/dev/null; fi } sha256_file() { if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi diff --git a/bin/fm-remote-job-lib.sh b/bin/fm-remote-job-lib.sh index f6ac2ad9b99..22e42a4b4ab 100755 --- a/bin/fm-remote-job-lib.sh +++ b/bin/fm-remote-job-lib.sh @@ -758,7 +758,7 @@ fm_remote_job_reap() { # <account-home> <id>; only removes an exact completed re fm_remote_job_path_mtime() { # <path> # The platform override controls worker shape in isolated tests, not the host # kernel's stat syntax. - if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi + if [ "$(uname -s 2>/dev/null || true)" = Darwin ]; then /usr/bin/stat -f %m "$1" 2>/dev/null; else stat -c %Y "$1" 2>/dev/null; fi } fm_remote_job_stage_owner_alive() { # <stage-dir> diff --git a/bin/fm-remote-secondmate-control.sh b/bin/fm-remote-secondmate-control.sh index 82a39f1e69e..b40ad4d59b8 100755 --- a/bin/fm-remote-secondmate-control.sh +++ b/bin/fm-remote-secondmate-control.sh @@ -286,6 +286,7 @@ cmd_send() { case "$ring_rc" in 1) printf 'notice: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at %s\n' "$rec" >&2 ;; 2) printf 'notice: doorbell did not reach %s; the steer is durably recorded at %s\n' "$REMOTE_ENDPOINT_TARGET" "$rec" >&2 ;; + 3) printf 'notice: doorbell not typed because the agent in %s has exited; the steer is durably recorded at %s for recovery\n' "$REMOTE_ENDPOINT_TARGET" "$rec" >&2 ;; esac } diff --git a/bin/fm-secondmate-restart.sh b/bin/fm-secondmate-restart.sh index 452f614609f..9f3b90ceffa 100755 --- a/bin/fm-secondmate-restart.sh +++ b/bin/fm-secondmate-restart.sh @@ -324,20 +324,36 @@ done while [ "$((pending_count + restart_active_count))" -gt 0 ]; do now=$(date +%s) next_wait=$PERSIST_POLL + # Resolve every arrived answer before processing any timeout. Delivery of a + # later fleet request can outlast an earlier mate's deadline under load; that + # expired mate must not hold an already-confirmed mate behind its fallback. + i=0 + while [ "$i" -lt "${#IDS[@]}" ]; do + if [ "${PLAN[i]}" = persisted-pending ] \ + && fm_pending_reply_try_resolve "$STATE" "${CORR[i]}"; then + pending_count=$((pending_count - 1)) + launch_restart "$i" + fi + i=$((i + 1)) + done i=0 while [ "$i" -lt "${#IDS[@]}" ]; do if [ "${PLAN[i]}" != persisted-pending ]; then i=$((i + 1)) continue fi - if fm_pending_reply_try_resolve "$STATE" "${CORR[i]}"; then - pending_count=$((pending_count - 1)) - launch_restart "$i" - elif [ "$now" -ge "${DEADLINE[i]}" ]; then - fall_back_to_nudge "${IDS[$i]}" \ - "it did not confirm within ${PERSIST_WAIT}s that its open work is written down, so its conversation was not spent" - PLAN[i]="done" - pending_count=$((pending_count - 1)) + if [ "$now" -ge "${DEADLINE[i]}" ]; then + # A reply can land after the fleet-wide resolution pass. Recheck at the + # timeout decision so an answer already on disk wins over the fallback. + if fm_pending_reply_try_resolve "$STATE" "${CORR[i]}"; then + pending_count=$((pending_count - 1)) + launch_restart "$i" + else + fall_back_to_nudge "${IDS[$i]}" \ + "it did not confirm within ${PERSIST_WAIT}s that its open work is written down, so its conversation was not spent" + PLAN[i]="done" + pending_count=$((pending_count - 1)) + fi else remaining=$((DEADLINE[i] - now)) [ "$remaining" -ge "$next_wait" ] || next_wait=$remaining diff --git a/bin/fm-send.sh b/bin/fm-send.sh index 91df0b9fb2f..79bfd8480f5 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -43,9 +43,11 @@ # instruction. There is no delivered-unconfirmed # outcome on this plane: "did the doorbell land" is no longer the question - # "was the message acted on" is, and that is answered asynchronously for an -# ordinary record by the worker's acknowledgement move into handled/, with the -# watcher re-ringing an unacknowledged message and escalating a stuck one. An -# explicit fire-and-forget record is excluded from that ladder. +# ordinary record by the worker's acknowledgement move into handled/. The +# watcher re-rings an unacknowledged message while its endpoint remains +# available, escalates after the bounded ladder, and instead routes a positively +# dead or missing endpoint directly to recovery without typing. An explicit +# fire-and-forget record is excluded from that ladder. # bin/fm-task-inbox-lib.sh owns the record format, the doorbell line, and the # re-ring ladder. The composer pre-check before the ring is ADVISORY only: when # the composer visibly holds pending text the ring is skipped with a notice and @@ -1015,12 +1017,14 @@ else fm_send_feed_resolved_holds "$RESOLVE_ANSWER_TEXT" || exit 1 fi # Ring the doorbell, best-effort: no ring outcome changes the exit status, - # because the watcher's re-ring ladder owns loss detection from here. + # because the watcher owns loss detection from here, either through its + # bounded re-ring ladder or direct unavailable-endpoint recovery. ring_rc=0 fm_task_inbox_ring "$TARGET_BACKEND" "$T" "$INBOX_RECORD" "$EXPECTED_LABEL" || ring_rc=$? case "$ring_rc" in 1) echo "fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;; 2) echo "fm-send: doorbell did not reach $T; the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;; + 3) echo "fm-send: doorbell not typed because the agent in $T has exited; the steer is durably recorded at $INBOX_RECORD for recovery (stuck-crewmate-recovery), and the watcher will not re-ring a dead pane" >&2 ;; esac exit 0 fi diff --git a/bin/fm-session-lock-lib.sh b/bin/fm-session-lock-lib.sh index debb92d389a..471feaf4d69 100644 --- a/bin/fm-session-lock-lib.sh +++ b/bin/fm-session-lock-lib.sh @@ -16,14 +16,16 @@ # shellcheck source=bin/fm-cursor-lib.sh . "$(dirname -- "${BASH_SOURCE[0]}")/fm-cursor-lib.sh" -# Known harness command names; extend when a new adapter is verified. -FM_HARNESS_RE='claude|codex|opencode|grok|kimi|^copilot(\.exe)?$|^pi$|^pi-signed$' +# Known harness command names; extend when a new adapter is verified. omp is +# anchored exactly like pi: its process name is the bare word `omp` (verified, +# omp 18.1.11), and a substring match would claim ompd or comp. +FM_HARNESS_RE='claude|codex|opencode|grok|kimi|^copilot(\.exe)?$|^pi$|^pi-signed$|^omp$' # The same harnesses as exact executable names. Keep in sync with # FM_HARNESS_RE. Used only for the stricter path evidence below, where the # loose regex would also match ordinary firstmate paths such as # bin/fm-claude-stop-autoarm.sh. -FM_HARNESS_NAMES=(claude codex opencode grok kimi copilot copilot.exe pi-signed pi) +FM_HARNESS_NAMES=(claude codex opencode grok kimi copilot copilot.exe pi-signed pi omp) fm_session_process_comm() { # <pid> local pid=$1 proc_root=${FM_PROC_ROOT_OVERRIDE:-/proc} @@ -125,7 +127,6 @@ fm_copilot_loader_pid() { fi printf '%s\n' "$pid" } - # Print the exact harness name carried by executable path $1 - its own basename # or any directory component - or return 1. # diff --git a/bin/fm-session-start.sh b/bin/fm-session-start.sh index 7e4b438008d..a790916a4ec 100755 --- a/bin/fm-session-start.sh +++ b/bin/fm-session-start.sh @@ -819,6 +819,24 @@ if [ "$PRIMARY_HARNESS" = pi ] || [ "$PRIMARY_HARNESS" = pi-signed ]; then printf 'PI_WATCH_EXTENSION: not loaded - approve Pi project trust once per clone, then restart %s so %s and %s auto-load for turn-end guard and background wake coverage; use -e %s -e %s only if project hooks are not trusted\n' "$PI_RESTART_COMMAND" "$PI_TURNEND_EXT" "$PI_EXT" "$PI_TURNEND_EXT" "$PI_EXT" fi fi +# omp (Oh My Pi) has no project-trust gate: it auto-discovers <cwd>/.omp/extensions +# with no dialog, so the only ways both tracked primary extensions fail to load +# are a session started outside this home, an extension disabled in the omp +# config, or a build older than the tracked file. The markers carry the loaded +# build plus the loading pid, exactly as the Pi ones do (bin/fm-wake-lib.sh). +if [ "$PRIMARY_HARNESS" = omp ]; then + OMP_EXT="$FM_ROOT/.omp/extensions/fm-primary-omp-watch.ts" + OMP_TURNEND_EXT="$FM_ROOT/.omp/extensions/fm-primary-turnend-guard.ts" + OMP_WATCH_MARKER="$STATE/.omp-watch-extension-loaded" + OMP_TURNEND_MARKER="$STATE/.omp-turnend-extension-loaded" + OMP_LOCK="$STATE/.lock" + OMP_WATCH_VERSION=$(fm_pi_extension_version "$OMP_EXT" || printf '') + OMP_TURNEND_VERSION=$(fm_pi_extension_version "$OMP_TURNEND_EXT" || printf '') + if ! fm_pi_extension_loaded "$OMP_WATCH_MARKER" "$OMP_WATCH_VERSION" "$OMP_LOCK" \ + || ! fm_pi_extension_loaded "$OMP_TURNEND_MARKER" "$OMP_TURNEND_VERSION" "$OMP_LOCK"; then + printf 'OMP_WATCH_EXTENSION: not loaded - restart omp with this home as its working directory so %s and %s auto-load from .omp/extensions/ for turn-end guard and background wake coverage; pass -e %s -e %s only when omp must start from another directory, never together with auto-discovery (omp loads a file named both ways twice)\n' "$OMP_TURNEND_EXT" "$OMP_EXT" "$OMP_TURNEND_EXT" "$OMP_EXT" + fi +fi "$SCRIPT_DIR/fm-supervision-instructions.sh" \ --harness "$PRIMARY_HARNESS" \ --read-only "$READ_ONLY" \ diff --git a/bin/fm-spawn.sh b/bin/fm-spawn.sh index de01932a6da..c6f3fa1ca2e 100755 --- a/bin/fm-spawn.sh +++ b/bin/fm-spawn.sh @@ -14,15 +14,18 @@ # scaffolded before that line existed warns once and launches on the flag. A # ship or scout spawn also refuses leftover `{TASK}` / `{FIRSTMATE_SPEC}` # placeholders, an empty Task, or an incomplete pair of Task subsections. -# For a no-mistakes ship, spawn renders `launch-brief.md` with the current -# `--intent` contract and the extracted captain intent. A legacy mixed Task is -# accepted there only under bin/fm-dod-lib.sh's provenance-marking rules; -# unmarked legacy Tasks stop for migration rather than becoming intent. That -# library owns the parsing and intent rules. When the explicit mode carries -# less rigor than the project's standing posture, a loud one-line deviation -# notice is printed and the spawn continues. +# Every ship or scout spawn renders `launch-brief.md`; for a no-mistakes ship +# it also carries the current `--intent` contract and the extracted captain +# intent. A legacy mixed Task is accepted there only under bin/fm-dod-lib.sh's +# provenance-marking rules; unmarked legacy Tasks stop for migration rather +# than becoming intent. That library owns the parsing and intent rules. When +# the explicit mode carries less rigor than the project's standing posture, a +# loud one-line deviation notice is printed and the spawn continues. # no-mistakes-prod-only is a registry policy rather than a task mode and is # refused as a flag value. +# Ship/scout launches always supply fm-dod-lib.sh's current worker role scope +# using the same private launch-brief overlay. This never rewrites a project's +# instruction files or a secondmate's charter. # fm-spawn.sh <task-id> --relaunch [--harness <name>] [--model <name>] [--effort <level>] # --relaunch launches a replacement agent for an EXISTING task into that # task's own recorded endpoint and worktree instead of creating either. It is @@ -108,13 +111,21 @@ # even when they select different backends. A fresh spawn first takes the # per-home task-set lock and refuses rather than waits when forced teardown owns # it; relaunch is exempt because the existing task's control lock covers it. +# A fresh Treehouse-backed spawn also takes the project-identity lock in the local +# root Firstmate home's state directory before slot allocation and holds it through +# task metadata publication. Teardown holds that same lock while proving and +# returning a slot, so allocation cannot reuse a slot before its owner record +# is published. The local root is whatever bin/fm-wake-lib.sh's +# fm_firstmate_root_home resolves, so a home seeded from another machine anchors +# that lock itself rather than failing to resolve one; +# contention refuses rather than waits. # With no harness arg, a crewmate/scout spawn resolves the CREW harness only when # config/crew-dispatch.json is absent. When that file exists, crewmate/scout # spawns require an explicit harness so firstmate cannot silently skip dispatch # profile consultation. A --secondmate spawn is exempt and resolves the SECONDMATE # harness (config/secondmate-harness -> config/crew-harness -> own), so the # secondmate-vs-crewmate split is DURABLE across every respawn (recovery, -# /updatefirstmate, restart). A bare adapter name (claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse) +# /updatefirstmate, restart). A bare adapter name (claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) # overrides it for this spawn (either kind). A non-flag string containing # whitespace is treated as a RAW launch command - the escape hatch for verifying # new adapters. For pi and pi-signed, fm-spawn resolves the selected executable @@ -123,6 +134,27 @@ # a failed or inconclusive probe omits it so older Pi versions remain launchable. # A missing selected executable refuses before endpoint creation, and pi-signed # never falls back to pi. +# For omp (Oh My Pi), fm-spawn resolves the `omp` executable from PATH once and +# refuses when it is absent. Every omp launch clears the foreign harness +# markers (omp publishes none of its own), sets the Firstmate-owned +# FM_OMP_HARNESS=omp detection marker, suppresses the first-run provider +# wizard with OMP_SKIP_SETUP=1, forces --auto-approve, pins the working +# directory with --cwd, and passes the tracked worker posture overlay +# .omp/fm-worker-overlay.yml through --config. That overlay pins composer +# shape, plan mode off, prewalk off, and the non-interactive usage-reserve +# policy for the one session only (--auto-approve alone owns approval); the +# captain's own ~/.omp/agent/config.yml (model roles, providers, theme) is +# never written. +# A model written as <provider>/<id> is validated against `omp models --json` +# only when that provider appears in the listing; a provider absent from the +# listing (an extension-registered provider such as claude-bridge, which omp +# never lists) passes through unvalidated with a stderr notice, and a bare +# fuzzy pattern is left to omp's own matcher. A crewmate or scout loads its +# per-task busy-state extension with -e from state/ (outside the worktree, so +# auto-discovery cannot load it a second time); a secondmate passes no -e at +# all and relies on omp auto-discovering the home's tracked .omp/extensions/ +# (verified, omp 18.1.11: a file named both ways loads twice, and discovery is +# cwd-only with no trust dialog). # config/secondmate-harness may also carry an optional model and effort as extra # whitespace-separated tokens ("<harness> [<model>] [<effort>]"). For a # --secondmate spawn, those tokens apply only when this spawn also resolves its @@ -146,11 +178,27 @@ # Ship/scout spawns refuse to launch unless the resolved task path is a real # git worktree root distinct from both the spawning project and its repository's # primary checkout, including when the spawning project is a linked worktree. -# Only after this isolation check, a fresh ship or scout's clean task worktree -# fetches origin, resolves the current remote default branch, and resets to its tip. -# Relaunch reuses the recorded worktree without fetching or resetting its base. -# An unreachable origin, unresolved default branch, or non-clean worktree -# refuses a fresh spawn rather than risking a PR based on stale history. +# On the backends that discover that path by reading the task pane's own cwd, +# the same isolation test screens every read: a pane still showing the project +# or the repository primary while `treehouse get` prepares the slot is waited +# out as a transient rather than adopted and then refused, so a home that is +# itself a linked worktree of the project repository still launches. A pane +# that never reaches an isolated worktree refuses at the end of that wait, +# naming the last path seen and why it was rejected. +# That placement is proven only at launch. Every ship or scout pane therefore +# also receives `export FM_TASK_ID=<task-id>` before the launch command, on +# the same channel as GOTMPDIR, and bin/fm-test-run.sh refuses to execute the +# behavior suite from the repository primary checkout while that marker is +# set (its header owns the refusal). A secondmate runs in its own home and is +# not marked. +# Only after this isolation check, every fresh ship or scout requires a clean +# task worktree. When an origin configuration is detected, spawn fetches it, +# resolves the current remote default branch, and resets to its tip. When none +# is detected, spawn skips that remote freshness check and launches from the +# clean worktree's current HEAD. Relaunch reuses the recorded worktree without +# fetching or resetting its base. An unreachable detected origin, unresolved +# default branch, or non-clean worktree refuses a fresh spawn rather than +# risking a PR based on stale history or discarding local work. # A slot whose only deviation is a stale submodule gitlink is refused by that # same clean check, but is reported as a stale checkout naming each submodule # and both pins; nothing is converged or removed, and no remedy is suggested. @@ -170,6 +218,29 @@ # and scout batches. The loop lives here, in bash, so callers never hand-write a # multi-task shell loop (the tool shell is zsh, which does not word-split unquoted # $vars and silently breaks ad-hoc `for ... in $pairs` loops). +# Launch environment (config/launch-env-allowlist): +# Absent means unchanged ambient inheritance. A present readable regular file +# opts every launch (ship, scout, secondmate, raw command, and relaunch) into +# /usr/bin/env -i followed by /bin/sh -c of the existing launch command. +# Each line is one POSIX environment name, never a value or shell expression; +# blank lines and lines beginning with # are ignored. Invalid input refuses +# before launch, as do path inspection errors such as inaccessible config +# directories. An empty file retains only the operational floor below. +# Names are read once per spawn; values are expanded in the destination pane, +# not copied from the invoking process or written into the launch text. +# Unset names stay unset and empty values stay empty. +# The fixed operational floor is HOME PATH USER LOGNAME SHELL TERM COLORTERM +# LANG LC_ALL LC_CTYPE TMPDIR TMP TEMP GOTMPDIR, plus backend identity/routing: +# TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH HERDR_PANE_ID +# CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID CMUX_SOCKET_PATH +# ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION, plus the task +# marker FM_TASK_ID that ship and scout panes receive above. +# An enabled task trace also retains TRACEPARENT. Explicit Firstmate launch +# assignments still apply inside the filtered environment. Raw commands must +# be POSIX sh compatible under this opt-in; the absent-file path is unchanged. +# This is an exec environment boundary, not a sandbox for the pane's startup +# shell, credential files, same-user processes, or later shell initialization. +# See docs/configuration.md for provider/Git setup and supported limits. # Launch templates live in launch_template() below; placeholders replaced before launch: # __BRIEF__ absolute path to data/<task-id>/brief.md # __PIBIN__ quoted concrete Pi-family executable path resolved from PATH @@ -180,10 +251,16 @@ # written by this script; outside the worktree to avoid pi's trust gate) # __PITURNEND__ absolute path to .pi/extensions/fm-primary-turnend-guard.ts in a pi secondmate home # __PIWATCH__ absolute path to .pi/extensions/fm-primary-pi-watch.ts in a pi secondmate home +# __OMPBIN__ quoted concrete omp executable path resolved from PATH +# __OMPEXT__ absolute path to state/<task-id>.omp-ext.ts (omp busy-state and +# turn-end extension, written by this script; outside the worktree so +# omp's cwd-only auto-discovery cannot load it a second time) +# __OMPWORKERCFG__ absolute path to the tracked .omp/fm-worker-overlay.yml posture overlay # __OPINPUT__ absolute path to the canonical operational-input encoder # __WORKTREE__ absolute path to the task worktree # __CURSORBIN__ resolved, cursor-verified executable for a cursor launch # __GEMINISETTINGS__ firstmate-owned per-task gemini settings file (busy-state hooks) +# __ROVOBIN__ resolved, rovo-verified executable for a rovo launch # __COPILOTBIN__ resolved GitHub Copilot CLI executable # Verified per-harness turn-end hooks are installed automatically where enabled; some live outside the worktree. # Kimi uses one surgically installed Firstmate region in $HOME/.kimi-code/config.toml, @@ -193,6 +270,13 @@ # muse installs no hook at all - its plugin engine is off in the default build - so # it writes state/<id>.muse-session to bind the pane to muse's own session event # log; muse and gemini are crewmate/scout only and are refused for --secondmate. +# rovo installs no hook either - its eventHooks fire at tool granularity only, +# never turn-end - so it carries no busy-source wiring at all and no turn-end +# hook. A positional brief is dead-on-arrival (rovo loads, never works, and drops +# to an idle shell), so rovo launches BARE and receives an absolute brief pointer +# only after a TUI readiness gate, then a delivery-confirmation gate - the same +# launch-then-send shape as kimi. Its busy state is a screen-scrape fallback like +# grok. rovo is crewmate/scout only and is refused for --secondmate, like muse. # cursor installs no per-task hook either: it writes state/<id>.cursor-session to # bind the pane to cursor's own conversation transcript (projects root, the exact # workspace path cursor records in .workspace-trusted, and the conversations that @@ -209,6 +293,11 @@ # and every refusal; a failed registration stops this spawn rather than launching # a worker that would wedge on the dialog. A --secondmate launch never runs it, # so a claude secondmate home keeps its own one-time trust decision. +# Every claude launch also carries the attribution-off policy in its per-launch +# --settings JSON, so a spawned worker never writes a Co-Authored-By trailer, +# Claude-Session link, or generated-with line into a commit or PR body; +# launch_template() below owns the reason it cannot come from the captain's own +# settings. # Publishing the record and moving this home's backlog item to In flight are one # step, not two: bin/fm-backlog-transition-lib.sh owns that invariant, and this # script performs the transition under the task's own meta lock before it reports @@ -292,6 +381,26 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" PROJECTS="${FM_PROJECTS_OVERRIDE:-$FM_HOME/projects}" CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" +# shellcheck source=bin/fm-config-inherit-lib.sh +. "$SCRIPT_DIR/fm-config-inherit-lib.sh" +if ! LAUNCH_ENV_ENABLED=$(fm_config_source_present "$CONFIG/launch-env-allowlist"); then + exit 1 +fi +LAUNCH_ENV_NAMES= +if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then + if [ ! -f "$CONFIG/launch-env-allowlist" ] || [ ! -r "$CONFIG/launch-env-allowlist" ]; then + echo "error: config/launch-env-allowlist must be a readable regular file" >&2 + exit 1 + fi + if ! LAUNCH_ENV_NAMES=$(jq -Rrs ' + split("\n") | map(select(. != "" and (startswith("#") | not))) | + if all(.[]; test("^[A-Za-z_][A-Za-z0-9_]*$")) then .[] + else error("expected environment names only") end + ' "$CONFIG/launch-env-allowlist" 2>/dev/null); then + echo "error: config/launch-env-allowlist must contain one environment name per line, blank lines, or # comments" >&2 + exit 1 + fi +fi SUB_HOME_MARKER=".fm-secondmate-home" if [ -e "$STATE" ] || [ -L "$STATE" ]; then fm_backlog_directory_present "$STATE" "state directory" || { @@ -309,8 +418,6 @@ fm_backlog_directory_present "$STATE" "state directory" || { } # shellcheck source=bin/fm-secondmate-nudge-lib.sh . "$SCRIPT_DIR/fm-secondmate-nudge-lib.sh" -# shellcheck source=bin/fm-config-inherit-lib.sh -. "$SCRIPT_DIR/fm-config-inherit-lib.sh" # shellcheck source=bin/fm-backend.sh . "$SCRIPT_DIR/fm-backend.sh" # shellcheck source=bin/fm-control-lib.sh @@ -763,6 +870,8 @@ SPAWN_META_PUBLISH_STARTED=0 SPAWN_FRESH_COMMIT_PENDING=0 SPAWN_TASK_SET_LOCK= SPAWN_TASK_SET_LOCK_HELD=0 +SPAWN_TREEHOUSE_PROJECT_LOCK= +SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 RELAUNCH_REPLACEMENT_PENDING=0 RELAUNCH_REPLACEMENT_BUSY_GEN= RELAUNCH_REPLACEMENT_HARNESS= @@ -938,6 +1047,10 @@ spawn_abort_cleanup() { SPAWN_META_LOCK_HELD=0 fm_lock_release "$SPAWN_META_LOCK" || true fi + if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then + SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 + fm_lock_release "$SPAWN_TREEHOUSE_PROJECT_LOCK" || true + fi if [ "$SPAWN_TASK_SET_LOCK_HELD" = 1 ]; then SPAWN_TASK_SET_LOCK_HELD=0 fm_lock_release "$SPAWN_TASK_SET_LOCK" || true @@ -1259,7 +1372,7 @@ if [ "$RELAUNCH" -eq 1 ]; then } elif [ "$KIND" = secondmate ]; then case "${POS[1]:-}" in - ''|claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse) + ''|claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) ARG3=${POS[1]:-} ;; *' '*) @@ -1319,6 +1432,34 @@ pi_supports_tui_mode() { printf '%s\n' "$help" | grep -Eq -- '(^|[[:space:]])--tui-mode([[:space:]=]|$)' } +# omp pre-launch model validation. `omp models --json` (omp 18.1.11) prints +# {"models":[{"provider","id","selector":"<provider>/<id>",...}]} for built-in and +# auto-discovered providers only; it never lists a provider an extension +# registers at runtime (claude-bridge is the verified example), so the check is +# scoped exactly to what the listing can prove: a <provider>/<id> whose provider +# IS listed must be listed too, a provider the listing does not know passes +# through with a notice, a bare fuzzy pattern is omp's own matcher's job, and an +# unreadable listing establishes nothing (harness-adapters model-and-effort.md). +omp_model_validate() { # <omp-bin> <model> + local bin=$1 model=$2 provider listing providers + [ -n "$model" ] && [ "$model" != default ] || return 0 + case "$model" in */*) ;; *) return 0 ;; esac + command -v jq >/dev/null 2>&1 || return 0 + listing=$(OMP_SKIP_SETUP=1 "$bin" models --json 2>/dev/null) || return 0 + providers=$(printf '%s' "$listing" | jq -r '.models[]?.provider // empty' 2>/dev/null | sort -u) || return 0 + [ -n "$providers" ] || return 0 + provider=${model%%/*} + if ! printf '%s\n' "$providers" | grep -qxF -- "$provider"; then + echo "notice: omp provider '$provider' is not in 'omp models --json' (extension-registered providers are never listed); launching '$model' unvalidated" >&2 + return 0 + fi + if printf '%s' "$listing" | jq -e --arg m "$model" '.models[]? | select(.selector == $m)' >/dev/null 2>&1; then + return 0 + fi + echo "error: omp model '$model' is not listed by 'omp models --json' although provider '$provider' is; choose a listed <provider>/<id> or omit --model" >&2 + return 1 +} + # The verified launch command per adapter. The knowledge half of each adapter # (busy-state source, exit command, dialogs, quirks) lives in the harness-adapters skill. launch_template() { @@ -1344,7 +1485,15 @@ launch_template() { # alone disables the feature; keep both so a managed override of one still # leaves the other in force. Both are per-launch, scoped to this invocation only, # and never touch the captain's global ~/.claude/settings.json. - claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off"}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # The same inline --settings JSON also carries the attribution policy + # ("attribution": {"commit": "", "pr": "", "sessionUrl": false}), which + # suppresses Claude Code's Co-Authored-By trailer, Claude-Session link, and + # generated-with line in commits and PR bodies. The captain sets that + # policy in the `user` settings scope, but a launched worker's settings + # sources are not guaranteed to load that scope, so a worker would + # otherwise run with attribution back on; carrying it per launch keeps the + # policy in force regardless of which settings scopes end up loaded. + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '\''{"feedbackDrafts":"off","attribution":{"commit":"","pr":"","sessionUrl":false}}'\'' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; codex) if [ "$kind" = secondmate ]; then printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' @@ -1362,6 +1511,26 @@ launch_template() { printf '%s' ' __MODELFLAG____EFFORTFLAG__-e __PIEXT__ "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' fi ;; + # omp (Oh My Pi), a Pi fork. Same one-positional-brief, --model, --thinking, + # and -e shape as Pi, verified on omp 18.1.11. The differences are all at + # the launch boundary and documented in the header above: foreign markers + # cleared (omp has none of its own, so an inherited CLAUDECODE would win), + # FM_OMP_HARNESS=omp established for bin/fm-harness.sh, OMP_SKIP_SETUP=1 + # against the fresh-profile provider wizard, --auto-approve so no approval + # prompt can park an unattended worker, the tracked posture overlay so a + # captain-level plan, prewalk, or usage dialog cannot either, and --cwd + # pinned to the worktree because omp's extension discovery is cwd-only. A + # secondmate loads its two primary extensions by that discovery alone: + # naming them with -e as well loads each twice (verified), doubling every + # session_stop continuation. + omp) + printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS -u GEMINI_CLI -u CURSOR_AGENT -u CURSOR_INVOKED_AS FM_OMP_HARNESS=omp OMP_SKIP_SETUP=1 __OMPBIN__ --config __OMPWORKERCFG__ --auto-approve --cwd __WORKTREE__' + if [ "$kind" = secondmate ]; then + printf '%s' ' __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + else + printf '%s' ' __MODELFLAG____EFFORTFLAG__-e __OMPEXT__ "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + fi + ;; # grok (Grok Build TUI): a positional prompt starts the supervised interactive # session. --always-approve auto-approves every tool execution (verified: the # crewmate runs fully autonomously, no permission gate), which an unattended @@ -1445,6 +1614,33 @@ launch_template() { # written below. Nothing to place in the template for it. # codex, opencode, and kimi are also markerless and share this inherited-marker hazard; changing their verified launch boundaries belongs in follow-up work. muse) printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS XDG_CONFIG_HOME=__MUSECONFIG__ XDG_DATA_HOME=__MUSEDATA__ MUSE_EXPERIMENTAL_FOREIGN_PERSONAL_CONTEXT_KILL=on __MUSEBIN__ --yolo __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # rovo (Atlassian Rovo CLI): a positional brief is dead-on-arrival - rovo + # loads, never enters a working state, and drops back to an idle shell within + # about 10-15 seconds (confirmed live four times over a raw PTY and once under + # real tmux with the exact send-keys shape below). So rovo launches BARE, + # exactly like kimi, and receives an absolute brief pointer only after the TUI + # readiness gate below. --disable-permission-checks/--yolo makes every file + # CRUD operation and bash command run without confirmation; Atlassian-data and + # user MCP-server tools still prompt per its own printed caveat, which crew and + # scout tasks never touch. --startup-receipt is not used either: it requires + # "prompt-free interactive mode", so it cannot gate a launch that will have a + # message typed into it. rovo does NOT scrub an inherited + # CLAUDECODE/CURSOR_AGENT/etc, so foreign primary markers are cleared here as + # defense in depth alongside the marker-ordering fix in bin/fm-harness.sh + # (issue #3517); CURSOR_AGENT/CURSOR_INVOKED_AS are cleared by the shared + # outer wrap below, like every other non-cursor harness. rovo has no + # turn-end hook (its eventHooks fire at tool granularity only, never + # turn-end), so no launch placeholder for one exists. + # __ROVOCONFIGOVERRIDE__ (not __EFFORTFLAG__) carries rovo's single + # --config-override flag: it always grants allowedExternalPaths for this + # task's home-side brief dir, steering inbox, and status file - the file + # tool confinement that otherwise blocks the standard + # instructions/steering/status/report loop (rovo's bash tool has no such + # grant and stays confined to the worktree; the worker's own file tools do + # respect the grant, confirmed live) - merged with agent.efficiencyLevel + # when a supported effort is requested, since a second --config-override + # would silently discard the first (confirmed live). + rovo) printf '%s' 'env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS __ROVOBIN__ run --yolo __MODELFLAG____ROVOCONFIGOVERRIDE__' ;; *) return 1 ;; esac } @@ -1501,6 +1697,15 @@ if [ "$KIND" = secondmate ] && { [ "$HARNESS" = muse ] || [ "$HARNESS" = gemini exit 1 fi +# rovo carries the same primary-supervision gap as muse: no turn-end hook, no +# verified primary integration, so a secondmate (a firstmate instance that must +# itself act as a primary) could never be supervised. Refuse loudly rather than +# standing one up with no way to arm its watch cycle. +if [ "$KIND" = secondmate ] && [ "$HARNESS" = rovo ]; then + echo "error: rovo is a verified crewmate/scout adapter only and cannot run a secondmate; it has no primary supervision protocol. Select a harness verified for secondmates." >&2 + exit 1 +fi + case "$HARNESS" in pi|pi-signed) PI_BIN=$(resolve_pi_executable "$HARNESS") || { @@ -1530,6 +1735,17 @@ case "$HARNESS" in fi fi ;; + omp) + OMP_BIN=$(resolve_pi_executable omp) || { + echo "error: omp executable not found on PATH; install Oh My Pi or select a different verified harness" >&2 + exit 1 + } + OMP_WORKER_CFG="$FM_ROOT/.omp/fm-worker-overlay.yml" + [ -f "$OMP_WORKER_CFG" ] || { + echo "error: omp worker posture overlay missing at $OMP_WORKER_CFG; a worker launched without it can park on the captain's own approval or plan-mode settings" >&2 + exit 1 + } + ;; copilot) COPILOT_BIN=$(resolve_executable copilot) || { echo "error: copilot executable not found on PATH; install GitHub Copilot CLI or select a different verified harness" >&2 @@ -1559,6 +1775,9 @@ if [ "$KIND" = secondmate ] && [ -z "$ARG3" ]; then fi fi fi +if [ "$HARNESS" = omp ]; then + omp_model_validate "$OMP_BIN" "$MODEL" || exit 1 +fi secondmate_registry_value() { secondmate_registry_field "$DATA/secondmates.md" "$1" "$2" @@ -1607,6 +1826,30 @@ resolve_muse_binary() { return 1 } +resolve_rovo_binary() { + local candidate dir fallback + candidate=$(command -v rovo 2>/dev/null || true) + if [ -n "$candidate" ] && [ -x "$candidate" ]; then + case "$candidate" in + /*) printf '%s\n' "$candidate"; return 0 ;; + *) + dir=$(cd "$(dirname "$candidate")" 2>/dev/null && pwd -P) || dir= + if [ -n "$dir" ]; then + printf '%s/%s\n' "$dir" "$(basename "$candidate")" + return 0 + fi + ;; + esac + fi + fallback="${HOME:-}/.local/bin/rovo" + if [ -n "${HOME:-}" ] && [ -x "$fallback" ]; then + printf '%s\n' "$fallback" + return 0 + fi + echo "error: rovo executable not found; searched PATH for 'rovo' and fallback '$fallback'" >&2 + return 1 +} + # muse_credential_present: 0 when a launched muse pane can reach its provider # without an interactive login. muse offers exactly two credential paths # (verified, muse 0.1.0-R708.1): the META_API_KEY environment variable, which @@ -1618,6 +1861,12 @@ resolve_muse_binary() { # supervision like a wedged worker rather than a missing credential. muse_worker_meta_api_key_present() { local session worker_env + if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then + case $'\n'"$LAUNCH_ENV_NAMES"$'\n' in + *$'\nMETA_API_KEY\n'*) ;; + *) return 1 ;; + esac + fi [ "$BACKEND" = tmux ] || return 1 if [ -n "${TMUX:-}" ]; then session=$(tmux display-message -p '#S' 2>/dev/null) || return 1 @@ -1641,7 +1890,7 @@ model_flag_for_harness() { local harness=$1 model=$2 [ -n "$model" ] && [ "$model" != default ] || return 0 case "$harness" in - claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse) + claude|codex|copilot|opencode|pi|pi-signed|grok|kimi|cursor|gemini|muse|rovo|omp) printf -- '--model %s ' "$(shell_quote "$model")" ;; esac @@ -1685,6 +1934,13 @@ effort_flag_for_harness() { low|medium|high|xhigh|max) printf -- '--thinking %s ' "$(shell_quote "$effort")" ;; esac ;; + omp) + # omp 18.1.11 --thinking accepts off|minimal|low|medium|high|xhigh|max|auto, + # a superset of the shared vocabulary, so every level maps straight across. + case "$effort" in + low|medium|high|xhigh|max) printf -- '--thinking %s ' "$(shell_quote "$effort")" ;; + esac + ;; muse) # muse 0.1.0-R708.1 --reasoning-effort accepts none|minimal|low|medium| # high|xhigh|ultra and defaults to high, so low..xhigh map straight across. @@ -1699,6 +1955,10 @@ effort_flag_for_harness() { max) printf -- '--reasoning-effort %s ' "$(shell_quote ultra)" ;; esac ;; + # rovo has no --effort flag on `run`; its effort mapping rides + # --config-override, but that flag is single-value (see + # rovo_config_override_flag below) so it is built there, merged with the + # mandatory allowedExternalPaths grant, rather than here. # opencode's interactive `opencode --prompt` launch has a verified --model # flag but no verified effort flag. Its `opencode run --variant` flag belongs # to a different, non-interactive launch mode, so fm-spawn does not pass it. @@ -1742,10 +2002,50 @@ case "$LAUNCH" in ;; esac +case "$LAUNCH" in + *__ROVOBIN__*) + ROVO_BIN=$(resolve_rovo_binary) || exit 1 + LAUNCH=${LAUNCH//__ROVOBIN__/$(shell_quote "$ROVO_BIN")} + ;; +esac + json_escape() { printf '%s' "$1" | sed 's/\\/\\\\/g; s/"/\\"/g' } +# rovo confines every file-tool operation (open_files, create_file, grep, ...) +# to its worktree by default; toolPermissions.allowedExternalPaths +# (~/.rovo/config.yml) is the only lift, and it must be granted at launch +# through --config-override since there is no per-session escalation once +# the process is running. rovo's bash tool is NOT covered by this grant and +# stays confined to the worktree regardless (confirmed live) - the standard +# crewmate flow's literal `echo ... >> status file` bash line therefore still +# fails under rovo, but the worker recovers by falling back to its own file +# tools for the same append (confirmed live), which the grant below does cover. +# --config-override itself is single-value (a second occurrence silently +# discards the first, confirmed live), so this is the ONE place that must +# also fold in agent.efficiencyLevel when a supported effort was requested. +# Granted paths are real (symlink-resolved) directories/files under this +# task's home, matching BRIEF_REAL's own resolution: the brief dir (covers +# brief.md/launch-brief.md/report.md), the steering inbox directory (covers +# every steer and its handled/ acknowledgement), and the status file itself. +rovo_config_override_flag() { + local effort=$1 data_dir=$2 state_dir=$3 id=$4 + local data_real state_real agent_json paths_json config_json + data_real=$(cd "$data_dir" && pwd -P) || return 1 + state_real=$(cd "$state_dir" && pwd -P) || return 1 + agent_json= + case "$effort" in + low|medium|high|max) agent_json="\"agent\":{\"efficiencyLevel\":\"$(json_escape "$effort")\"}," ;; + esac + paths_json=$(printf '"%s","%s","%s"' \ + "$(json_escape "$data_real/$id")" \ + "$(json_escape "$state_real/$id.inbox")" \ + "$(json_escape "$state_real/$id.status")") + config_json="{${agent_json}\"toolPermissions\":{\"allowedExternalPaths\":[$paths_json]}}" + printf -- '--config-override %s ' "$(shell_quote "$config_json")" +} + resolved_existing_dir() { local path=$1 [ -d "$path" ] || { echo "error: firstmate home does not exist or is not a directory: $path" >&2; return 1; } @@ -1936,6 +2236,17 @@ else WT="" BRIEF="$DATA/$ID/brief.md" fi +if [ "$RELAUNCH" -eq 0 ] && [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then + SPAWN_TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$PROJ_ABS") || { + echo "error: could not resolve the shared Treehouse project lock for $PROJ_ABS" >&2 + exit 1 + } + if ! fm_lock_try_acquire "$SPAWN_TREEHOUSE_PROJECT_LOCK"; then + echo "error: another Treehouse slot allocation or return is in progress for $PROJ_ABS; refusing to race it" >&2 + exit 1 + fi + SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=1 +fi [ -f "$BRIEF" ] || { echo "error: task $ID has no brief at inaccessible data path $BRIEF" >&2; exit 1; } if [ "$KIND" = ship ] || [ "$KIND" = scout ]; then if fm_brief_task_placeholders_present "$BRIEF"; then @@ -1957,18 +2268,24 @@ if [ "$KIND" = ship ] || [ "$KIND" = scout ]; then exit 1 fi fi - SOURCE_BRIEF=$BRIEF - BRIEF="$DATA/$ID/launch-brief.md" - BRIEF_TMP="$DATA/$ID/.launch-brief.md.${BASHPID:-$$}" - { - cat "$SOURCE_BRIEF" - fm_brief_intent_overlay "$CAPTAIN_INTENT" - } > "$BRIEF_TMP" || { rm -f -- "$BRIEF_TMP"; echo "error: could not render current intent contract for $SOURCE_BRIEF" >&2; exit 1; } - if ! mv "$BRIEF_TMP" "$BRIEF"; then - rm -f -- "$BRIEF_TMP" - echo "error: could not publish current intent contract for $SOURCE_BRIEF" >&2 - exit 1 - fi + fi + # Use the existing launch-brief overlay for every worker kind, including + # pre-scope briefs and relaunches. Charters never enter this worker path. + SOURCE_BRIEF=$BRIEF + BRIEF="$DATA/$ID/launch-brief.md" + BRIEF_TMP="$DATA/$ID/.launch-brief.md.${BASHPID:-$$}" + { + cat "$SOURCE_BRIEF" && + printf '\n' && + fm_brief_worker_role && + if [ "$KIND" = ship ] && [ "$MODE" = no-mistakes ]; then + fm_brief_intent_overlay "$CAPTAIN_INTENT" + fi + } > "$BRIEF_TMP" || { rm -f -- "$BRIEF_TMP"; echo "error: could not render current launch contract for $SOURCE_BRIEF" >&2; exit 1; } + if ! mv "$BRIEF_TMP" "$BRIEF"; then + rm -f -- "$BRIEF_TMP" + echo "error: could not publish current launch contract for $SOURCE_BRIEF" >&2 + exit 1 fi fi @@ -2038,30 +2355,78 @@ real_path_or_raw() { # <path> # herdr-sm-spaces-k4). Both branches converge on the same $T ("target") string # that every downstream operation (send/capture/kill) already treats as opaque # per-backend routing (fm_backend_resolve_selector). -validate_spawn_worktree() { # <source> <inspect-target> - local source=$1 inspect_target=$2 wt_real proj_real wt_top wt_top_real - local wt_git_dir proj_common + +# True when <path> is an isolated worktree of the spawning project: a real +# directory that is its own worktree root, is not the spawning project itself, +# and does not share the project repository's common git dir. SPAWN_WT_TOP is +# left holding the worktree root the check read, and SPAWN_WT_REASON a short +# phrase naming why a rejected path failed, both for the refusal messages. +# +# The worktree-discovery poll below reads this same predicate, so it can never +# adopt a path the guard would then refuse. That matters because a pane's cwd +# read is a snapshot of whatever process is in the foreground: while `treehouse +# get` is still fetching and checking a slot out, it reports the REPOSITORY's +# primary checkout as its own cwd. That path differs from a linked spawning +# project, so a poll comparing only against the project accepted it, and the +# guard then refused a launch whose slot treehouse went on to create normally. +# A read like that is a transient, not a destination: the poll keeps waiting. +SPAWN_WT_TOP= +SPAWN_WT_REASON= +spawn_worktree_isolated() { # <path> + local path=$1 wt_real wt_top_real wt_git_dir proj_common + SPAWN_WT_TOP= + SPAWN_WT_REASON= wt_real= - if ! wt_real=$(cd "$WT" 2>/dev/null && pwd -P); then + if ! wt_real=$(cd "$path" 2>/dev/null && pwd -P); then wt_real= fi - proj_real=$PROJ_ABS_REAL - wt_top=$(git -C "$WT" rev-parse --show-toplevel 2>/dev/null || true) + if [ -z "$wt_real" ]; then + SPAWN_WT_REASON="it is not a readable directory" + return 1 + fi + SPAWN_WT_TOP=$(git -C "$path" rev-parse --show-toplevel 2>/dev/null || true) + # A path in no repository leaves the toplevel empty, and that empty value must + # never reach `cd`: bash before 5.3 accepts `cd ""` as a successful no-op, so + # it would resolve to fm-spawn's OWN cwd and report the path as a subdirectory + # of whatever checkout firstmate happens to be running from. wt_top_real= - if ! wt_top_real=$(cd "$wt_top" 2>/dev/null && pwd -P); then + if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then wt_top_real= fi + if [ -z "$wt_top_real" ]; then + SPAWN_WT_REASON="it is not inside a git worktree" + return 1 + fi + if [ "$wt_real" != "$wt_top_real" ]; then + SPAWN_WT_REASON="it is a subdirectory of worktree root '$wt_top_real', not a worktree root" + return 1 + fi + if [ "$wt_real" = "$PROJ_ABS_REAL" ]; then + SPAWN_WT_REASON="it is the spawning project itself" + return 1 + fi # The primary checkout uses the repository's common git dir as its own git # dir. A linked spawning home has a different top-level, but the same common # dir, so comparing only the two working directories cannot protect primary. - wt_git_dir=$(git -C "$WT" rev-parse --absolute-git-dir 2>/dev/null) \ + wt_git_dir=$(git -C "$path" rev-parse --absolute-git-dir 2>/dev/null) \ && wt_git_dir=$(cd "$wt_git_dir" 2>/dev/null && pwd -P) || wt_git_dir= proj_common=$(git -C "$PROJ_ABS" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) \ && proj_common=$(cd "$proj_common" 2>/dev/null && pwd -P) || proj_common= - if [ -z "$wt_real" ] || [ -z "$wt_top_real" ] || [ "$wt_real" != "$wt_top_real" ] \ - || [ "$wt_real" = "$proj_real" ] || [ -z "$wt_git_dir" ] || [ -z "$proj_common" ] \ - || [ "$wt_git_dir" = "$proj_common" ]; then - echo "error: $source did not yield an isolated worktree (resolved '$WT'; worktree root '${wt_top:-none}'; spawning project '$PROJ_ABS'); refusing to launch to avoid tangling the primary checkout. Inspect target $inspect_target" >&2 + if [ -z "$wt_git_dir" ] || [ -z "$proj_common" ]; then + SPAWN_WT_REASON="its git directory could not be resolved" + return 1 + fi + if [ "$wt_git_dir" = "$proj_common" ]; then + SPAWN_WT_REASON="it is the repository's primary checkout (its git dir is the spawning project's common git dir)" + return 1 + fi + return 0 +} + +validate_spawn_worktree() { # <source> <inspect-target> + local source=$1 inspect_target=$2 + if ! spawn_worktree_isolated "$WT"; then + echo "error: $source did not yield an isolated worktree (resolved '$WT'; worktree root '${SPAWN_WT_TOP:-none}'; spawning project '$PROJ_ABS'); refusing to launch to avoid tangling the primary checkout. Inspect target $inspect_target" >&2 exit 1 fi } @@ -2107,8 +2472,37 @@ EOF printf '%s' "$lines" >&2 } +spawn_worktree_has_origin_config() { # <worktree> + # Resolved remote.origin.* variables cover Git's effective include/includeIf chain; raw headers are also detected in the worktree config and any included file Git names through another variable. Git cannot enumerate a variable-less included file, so an empty origin section that is its only content remains indistinguishable from absence and intentionally proceeds rather than reimplementing Git's config parser. + local worktree=$1 config origin key seen=$'\n' + git -C "$worktree" config --get-regexp '^remote\.origin\.' >/dev/null 2>&1 && return 0 + while IFS=$'\t' read -r origin key; do + case $origin in file:*) config=${origin#file:} ;; *) continue ;; esac + [ -f "$config" ] || continue + case $seen in *$'\n'"$config"$'\n'*) continue ;; esac + seen+="$config"$'\n' + awk '/^[[:space:]]*\[[[:space:]]*[Rr][Ee][Mm][Oo][Tt][Ee][[:space:]]+"origin"[[:space:]]*\][[:space:]]*([#;].*)?$/ || /^[[:space:]]*\[[[:space:]]*[Rr][Ee][Mm][Oo][Tt][Ee]\.origin[[:space:]]*\][[:space:]]*([#;].*)?$/ { found=1 } END { exit !found }' "$config" && return 0 + done < <(git -C "$worktree" config --list --show-origin 2>/dev/null || true) + return 1 +} + freshen_spawn_worktree_base() { # <worktree> local worktree=$1 default target expected actual status + status=$(git -C "$worktree" -c core.quotePath=false status --porcelain) || { + echo "error: could not inspect pooled worktree '$worktree' before refreshing its base" >&2 + return 1 + } + if [ -n "$status" ]; then + if describe_stale_submodule_pins "$worktree" "$status"; then + echo "error: pooled worktree '$worktree' has a stale submodule checkout, not uncommitted work; refusing to launch and leaving it untouched" >&2 + else + echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 + fi + return 1 + fi + if ! spawn_worktree_has_origin_config "$worktree"; then + return 0 + fi if ! git -C "$worktree" fetch --quiet origin; then echo "error: could not fetch origin for pooled worktree '$worktree'; refusing to launch from a potentially stale base" >&2 return 1 @@ -2130,18 +2524,6 @@ freshen_spawn_worktree_base() { # <worktree> echo "error: '$target' is not a commit for pooled worktree '$worktree'; refusing to launch from a potentially stale base" >&2 return 1 } - status=$(git -C "$worktree" -c core.quotePath=false status --porcelain) || { - echo "error: could not inspect pooled worktree '$worktree' before refreshing its base" >&2 - return 1 - } - if [ -n "$status" ]; then - if describe_stale_submodule_pins "$worktree" "$status"; then - echo "error: pooled worktree '$worktree' has a stale submodule checkout, not uncommitted work; refusing to launch and leaving it untouched" >&2 - else - echo "error: pooled worktree '$worktree' is not clean; refusing to discard uncommitted work while refreshing its base" >&2 - fi - return 1 - fi if ! git -C "$worktree" reset --hard "$target" >/dev/null; then echo "error: could not reset pooled worktree '$worktree' to '$target'; refusing to launch from a potentially stale base" >&2 return 1 @@ -2633,6 +3015,88 @@ kimi_spawn_fail() { # <detail> echo "error: $1; inspect window $T" >&2 } +# rovo mirrors kimi's launch-then-send shape exactly: a positional brief is +# dead-on-arrival, so rovo launches bare and takes its brief pointer only after a +# readiness gate, then a delivery-confirmation gate. Both route their +# composer-emptiness half through the shared classifier (fm_backend_composer_state) +# like kimi. The banner and context-usage greps are launch-progress signals, not +# composer shapes. +rovo_capture() { + fm_backend_capture "$BACKEND" "$T" 120 "$W" 2>/dev/null || true +} + +rovo_composer_is_empty() { + [ "$(fm_backend_composer_state "$BACKEND" "$T" "$W" 2>/dev/null)" = empty ] +} + +rovo_wait_for_ready() { + local pane i=0 max=${FM_ROVO_READY_POLLS:-60} interval=${FM_ROVO_POLL_INTERVAL:-0.5} + while [ "$i" -lt "$max" ]; do + pane=$(rovo_capture) + # Lead with rovo's fresh-launch ASCII welcome banner (confirmed live), the + # same primary evidence kimi's own 'Welcome to Kimi Code!' match uses. The + # composer-empty fallback is WEAKER for rovo than for kimi: rovo's idle + # composer renders an inline placeholder chip (luminance ~163, above the + # ghost-strip threshold) that bin/fm-composer-lib.sh does not currently strip + # (see the deliberately-unfixed composer-ghost gap in rovo.md), so it can read + # non-empty - hence the banner is the primary signal. + if printf '%s\n' "$pane" | grep -Fq 'Welcome to Rovo!' \ + || rovo_composer_is_empty; then + return 0 + fi + i=$((i + 1)) + [ "$i" -ge "$max" ] || sleep "$interval" + done + return 1 +} + +rovo_delivery_is_confirmed() { # <plain-pane-capture> + local pane=$1 + rovo_composer_is_empty || return 1 + # rovo's real footer is `Context: <bar> N.N% NN.NK/NNNK` (e.g. + # "Context: ▎ 3.3% 30.1K/922K"). Confirm delivery when the sent pointer has + # scrolled into view OR the context-usage PERCENTAGE has advanced off zero. The + # regex tolerates the bar glyph and arbitrary spacing between the colon and the + # number (the [^%]* runs, unlike kimi's exact spacing) but is anchored to the + # digits BEFORE the % sign, so the always-nonzero total in the denominator + # (e.g. .../922K) can never masquerade as a nonzero usage percentage. + if printf '%s\n' "$pane" | grep -Fq 'Read the brief at' \ + || printf '%s\n' "$pane" | grep -qiE 'context:[^%]*[1-9][^%]*%'; then + return 0 + fi + return 1 +} + +rovo_wait_for_delivery() { + local pane i=0 max=${FM_ROVO_DELIVERY_POLLS:-40} interval=${FM_ROVO_POLL_INTERVAL:-0.5} + while [ "$i" -lt "$max" ]; do + pane=$(rovo_capture) + rovo_delivery_is_confirmed "$pane" && return 0 + i=$((i + 1)) + [ "$i" -ge "$max" ] || sleep "$interval" + done + return 1 +} + +rovo_spawn_fail() { # <detail> + printf 'failed: %s\n' "$1" >> "$STATE/$ID.status" + echo "error: $1; inspect window $T" >&2 + rovo_endpoint_cleanup +} + +# No task record is ever published on this failure path, so nothing else +# (teardown, the watcher) will ever learn this endpoint exists to close it: +# without this, the already-launched --yolo rovo process keeps running as an +# orphaned autonomous agent outside task control. Mirrors fm-teardown.sh's own +# generic non-orca kill call; orca's worktree+terminal are owned by the +# separate ORCA_ABORT_CLEANUP trap path and are out of scope here. +rovo_endpoint_cleanup() { + [ "$BACKEND" = orca ] && return 0 + local tab_id= + [ "$BACKEND" = zellij ] && tab_id=$ZELLIJ_TAB_ID + fm_backend_kill "$BACKEND" "$T" "$tab_id" "fm-$ID" 2>/dev/null || true +} + if [ "$RELAUNCH" -eq 1 ]; then # No worktree is acquired: the recorded one is reused as-is. What must be # proven instead is that the adopted endpoint's shell is actually sitting in @@ -2707,42 +3171,55 @@ elif [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ]; then # automatic-rename slips through), display-message -t <bad-name> falls back to the # active client's window, which would misread firstmate's OWN pane path as the # worktree and tangle a hook into the primary checkout. The window id never lies. - # Compare against PROJ_ABS_REAL (physical), not PROJ_ABS: a symlinked project - # prefix would otherwise make the pane's OS-level cwd read differ from - # PROJ_ABS on the very first poll, before the pane has actually moved. + # The project comparison is physical: spawn_worktree_isolated screens each + # read against PROJ_ABS_REAL, not PROJ_ABS, because a symlinked project prefix + # would otherwise make the pane's OS-level cwd read differ from PROJ_ABS on + # the very first poll, before the pane has actually moved. # - # A single read that already differs from PROJ_ABS_REAL is not proof the pane - # settled there: on some tmux/WSL setups a brand-new window's pane_current_path + # A single read that already looks isolated is not proof the pane settled + # there: on some tmux/WSL setups a brand-new window's pane_current_path # transiently reports an unrelated stale path (seen live as another real git # checkout entirely) before the shell catches up with treehouse get's cd. That - # stale path still passes the PROJ_ABS_REAL comparison and validate_spawn_worktree - # below (it resolves to a real, distinct worktree top-level too), so accepting it - # on one read alone silently records the wrong worktree= in state/<id>.meta. Require - # two consecutive reads to agree on the same non-project path before accepting it; - # a mismatch just becomes the new candidate rather than resetting the wait, so a - # pane that is already settled by the first real read only costs the one existing + # stale path passes spawn_worktree_isolated too (it resolves to a real, + # distinct worktree top-level), so accepting it on one read alone silently + # records the wrong worktree= in state/<id>.meta. Require two consecutive + # reads to agree on the same isolated path before accepting it; a mismatch + # just becomes the new candidate rather than resetting the wait, so a pane + # that is already settled by the first real read only costs the one existing # inter-poll sleep as confirmation, not a whole extra cycle on top. + # + # Every candidate is screened with the isolation guard's own predicate, so a + # read of the project itself or of the repository primary checkout is treated + # as the transient it is and the wait continues, instead of being adopted and + # then refused by the guard. + # A candidate the screen rejects is never adopted, so a host where the pane + # never reaches an isolated worktree spends the whole window before refusing. + # That wait is deliberate - telling a transient apart from a terminal + # misconfiguration would need machinery this path does not want - so the + # refusal has to be self-explaining instead: carry the last path seen and the + # reason it was rejected, and report both at the deadline. candidate="" + last_seen="" + last_reason="the pane reported no path" for _ in $(seq 1 60); do p=$(spawn_current_path "$WT_TARGET" || true) - if [ -n "$p" ]; then + [ -z "$p" ] || last_seen="$p" + if [ -n "$p" ] && spawn_worktree_isolated "$p"; then p_real=$(real_path_or_raw "$p") - if [ "$p_real" != "$PROJ_ABS_REAL" ]; then - if [ -n "$candidate" ] && [ "$p_real" = "$candidate" ]; then - WT="$p" - break - fi - candidate="$p_real" - else - candidate="" + last_reason="it is an isolated worktree, but no second read agreed with it" + if [ -n "$candidate" ] && [ "$p_real" = "$candidate" ]; then + WT="$p" + break fi + candidate="$p_real" else candidate="" + [ -z "$p" ] || last_reason=$SPAWN_WT_REASON fi sleep 1 done if [ -z "$WT" ]; then - echo "error: treehouse get did not enter a worktree within 60s; inspect window $T" >&2 + echo "error: treehouse get did not enter an isolated worktree within 60s (last seen '${last_seen:-none}': $last_reason; spawning project '$PROJ_ABS'); inspect window $T" >&2 exit 1 fi fi @@ -2819,10 +3296,11 @@ if [ "$KIND" != secondmate ] || [ "$HARNESS" = copilot ]; then # submission acknowledgement. The launch brief sent below IS a # submitted turn, so the seed record is busy/fm-spawn. The minted gen is # embedded into each adapter's wiring so an event from a superseded - # incarnation is rejected as stale. Grok stays on its isolated rendered-tail - # fallback and standalone Kimi stays unknown until fm_busy_kimi_verified - # opens, so neither is armed here. Gemini IS armed: its BeforeAgent / - # AfterAgent / SessionEnd hooks are a verified open-close pair. + # incarnation is rejected as stale. Grok and rovo stay on their isolated + # rendered-tail fallbacks and standalone Kimi stays unknown until + # fm_busy_kimi_verified opens, so none of the three is armed here. Gemini IS + # armed: its BeforeAgent / AfterAgent / SessionEnd hooks are a verified + # open-close pair. BUSY_GEN= case "$HARNESS" in codex*) @@ -2833,7 +3311,7 @@ if [ "$KIND" != secondmate ] || [ "$HARNESS" = copilot ]; then ;; esac case "$HARNESS" in - claude*|copilot*|opencode*|pi|pi-signed) + claude*|copilot*|opencode*|pi|pi-signed|omp) BUSY_GEN=$("$FM_ROOT/bin/fm-busy-event.sh" arm "$STATE_REAL" "$ID") || { echo "error: failed to arm the busy-state contract for $ID" >&2 exit 1 @@ -3020,6 +3498,43 @@ export default function (pi: any) { }); pi.on("turn_end", () => execFile("touch", ["$TURNEND"])); } +EOF + ;; + omp) + # Written OUTSIDE the worktree like Pi's, but for a different reason: omp + # has no trust gate, yet its cwd-only extension auto-discovery would load a + # worktree-resident copy a SECOND time next to the explicit -e (verified, + # omp 18.1.11). Lives in state/, cleaned by teardown. + cat > "$STATE/$ID.omp-ext.ts" <<EOF +// Firstmate semantic busy-state events + turn-end notification for omp (Oh My +// Pi); written by fm-spawn under the contract owned by bin/fm-busy-lib.sh. +// Semantic state: "agent_start" -> busy when a low-level agent run begins; +// "agent_end" -> idle only when event.willContinue is not true. omp has no +// agent_settled at all (verified, omp 18.1.2 and 18.1.11: zero occurrences in +// the binary); agent_end is its loop boundary and willContinue is the reliable +// "another loop is coming" flag, covering auto-retries, compaction retries, +// queued follow-ups, and a session_stop-forced continuation. ctx.isIdle() is +// deliberately NOT consulted: at a natural TUI agent_end it still reads false +// because session_stop is awaited before the session settles, so gating on it +// would leave every completed turn recorded busy. "turn_end" fires at every +// inner turn boundary and stays a wake NOTIFICATION touch for the watcher, +// never current-state truth. +import { execFile } from "node:child_process"; +const busyEvent = (state: string, event: string) => + new Promise<void>((resolve) => { + execFile("$FM_ROOT/bin/fm-busy-event.sh", [ + "apply", "$STATE_REAL", "$ID", state, + "--gen", "$BUSY_GEN", "--source", "omp-ext", "--event", event, + ], () => resolve()); + }); +export default function (pi: any) { + pi.on("agent_start", () => busyEvent("busy", "agent-start")); + pi.on("agent_end", (event: any) => { + if (event && event.willContinue === true) return; + return busyEvent("idle", "agent-end"); + }); + pi.on("turn_end", () => execFile("touch", ["$TURNEND"])); +} EOF ;; codex*) @@ -3289,6 +3804,44 @@ spawn_commit_backlog_transition() { fm_backlog_atomic_transition dispatch "$STATE/$ID.meta" "$DATA" "$ID" "$STATE" } +# The deferred-signal exit path's preservation report. A claim about preserved +# state is only trustworthy if that state is read back after the commit: the +# commit's own exit status has been observed to agree with a row that did not +# actually move (fm-yi4j evidence, 2026-09-05). This re-reads the paired record +# and the backlog row under the same per-task lock as the commit, repairs a row +# the commit believed it moved, and sets SPAWN_PRESERVED_CLAIM to exactly what +# was verified or attempted - never intent phrased as outcome. +spawn_report_preserved_state() { + local repair_error= + if ! fm_backlog_record_present "$STATE/$ID.meta" "task record" "$STATE"; then + SPAWN_PRESERVED_CLAIM="preservation could not be verified: its paired task record is missing; close out its backlog item by hand" + return 1 + fi + if ! fm_backlog_row_probe "$DATA" "$ID"; then + if [ "$FM_BACKLOG_ROW_RESULT" = not_found ]; then + SPAWN_PRESERVED_CLAIM="preservation could not be verified: its backlog item was not found; close out its paired task record by hand" + else + SPAWN_PRESERVED_CLAIM="preservation could not be verified: its backlog item state is unreadable (${FM_BACKLOG_ROW_ERROR:-no error recorded}); close out its paired task record and backlog item by hand" + fi + return 1 + fi + if [ "$FM_BACKLOG_ROW_STATE" = "in_flight no no" ]; then + SPAWN_PRESERVED_CLAIM="verified preserved: its paired task record is present and its backlog item is In flight" + return 0 + fi + # The commit reported success, but the row does not read back In flight: + # move it now under the same lock and verify the result before naming it. + fm_backlog_start "$DATA" "$ID" || repair_error=$FM_BACKLOG_TRANSITION_ERROR + if [ -z "$repair_error" ] \ + && fm_backlog_row_probe "$DATA" "$ID" \ + && [ "$FM_BACKLOG_ROW_STATE" = "in_flight no no" ]; then + SPAWN_PRESERVED_CLAIM="its backlog item did not read back In flight after the commit; it was moved to In flight now and verified, together with its paired task record" + return 0 + fi + SPAWN_PRESERVED_CLAIM="preservation could not be verified: its backlog item reads ${FM_BACKLOG_ROW_STATE:-unreadable}${repair_error:+, and moving it to In flight failed ($repair_error)}; close out its paired task record and backlog item by hand" + return 1 +} + if [ "$RELAUNCH" -eq 1 ]; then SPAWN_META_PUBLISH_STARTED=1 if ! fm_backlog_atomic_transition publish "$SPAWN_META_TMP" "$STATE/$ID.meta" "task record" "$STATE"; then @@ -3306,6 +3859,10 @@ TREEHOUSE_ABORT_RETURN=0 # still being delivered, cannot observe or complete a fresh provisional record # between its state check and `tasks-axi start`, and a delivery failure cannot # follow a committed In-flight transition. +if [ "$SPAWN_TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then + SPAWN_TREEHOUSE_PROJECT_LOCK_HELD=0 + fm_lock_release "$SPAWN_TREEHOUSE_PROJECT_LOCK" +fi if [ "$SPAWN_TASK_SET_LOCK_HELD" = 1 ]; then # The record is published, so this task is now part of the set a teardown # enumerates and locks per task. The set lock is only needed across that @@ -3321,27 +3878,39 @@ sq_turnend=$(shell_quote "$TURNEND") sq_piext=$(shell_quote "$STATE/$ID.pi-ext.ts") sq_piturnend=$(shell_quote "$PROJ_ABS/.pi/extensions/fm-primary-turnend-guard.ts") sq_piwatch=$(shell_quote "$PROJ_ABS/.pi/extensions/fm-primary-pi-watch.ts") +sq_ompext=$(shell_quote "$STATE/$ID.omp-ext.ts") +sq_ompcfg=$(shell_quote "${OMP_WORKER_CFG:-$FM_ROOT/.omp/fm-worker-overlay.yml}") sq_opinput=$(shell_quote "$FM_ROOT/bin/fm-operational-input.sh") sq_worktree=$(shell_quote "$WT") MODELFLAG=$(model_flag_for_harness "$HARNESS" "$MODEL") EFFORTFLAG=$(effort_flag_for_harness "$HARNESS" "$EFFORT") LAUNCH=${LAUNCH//__MODELFLAG__/$MODELFLAG} LAUNCH=${LAUNCH//__EFFORTFLAG__/$EFFORTFLAG} +if [ "$HARNESS" = rovo ]; then + ROVOCONFIGOVERRIDE=$(rovo_config_override_flag "$EFFORT" "$DATA" "$STATE" "$ID") || { + echo "error: could not resolve this task's home paths for rovo's allowedExternalPaths grant" >&2 + exit 1 + } + LAUNCH=${LAUNCH//__ROVOCONFIGOVERRIDE__/$ROVOCONFIGOVERRIDE} +fi LAUNCH=${LAUNCH//__BRIEF__/$sq_brief} LAUNCH=${LAUNCH//__TURNEND__/$sq_turnend} LAUNCH=${LAUNCH//__PIEXT__/$sq_piext} LAUNCH=${LAUNCH//__PITURNEND__/$sq_piturnend} LAUNCH=${LAUNCH//__PIWATCH__/$sq_piwatch} +LAUNCH=${LAUNCH//__OMPEXT__/$sq_ompext} +LAUNCH=${LAUNCH//__OMPWORKERCFG__/$sq_ompcfg} LAUNCH=${LAUNCH//__OPINPUT__/$sq_opinput} case "$HARNESS" in pi|pi-signed) LAUNCH=${LAUNCH//__PIBIN__/"$(shell_quote "$PI_BIN")"} ;; cursor) LAUNCH=${LAUNCH//__CURSORBIN__/"$(shell_quote "$CURSOR_BIN")"} ;; gemini) LAUNCH=${LAUNCH//__GEMINISETTINGS__/"$(shell_quote "$STATE_REAL/$ID.gemini-settings.json")"} ;; + omp) LAUNCH=${LAUNCH//__OMPBIN__/"$(shell_quote "$OMP_BIN")"} ;; copilot) LAUNCH=${LAUNCH//__COPILOTBIN__/"$(shell_quote "$COPILOT_BIN")"} ;; esac LAUNCH=${LAUNCH//__WORKTREE__/$sq_worktree} case "$HARNESS" in - claude|codex|opencode|pi|pi-signed|grok|kimi|gemini|muse) + claude|codex|opencode|pi|pi-signed|grok|kimi|gemini|muse|rovo|omp) LAUNCH="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI -u COPILOT_CLI -u COPILOT_AGENT_SESSION_ID -u COPILOT_LOADER_PID $LAUNCH" ;; cursor) @@ -3370,8 +3939,12 @@ if [ "$KIND" = secondmate ]; then # Keep this in step with fm_supervision_model (bin/fm-wake-lib.sh): Claude's # Stop auto-arm and Cursor's stop-hook park both run the watcher only BETWEEN # turns, so a fresh beacon with no live watcher is their healthy mid-turn state. + # Pi and pi-signed secondmates previously received persistent here and now + # receive extension to match fm_supervision_model's own table, so their pull + # guard tolerates the extension hand-off exactly as a Pi primary does. case "$HARNESS" in claude|copilot|cursor) supervision_model=autoarm ;; + pi|pi-signed|omp) supervision_model=extension ;; *) supervision_model=persistent ;; esac # Deliver the primary's EFFECTIVE trace-context decision as a normalized on/off @@ -3428,6 +4001,18 @@ if [ "$WINDOWS_HERDR_POWERSHELL" = 1 ]; then } fi spawn_send_text_line "$T" "$GOTMPDIR_COMMAND" +# Mark ship and scout panes so the runner refuses their primary checkout. +# Secondmates run their own homes rather than an assigned task worktree. +if [ "$KIND" = ship ] || [ "$KIND" = scout ]; then + TASK_ID_COMMAND="export FM_TASK_ID=$ID" + if [ "$WINDOWS_HERDR_POWERSHELL" = 1 ]; then + TASK_ID_COMMAND=$(fm_backend_herdr_windows_set_environment_command FM_TASK_ID "$ID") || { + echo "error: could not compose the native Windows FM_TASK_ID assignment for $ID" >&2 + exit 1 + } + fi + spawn_send_text_line "$T" "$TASK_ID_COMMAND" +fi # Send through the exact channel that already ships GOTMPDIR, so every backend # and harness - ship, scout, and secondmate - gets it before launch. Skipped # entirely when trace context is off. @@ -3453,6 +4038,26 @@ if [ -n "$SPAWN_TRACEPARENT" ]; then LAUNCH="unset TRACEPARENT; $LAUNCH" fi fi +if [ "$LAUNCH_ENV_ENABLED" = 1 ]; then + LAUNCH_ENV_PREFIX='/usr/bin/env -i' + for env_name in HOME PATH USER LOGNAME SHELL TERM COLORTERM LANG LC_ALL LC_CTYPE \ + TMPDIR TMP TEMP GOTMPDIR TMUX TMUX_PANE HERDR_ENV HERDR_SESSION HERDR_SOCKET_PATH \ + HERDR_PANE_ID CMUX_WORKSPACE_ID CMUX_SURFACE_ID CMUX_TAB_ID CMUX_PANEL_ID \ + CMUX_SOCKET_PATH ZELLIJ ZELLIJ_SESSION_NAME ZELLIJ_PANE_ID FM_ZELLIJ_SESSION \ + FM_TASK_ID \ + $LAUNCH_ENV_NAMES; do + # Only validated names enter shell syntax. Values expand once, quoted, in + # the pane shell and never become source text or spawn-process snapshots. + # shellcheck disable=SC2016 + printf -v env_arg '${%s+"%s=$%s"}' "$env_name" "$env_name" "$env_name" + LAUNCH_ENV_PREFIX="$LAUNCH_ENV_PREFIX $env_arg" + done + if [ -n "$SPAWN_TRACEPARENT" ]; then + # shellcheck disable=SC2016 + LAUNCH_ENV_PREFIX="$LAUNCH_ENV_PREFIX "'${TRACEPARENT+"TRACEPARENT=$TRACEPARENT"}' + fi + LAUNCH="$LAUNCH_ENV_PREFIX /bin/sh -c $(shell_quote "$LAUNCH")" +fi if [ "$WINDOWS_HERDR_POWERSHELL" = 1 ]; then LAUNCH_SCRIPT="$TASK_TMP/launch.sh" LAUNCH_SCRIPT_TMP="$LAUNCH_SCRIPT.tmp.${BASHPID:-$$}" @@ -3500,6 +4105,30 @@ if [ "$HARNESS" = kimi ]; then exit 1 fi fi +if [ "$HARNESS" = rovo ]; then + if ! rovo_wait_for_ready; then + rovo_spawn_fail "rovo did not show a verified ready signal before brief delivery in window $T" + exit 1 + fi + ROVO_POINTER="Read the brief at $BRIEF_REAL and follow it exactly." + ROVO_SUBMIT_RETRIES=${FM_ROVO_SUBMIT_RETRIES:-3} + ROVO_SUBMIT_SLEEP=${FM_ROVO_SUBMIT_SLEEP:-${FM_ROVO_POLL_INTERVAL:-0.5}} + ROVO_SUBMIT_SETTLE=${FM_ROVO_SUBMIT_SETTLE:-0} + if ! ROVO_SUBMIT_VERDICT=$(fm_backend_send_text_submit \ + "$BACKEND" "$T" "$ROVO_POINTER" "$ROVO_SUBMIT_RETRIES" \ + "$ROVO_SUBMIT_SLEEP" "$ROVO_SUBMIT_SETTLE" "$W"); then + rovo_spawn_fail "rovo brief pointer could not be submitted into window $T" + exit 1 + fi + if [ "$ROVO_SUBMIT_VERDICT" = send-failed ]; then + rovo_spawn_fail "rovo brief pointer could not be submitted into window $T" + exit 1 + fi + if ! rovo_wait_for_delivery; then + rovo_spawn_fail "rovo brief pointer delivery was not confirmed in window $T" + exit 1 + fi +fi if [ "$KIND" = secondmate ] && [ "${FM_SKIP_SECONDMATE_INHERIT:-0}" != 1 ]; then if ! fm_config_reread_discard_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then if fm_config_reread_quarantine_pending "$PROJ_ABS" "$ID" "$FM_HOME"; then @@ -3525,6 +4154,14 @@ if [ "$BACKLOG_TRANSITION" = 1 ]; then trap 'SPAWN_DEFERRED_SIGNAL=TERM' TERM fi SPAWN_BACKLOG_COMMIT_STATUS=0 +# Both the commit and its preservation read-back run under this task's meta +# lock, so an unresponsive tasks-axi there would hold the lock - and every +# lifecycle operation waiting on it - open ended, with even the deferred +# signals parked in a trap. Bound each invocation +# (bin/fm-backlog-transition-lib.sh's fm_tasks_axi): a timed-out call +# fails through the ordinary error plumbing, and the interrupted exit path +# reports it as the reason the preservation could not be verified. +FM_TASKS_AXI_TIMEOUT=${FM_TASKS_AXI_TIMEOUT:-30} if spawn_commit_backlog_transition; then SPAWN_FRESH_COMMIT_PENDING=0 else @@ -3549,17 +4186,24 @@ trap - HUP INT TERM if [ "$SPAWN_BACKLOG_COMMIT_STATUS" -ne 0 ]; then exit "$SPAWN_BACKLOG_COMMIT_STATUS" fi -fm_lock_release "$SPAWN_META_LOCK" -SPAWN_META_LOCK_HELD=0 if [ -n "$SPAWN_DEFERRED_SIGNAL" ]; then case "$SPAWN_DEFERRED_SIGNAL" in HUP) SPAWN_DEFERRED_SIGNAL_STATUS=129 ;; INT) SPAWN_DEFERRED_SIGNAL_STATUS=130 ;; TERM) SPAWN_DEFERRED_SIGNAL_STATUS=143 ;; esac - echo "error: spawn of $ID was interrupted after launch delivery began; its paired task record and In-flight backlog state were preserved" >&2 + # Keep deferring further signals so the read-back below cannot itself be + # killed halfway through verifying or correcting the preserved state. + trap 'SPAWN_DEFERRED_SIGNAL=$SPAWN_DEFERRED_SIGNAL' HUP INT TERM + # Deliberately unguarded against errexit: a failed verification still set + # the honest attempted-preservation claim the exit below reports. + spawn_report_preserved_state || true + trap - HUP INT TERM + echo "error: spawn of $ID was interrupted after launch delivery began; $SPAWN_PRESERVED_CLAIM" >&2 exit "$SPAWN_DEFERRED_SIGNAL_STATUS" fi +fm_lock_release "$SPAWN_META_LOCK" +SPAWN_META_LOCK_HELD=0 SPAWN_DELIVERY= [ -z "$MODE" ] || SPAWN_DELIVERY=" mode=$MODE yolo=$YOLO" diff --git a/bin/fm-startup-memory-budget-lib.sh b/bin/fm-startup-memory-budget-lib.sh index f2c06014b8e..033bb69ba42 100644 --- a/bin/fm-startup-memory-budget-lib.sh +++ b/bin/fm-startup-memory-budget-lib.sh @@ -23,7 +23,7 @@ fm_startup_memory_budget_fail() { fm_startup_memory_budget_link_count() { if [ "$(uname)" = Darwin ]; then - stat -f %l "$1" 2>/dev/null + /usr/bin/stat -f %l "$1" 2>/dev/null else stat -c %h "$1" 2>/dev/null fi diff --git a/bin/fm-supervise-daemon.sh b/bin/fm-supervise-daemon.sh index 91174bc5baf..a0145067528 100755 --- a/bin/fm-supervise-daemon.sh +++ b/bin/fm-supervise-daemon.sh @@ -236,7 +236,7 @@ _state_root() { printf '%s' "${FM_STATE_OVERRIDE:-$FM_HOME/state}"; } # --- portable stat (same trap as fm-watch.sh: no `stat -f || stat -c`) ------- if [ "$(uname)" = Darwin ]; then - _stat_file_mtime() { stat -f %m "$1" 2>/dev/null; } + _stat_file_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } else _stat_file_mtime() { stat -c %Y "$1" 2>/dev/null; } fi diff --git a/bin/fm-supervision-instructions.sh b/bin/fm-supervision-instructions.sh index a70fe89abb9..542d8159a3c 100755 --- a/bin/fm-supervision-instructions.sh +++ b/bin/fm-supervision-instructions.sh @@ -81,7 +81,7 @@ if [ -z "$HARNESS" ]; then fi case "$HARNESS" in - claude|codex|copilot|opencode|pi|grok|cursor) SNIPPET="$DOC_DIR/$HARNESS.md" ;; + claude|codex|copilot|opencode|pi|grok|cursor|omp) SNIPPET="$DOC_DIR/$HARNESS.md" ;; pi-signed) SNIPPET="$DOC_DIR/pi.md" ;; *) HARNESS=unknown; SNIPPET="$DOC_DIR/unknown.md" ;; esac @@ -90,6 +90,8 @@ esac checkpoint_seconds=${FM_CODEX_WATCH_CHECKPOINT:-180} pi_ext="$FM_ROOT/.pi/extensions/fm-primary-pi-watch.ts" pi_turnend_ext="$FM_ROOT/.pi/extensions/fm-primary-turnend-guard.ts" +omp_ext="$FM_ROOT/.omp/extensions/fm-primary-omp-watch.ts" +omp_turnend_ext="$FM_ROOT/.omp/extensions/fm-primary-turnend-guard.ts" x_mode_env="$CONFIG/x-mode.env" shell_quote() { @@ -109,6 +111,8 @@ render_snippet() { while IFS= read -r line || [ -n "$line" ]; do line=${line//__FM_PI_EXT__/$pi_ext} line=${line//__FM_PI_TURNEND_EXT__/$pi_turnend_ext} + line=${line//__FM_OMP_EXT__/$omp_ext} + line=${line//__FM_OMP_TURNEND_EXT__/$omp_turnend_ext} line=${line//__FM_X_MODE_ENV_SH__/$x_mode_env_sh} line=${line//__FM_X_MODE_ENV__/$x_mode_env} printf '%s\n' "$line" @@ -143,6 +147,9 @@ repair_line() { pi|pi-signed) printf '%s%s%s%s%s%s\n' "$prefix" 'repair a missing or failed watcher cycle with the Pi tool fm_watch_arm_pi, or restart Pi with -e ' "$pi_turnend_ext" ' -e ' "$pi_ext" ' if the extensions are not loaded.' ;; + omp) + printf '%s%s%s%s%s%s\n' "$prefix" 'repair a missing or failed watcher cycle with the omp tool fm_watch_arm_omp, or restart omp inside this home so ' "$omp_turnend_ext" ' and ' "$omp_ext" ' auto-load from .omp/extensions/ (use -e with both paths only when starting omp from another directory).' + ;; opencode) printf '%s%s\n' "$prefix" 'repair missing watcher supervision by letting the OpenCode TUI plugin arm after idle; use bin/fm-watch-arm.sh only as a manual recovery probe if the plugin reports failure.' ;; @@ -172,6 +179,9 @@ ordinary_wake_line() { pi|pi-signed) printf '%s\n' '- Ordinary wake: the Pi extension already owns watcher continuity; do not arm another cycle.' ;; + omp) + printf '%s\n' '- Ordinary wake: the omp extension already owns watcher continuity; do not arm another cycle.' + ;; opencode) printf '%s\n' '- Ordinary wake: the OpenCode TUI plugin already owns watcher continuity; do not arm manually.' ;; diff --git a/bin/fm-supervision-lib.sh b/bin/fm-supervision-lib.sh index 3bbb13bdf8d..1bbc5708834 100644 --- a/bin/fm-supervision-lib.sh +++ b/bin/fm-supervision-lib.sh @@ -2,10 +2,10 @@ # Shared "supervision missing" predicate. # Usage: . bin/fm-supervision-lib.sh # -# Reports whether a firstmate home needs supervision because it has in-flight -# work (a state/<id>.meta exists) or an X-mode relay poll -# (state/x-watch.check.sh), and whether its watcher has a fresh liveness beacon -# (state/.last-watcher-beat, touched every poll cycle, within the grace window). +# Reports whether a firstmate home needs supervision (fm_supervision_status +# below is the single owner of that condition set), and whether its watcher has +# a fresh liveness beacon (state/.last-watcher-beat, touched every poll cycle, +# within the grace window). # bin/fm-turnend-guard.sh uses the PID-strict fm_watcher_healthy from # bin/fm-wake-lib.sh for its block decision. bin/fm-guard.sh uses the model-aware # fm_watcher_supervision_verdict (also in bin/fm-wake-lib.sh), which owns what a @@ -15,7 +15,7 @@ # Portable mtime; Linux stat lacks -f, macOS stat lacks -c. fm_sup_stat_mtime() { if [ "$(uname)" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi @@ -25,16 +25,27 @@ fm_sup_stat_mtime() { # Populates, for the state dir at $1: # FM_SUP_IN_FLIGHT count of state/*.meta (in-flight tasks) # FM_SUP_SOURCES count of registered process-to-event sources -# FM_SUP_NEEDED true/false - in-flight work, an X-mode relay poll, or a +# FM_SUP_CHECKS count of registered custom checks: a state/<id>.check.sh +# with the state/<id>.check-trust binding that +# bin/fm-check-register.sh writes. Task PR polls carry no +# such binding and are torn down with their task, and the +# relay shim keeps its own trust path, so neither counts +# here. Presence of the binding is the whole test: whether +# those bytes are still the registered ones is the check +# sweep's call at execution time, and a home whose check +# no longer validates needs the watcher precisely so the +# sweep can report the rejection instead of going quiet. +# FM_SUP_NEEDED true/false - in-flight work, an X-mode relay poll, a # registered event source (a source is a wait on an -# external process, not a task, so it has no metadata) +# external process, not a task, so it has no metadata), +# or a registered custom check # FM_SUP_WATCHER_FRESH true/false - a watcher beacon within the grace window # FM_SUP_BEACON_DESC human-readable beacon age, for banners ("never" if absent) # FM_SUP_QUEUE_PENDING true/false - state/.wake-queue has unread records # grace-seconds defaults to $FM_GUARD_GRACE, then 300, matching fm-guard.sh. # Always returns 0; callers read the vars, or use fm_supervision_unhealthy below. fm_supervision_status() { - local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} meta source beat m age + local state=$1 grace=${2:-${FM_GUARD_GRACE:-300}} meta source check id beat m age FM_SUP_IN_FLIGHT=0 FM_SUP_NEEDED=false FM_SUP_WATCHER_FRESH=false @@ -50,9 +61,21 @@ fm_supervision_status() { [ -e "$source" ] || continue FM_SUP_SOURCES=$((FM_SUP_SOURCES + 1)) done + FM_SUP_CHECKS=0 + for check in "$state"/*.check.sh; do + [ -e "$check" ] || continue + id=${check##*/} + id=${id%.check.sh} + if [ "$id" = x-watch ]; then + continue + fi + [ -e "$state/$id.check-trust" ] || continue + FM_SUP_CHECKS=$((FM_SUP_CHECKS + 1)) + done if [ "$FM_SUP_IN_FLIGHT" -gt 0 ] \ || [ -f "$state/x-watch.check.sh" ] \ - || [ "$FM_SUP_SOURCES" -gt 0 ]; then + || [ "$FM_SUP_SOURCES" -gt 0 ] \ + || [ "$FM_SUP_CHECKS" -gt 0 ]; then FM_SUP_NEEDED=true fi diff --git a/bin/fm-task-inbox-lib.sh b/bin/fm-task-inbox-lib.sh index 445ed95a499..6a0287ab78f 100644 --- a/bin/fm-task-inbox-lib.sh +++ b/bin/fm-task-inbox-lib.sh @@ -13,12 +13,14 @@ # # Design (captain-adopted, data/fm-send-reliability-reframe-s1/report.md): the # payload moves to the filesystem, which is reliable; the terminal carries only -# a short constant doorbell line, which does not need to be reliable because -# ringing it again is free. A duplicated doorbell is a no-op by construction -# (the worker finds the inbox empty or already handled), a swallowed doorbell -# is detected by the absence of the worker's acknowledgement and re-rung on a -# bounded schedule, and a worker that never acknowledges surfaces through the -# ordinary stale wake into stuck-crewmate-recovery. +# a short constant doorbell line. While the endpoint remains available, that +# line does not need to be reliable because ringing it again is free. A +# duplicated doorbell is a no-op by construction (the worker finds the inbox +# empty or already handled), and a swallowed doorbell is detected by the +# absence of the worker's acknowledgement and re-rung on a bounded schedule. +# A positively dead or missing endpoint bypasses that schedule without being +# typed into, and its unhandled record surfaces through the ordinary stale wake +# into stuck-crewmate-recovery. # # Layout under <state-dir>: # <task>.inbox/NNN.msg one durable steer, numeric sequence, atomic rename @@ -46,14 +48,18 @@ # FM_TASK_INBOX_GRACE_SECS is due one delivery attempt per grace period; an # attempt may ring or be skipped to protect proven pending composer text. After # FM_TASK_INBOX_RING_MAX attempts without an acknowledgement it escalates. The -# caller owns the busy check (a busy pane just waits - the record is durable and -# the worker reaches a turn boundary) and the wake emission; this library owns -# only the schedule. If attempt bookkeeping cannot be persisted while the record -# remains unhandled, the caller surfaces that failure instead of retrying -# silently; a concurrently removed inbox is a quiet no-op. Escalation -# deliberately queues the wake before writing the -# deduplication marker: normal polls surface a message once, while a crash or -# marker failure may produce a rare duplicate rather than silently lose a wake. +# caller owns the busy and recovery-grade endpoint checks: a busy pane waits, +# while a positively dead or missing endpoint skips delivery and the ladder and +# escalates directly. This library owns only the schedule and escalation marker. +# If attempt bookkeeping cannot be persisted while the record remains unhandled, +# the caller surfaces that failure instead of retrying silently; a concurrently +# removed inbox is a quiet no-op. Escalation deliberately queues the wake before +# writing the deduplication marker: normal polls surface a message once, while a +# crash or marker failure may produce a rare duplicate rather than silently lose +# a wake. +# +# Inbox paths containing bytes outside printable ASCII are unsupported. The +# doorbell refuses them rather than sending terminal control bytes to a pane. # # fm_task_inbox_ring requires bin/fm-backend.sh's dispatch (sourced below); the # other helpers are dependency-light. Sourced by bin/fm-send.sh, bin/fm-watch.sh, @@ -246,19 +252,30 @@ fm_task_inbox_body() { # <record-path> # The constant self-describing doorbell line for the inbox containing a record. # Self-describing on purpose: a worker whose brief predates the inbox contract -# still receives the complete instruction in the line itself. +# still receives the complete instruction in the line itself. The leading `: ` +# is the POSIX shell no-op, so the same line typed into a pane whose agent has +# exited (a bare shell) runs nothing; see the dead-pane note in the header. +# A non-printable path fails without output so terminal controls never reach +# the pane's line discipline. fm_task_inbox_doorbell_line() { # <record-path> - local dir=${1%/*} abs + local dir=${1%/*} abs quoted LC_ALL=C abs=$(cd "$dir" 2>/dev/null && pwd) || abs=$dir - printf 'Firstmate instruction waiting: list %s/*.msg and, in numeric order, read and act on each, then mv each handled file to %s/handled/.' \ - "$abs" "$abs" + case "$abs" in + *[![:print:]]*) return 1 ;; + esac + quoted=$(printf '%s' "$abs" | sed "s/'/'\\\\''/g") + printf ": Firstmate instruction waiting: list '%s'/*.msg and, in numeric order, read and act on each, then mv each handled file to '%s'/handled/." \ + "$quoted" "$quoted" } -# Ring the doorbell, best-effort: one advisory composer pre-check, then the -# backend's submit machinery with a minimal retry budget, verdict discarded. +# Ring the doorbell, best-effort: one endpoint-liveness pre-check, one advisory +# composer pre-check, then the backend's submit machinery with a minimal retry +# budget, verdict discarded. # Returns 0 rang, 1 skipped because the composer PROVENLY holds pending text -# (the watcher re-rings later), 2 the backend send failed. No return value is -# delivery proof; the acknowledgement move is the only delivery signal. +# (the watcher re-rings later), 2 the backend send failed, 3 skipped because +# the endpoint is positively dead or missing (nothing typed; recovery owns the +# record). No return value is delivery proof; the acknowledgement move is the +# only delivery signal. # The skip is deliberately narrow: only an exact `pending` verdict defers, # because there our Enter could submit someone's real half-typed content. # `pending-unproven` and `unknown` still ring - the worst outcome is a garbled @@ -267,11 +284,20 @@ fm_task_inbox_doorbell_line() { # <record-path> # positively identify (that classifier is advisory here by design). fm_task_inbox_ring() { # <backend> <target> <record-path> [expected-label] local backend=$1 target=$2 rec=$3 label=${4:-} line cstate verdict - line=$(fm_task_inbox_doorbell_line "$rec") + case "$(fm_backend_agent_state "$backend" "$target" 2>/dev/null || true)" in + dead|missing) return 3 ;; + esac + if ! line=$(fm_task_inbox_doorbell_line "$rec"); then + return 2 + fi cstate=$(fm_backend_composer_state "$backend" "$target" "$label" 2>/dev/null) || cstate=unknown case "$cstate" in pending) return 1 ;; esac + # Accepted residual race: terminal input and Enter are separate delivery + # steps, so an agent exiting after the liveness check could leave a bare + # shell only a suffix; the `: ` prefix protects complete lines only. Do not + # add process-bound atomic delivery here unless an incident reopens this. if ! verdict=$(fm_backend_send_text_submit "$backend" "$target" "$line" 1 0.4 0.3 "$label" 2>/dev/null); then return 2 fi @@ -338,8 +364,11 @@ fm_task_inbox_due_action() { # <state-dir> <task-id> IFS=$(printf '\t') read -r rec_base count last <<EOF $ladder EOF - if [ "$rec_base" != "$base" ]; then - # A different (or first) oldest message: the previous ladder is stale. + if [ -n "$rec_base" ] && [ "$rec_base" != "$base" ]; then + # A different oldest message: the previous ladder is stale. An absent + # ladder is left alone so a dead-pane escalation, which never rings and so + # never writes one, keeps its marker (the marker check below still ignores + # a marker naming some other message). count=0 last=0 rm -f "$dir/.escalated" 2>/dev/null || true @@ -364,10 +393,12 @@ EOF } # Advance the ladder after a delivery attempt. A failed ring or a composer- -# protected skip still consumes budget so neither a dead pane nor permanently -# blocked composer can retry silently forever. A concurrently removed inbox is -# a successful no-op; otherwise failure means the caller must surface the -# unwritable ladder while the record remains unhandled. +# protected skip still consumes budget so neither an unreadable pane nor a +# permanently blocked composer can retry silently forever. A positively dead or +# missing endpoint never enters the ladder: the watcher escalates it directly. +# A concurrently removed inbox is a successful no-op; otherwise failure means +# the caller must surface the unwritable ladder while the record remains +# unhandled. fm_task_inbox_record_ring() { # <state-dir> <task-id> <record-path> local dir base ladder rec_base count last dir=$(fm_task_inbox_dir "$1" "$2") diff --git a/bin/fm-teardown.sh b/bin/fm-teardown.sh index 9331f3f1018..03e0737ab1f 100755 --- a/bin/fm-teardown.sh +++ b/bin/fm-teardown.sh @@ -22,6 +22,9 @@ # The close - and only the close - is replaced by `tasks-axi reopen` with the # deliverable recorded while the backlog item is still an open captain call # (bin/fm-captain-hold.sh `open` owns that predicate), because the policy holds +# NOTE: this uses `open`'s silent default and depends only on its unchanged +# 0/1/2 exit-code contract. The optional `--identity` output that bin/fm-watch.sh +# asks for prints only on an exit 0 and changes nothing read here. # the very work item a question gates and cleanup must never retire the # captain's own question. The same pending-close record carries that intent as # `mode=retain`, so an interrupted cleanup replays the retention rather than a @@ -38,6 +41,13 @@ # already present in the up-to-date default branch. This recognizes the common # squash-merge-then-delete-branch flow, where the branch's own commits live nowhere # on a remote yet the change is fully in main. +# Squash merges collapse the branch's commits, so per-commit patch ids against main +# no longer match, and a pipeline rebase can leave the local worktree diverged from +# the PR head. A diverged copy is not treated as landed: path-set coverage, git +# cherry, and merge-tree containment each fail to prove content landed without also +# accepting unlanded edits to the same paths. Teardown still accepts a merged PR +# whose head contains the current local work (ancestor or equivalent patch ids), +# or a clean content-in-default tree match. Anything else refuses. # The PR itself is resolved from the task's recorded pr= when present, or - when # no pr= was ever recorded (e.g. a yolo-authorized merge on a repo with no PR CI, # where the usual "checks green" fm-pr-check.sh trigger never fires) - by looking @@ -59,6 +69,32 @@ # task state when that proof fails; otherwise it removes the task's check, # trust record, PR sidecar, and publication record with the rest of the # volatile state. +# Worktree-slot ownership (teardown-slot-collision): a treehouse pool slot is +# reused across tasks, so a stale, duplicated, or drifted worktree= record can +# name a slot a DIFFERENT live task now holds. Cleanup kills every process under +# that path and hard-resets it before returning it, so releasing a slot that is +# not genuinely this task's destroys another worker's live work. Before the first +# cleanup step, teardown verifies record exclusivity: no OTHER task record in +# this home or any locally registered Firstmate home may name the same live path +# in its worktree= or home=. One live path with two task records is the reuse +# collision itself, whichever record is stale. The recorded endpoint's exact +# task identity and the record's spawn incarnation are validated separately +# before cleanup. Its current working directory is only incidental process +# state: the same worker remains the owner after changing directory, so cwd can +# never veto teardown of that exact recorded endpoint. +# The scan and destructive return hold a project-identity lock in the local root +# Firstmate home's state directory, as resolved by bin/fm-wake-lib.sh's +# fm_firstmate_root_home; a home seeded from another machine is its own local +# root, since a lock on this filesystem cannot be held or observed across that +# boundary. Fresh Treehouse spawns for that project in +# every local Firstmate home hold the same lock from before slot allocation +# through metadata publication, closing the publication +# gap; forced secondmate teardown takes it and runs the same checks for every +# descendant Treehouse slot before touching any child. +# This refusal is not relaxed by --force: --force authorizes discarding THIS +# task's unlanded work, never another task's live work. Reconcile whichever +# record is wrong and re-run. Orca is not a pool slot and proves its path through +# require_orca_worktree_path_match instead. # Orca tasks use the same safety checks, then close the recorded terminal and # remove the recorded worktree through `orca worktree rm`; teardown never guesses # an Orca target from ambient CLI state. @@ -87,10 +123,20 @@ # releases its durable treehouse lease so the pool slot is freed, # never left leased forever. If the treehouse return fails, teardown leaves the # leased home and state in place instead of hiding a still-held lease. -# Usage: fm-teardown.sh <task-id> [--force] +# Usage: fm-teardown.sh <task-id> [--force] [--legacy-record] # --force skips ordinary-task dirty and landed-work checks, skips scout report # checks, and discards secondmate child work for kind=secondmate. Only use it # when the captain has explicitly said to discard the work. +# --legacy-record accepts a task record that predates the spawn_gen field: +# teardown then proceeds only when the recorded endpoint is confirmed dead or +# agent-less (bin/fm-backend.sh's recovery-grade classifier), and without +# --force the worktree still passes the ordinary landed-work checks. The +# accepted legacy incarnation is stamped into the record before its close is +# recorded and named in the teardown line; the flag never relaxes the +# unlanded-work refusal, which --force alone can authorize. A legacy- stamp +# an abandoned attempt left behind never counts as a published incarnation: +# the record still reads as a legacy record, so the endpoint gate runs again +# and the retry still needs --legacy-record. # # Transient / stale worktree git lock recovery (teardown-lock-race): a crew process # killed mid-git-operation can leave a .git/worktrees/<wt>/index.lock (or, for a @@ -218,7 +264,20 @@ if [ "$#" -lt 1 ] || ! fm_task_id_path_safe "$1"; then exit 2 fi ID=$1 -FORCE=${2:-} +FORCE= +LEGACY_RECORD_GIVEN=0 +shift +while [ "$#" -gt 0 ]; do + case "$1" in + --force) FORCE=--force ;; + --legacy-record) LEGACY_RECORD_GIVEN=1 ;; + *) + echo "error: invalid teardown request" >&2 + exit 2 + ;; + esac + shift +done fm_backlog_directory_present "$STATE" "state directory" || { echo "error: teardown refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 exit 1 @@ -239,6 +298,50 @@ if [ "$FORCE" = --force ] && [ "$(fm_lease_actor)" = branch ]; then exit "$FM_LEASE_REFUSE_EXIT" fi fm_lease_guard "$ID" "teardown (fm-teardown)" + +# A Treehouse slot has the managed pool's fixed <pool>/<slot>/<repo> layout. +# Require both its pool state and the same Git common directory as the recorded +# project; an ordinary linked worktree is not evidence that Treehouse owns it. +is_treehouse_pool_slot() { # <project> <worktree> + local project=$1 worktree=$2 slot pool state project_common slot_common + [ -d "$project" ] && [ -d "$worktree" ] || return 1 + slot=$(CDPATH='' cd -- "$worktree" 2>/dev/null && pwd -P) || return 1 + pool=$(dirname "$(dirname "$slot")") + state="$pool/treehouse-state.json" + [ -f "$state" ] && [ ! -L "$state" ] || return 1 + project_common=$(git -C "$project" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + slot_common=$(git -C "$slot" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) || return 1 + project_common=$(CDPATH='' cd -- "$project_common" 2>/dev/null && pwd -P) || return 1 + slot_common=$(CDPATH='' cd -- "$slot_common" 2>/dev/null && pwd -P) || return 1 + [ "$project_common" = "$slot_common" ] +} + +META="$STATE/$ID.meta" +TREEHOUSE_PROJECT_LOCK= +TREEHOUSE_PROJECT_LOCK_HELD=0 +TREEHOUSE_SLOT_LOCK_REQUIRED=0 +if [ -f "$META" ] && [ ! -L "$META" ]; then + TEARDOWN_LOCK_KIND=$(fm_meta_get "$META" kind) + [ -n "$TEARDOWN_LOCK_KIND" ] || TEARDOWN_LOCK_KIND=ship + TEARDOWN_LOCK_BACKEND=$(fm_meta_get "$META" backend) + [ -n "$TEARDOWN_LOCK_BACKEND" ] || TEARDOWN_LOCK_BACKEND=tmux + TEARDOWN_LOCK_WT=$(fm_meta_get "$META" worktree) + TEARDOWN_LOCK_PROJECT=$(fm_meta_get "$META" project) + if [ "$TEARDOWN_LOCK_KIND" != secondmate ] \ + && [ "$TEARDOWN_LOCK_BACKEND" != orca ] \ + && is_treehouse_pool_slot "$TEARDOWN_LOCK_PROJECT" "$TEARDOWN_LOCK_WT"; then + TREEHOUSE_SLOT_LOCK_REQUIRED=1 + TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$TEARDOWN_LOCK_PROJECT") || { + echo "REFUSED: cannot resolve the shared Treehouse project lock for ${TEARDOWN_LOCK_PROJECT:-<missing>}; nothing was changed" >&2 + exit 1 + } + fm_lock_try_acquire "$TREEHOUSE_PROJECT_LOCK" || { + echo "REFUSED: another Treehouse slot allocation or return is in progress for $TEARDOWN_LOCK_PROJECT; nothing was changed" >&2 + exit 1 + } + TREEHOUSE_PROJECT_LOCK_HELD=1 + fi +fi CONTROL_LOCK="$STATE/.control-$ID.lock" CONTROL_LOCK_HELD=0 META_LOCK= @@ -248,6 +351,7 @@ DESCENDANT_TASK_STATES=() DESCENDANT_TASK_IDS=() DESCENDANT_TASK_KINDS=() DESCENDANT_TASK_HOMES=() +DESCENDANT_TREEHOUSE_LOCK_PATHS=() teardown_release_locks() { local status=$? i if declare -F teardown_release_herdr_locks >/dev/null 2>&1; then @@ -277,6 +381,10 @@ teardown_release_locks() { fm_lock_release "$CONTROL_LOCK" || true CONTROL_LOCK_HELD=0 fi + if [ "$TREEHOUSE_PROJECT_LOCK_HELD" = 1 ]; then + fm_lock_release "$TREEHOUSE_PROJECT_LOCK" || true + TREEHOUSE_PROJECT_LOCK_HELD=0 + fi fm_lease_guard_release || true return "$status" } @@ -291,7 +399,6 @@ CONTROL_LOCK_HELD=1 fm_refuse_if_gate_agent FM_LOCK_LOG_PREFIX=teardown -META="$STATE/$ID.meta" fm_backlog_record_present "$META" "task record" "$STATE" || { echo "error: teardown refused: $FM_BACKLOG_TRANSITION_ERROR" >&2 exit 1 @@ -307,6 +414,11 @@ TEARDOWN_META_KIND=$(fm_meta_get "$META" kind) [ -n "$TEARDOWN_META_KIND" ] || TEARDOWN_META_KIND=ship TEARDOWN_CLEANUP_RECOVERY=$(fm_meta_get "$META" cleanup_recovery) TEARDOWN_META_SPAWN_GEN= +TEARDOWN_LEGACY_PENDING=0 +TEARDOWN_LEGACY_ACCEPTED=0 +TEARDOWN_LEGACY_ENDPOINT= +TEARDOWN_LEGACY_RETAINED_STAMP= +TEARDOWN_LEGACY_PRESTAMP_SIZE=0 TEARDOWN_BACKLOG_APPLIES=0 TEARDOWN_BACKLOG_SKIP_REASON= if [ "$TEARDOWN_CLEANUP_RECOVERY" != orca ]; then @@ -323,10 +435,39 @@ if [ "$TEARDOWN_CLEANUP_RECOVERY" != orca ]; then fi if [ "$TEARDOWN_BACKLOG_APPLIES" = 1 ]; then if ! fm_backlog_meta_spawn_gen "$META" "$STATE"; then - echo "error: task $ID's record has no spawn_gen that identifies one exact incarnation ($FM_BACKLOG_TRANSITION_ERROR); refusing automatic teardown - relaunch the task to publish an unambiguous incarnation, then retry teardown" >&2 - exit 1 + TEARDOWN_LEGACY_GEN_COUNT=$(LC_ALL=C awk -F= '$1 == "spawn_gen" { count++ } END { print count + 0 }' "$META" 2>/dev/null || printf '0\n') + if [ "$TEARDOWN_LEGACY_GEN_COUNT" = 0 ] && [ "$LEGACY_RECORD_GIVEN" = 1 ]; then + # A record that predates the incarnation field: acceptance is gated later, + # once the recorded endpoint is known, so its state can be confirmed dead + # or agent-less before any cleanup decision is made. + TEARDOWN_LEGACY_PENDING=1 + elif [ "$TEARDOWN_LEGACY_GEN_COUNT" = 0 ]; then + echo "error: task $ID's record has no spawn_gen that identifies one exact incarnation ($FM_BACKLOG_TRANSITION_ERROR); refusing automatic teardown - relaunch the task to publish an unambiguous incarnation, then retry teardown, or pass --legacy-record once its recorded endpoint is confirmed dead or agent-less" >&2 + exit 1 + else + echo "error: task $ID's record has an unreadable spawn_gen that identifies one exact incarnation ($FM_BACKLOG_TRANSITION_ERROR); refusing automatic teardown - fix the record, then retry teardown" >&2 + exit 1 + fi + else + case "$FM_BACKLOG_META_SPAWN_GEN" in + legacy-*) + # Only this teardown path mints a legacy- token; a launch publishes + # s<epoch>.<pid>.<random>. So one still on a retained record is the + # stamp an abandoned --legacy-record attempt could not roll back, not + # an incarnation any spawn ever published. The record is still the + # legacy record it was, and is treated as one: the dead-or-agent-less + # endpoint gate runs again on the retry instead of being skipped by + # the abandoned attempt's own stamp. + if [ "$LEGACY_RECORD_GIVEN" != 1 ]; then + echo "error: task $ID's record carries the legacy incarnation stamp $FM_BACKLOG_META_SPAWN_GEN left by an abandoned --legacy-record teardown, not an incarnation published by a spawn; refusing automatic teardown - relaunch the task to publish an unambiguous incarnation, then retry teardown, or pass --legacy-record once its recorded endpoint is confirmed dead or agent-less" >&2 + exit 1 + fi + TEARDOWN_LEGACY_PENDING=1 + TEARDOWN_LEGACY_RETAINED_STAMP=$FM_BACKLOG_META_SPAWN_GEN + ;; + esac fi - TEARDOWN_META_SPAWN_GEN=$FM_BACKLOG_META_SPAWN_GEN + [ "$TEARDOWN_LEGACY_PENDING" = 1 ] || TEARDOWN_META_SPAWN_GEN=$FM_BACKLOG_META_SPAWN_GEN fi # Cleanup never closes a captain call (see the header). Asked here, before any # destructive step, so "cannot tell" can refuse while everything is intact. @@ -668,7 +809,6 @@ remote_secondmate_teardown() { route_home=$SECONDMATE_REGISTRY_HOME [ "$route_host" = "$remote_host" ] && [ "$route_root" = "$remote_root" ] && [ "$route_home" = "$remote_home" ] \ || { echo "REFUSED: remote secondmate metadata does not match its registry route" >&2; return 1; } - [ -z "$FORCE" ] || [ "$FORCE" = --force ] || { echo "error: invalid teardown option: $FORCE" >&2; return 2; } handoff_wake_retire_validate || return 1 remote_recovery_paths_validate initial || return 1 if [ "$FORCE" != --force ] && [ "$REMOTE_OUTBOX_PRESENT" -eq 1 ]; then @@ -786,8 +926,51 @@ ORCA_PATH_MATCH_VERIFIED=0 CLEANUP_RECOVERY=$TEARDOWN_CLEANUP_RECOVERY KIND=$TEARDOWN_META_KIND +EXPECTED_TREEHOUSE_PROJECT_LOCK= +if [ "$KIND" != secondmate ] && [ "$BACKEND" != orca ] \ + && is_treehouse_pool_slot "$PROJ" "$WT"; then + EXPECTED_TREEHOUSE_PROJECT_LOCK=$(fm_treehouse_project_lock_path "$PROJ") || { + echo "REFUSED: cannot resolve the shared Treehouse project lock for ${PROJ:-<missing>}; nothing was changed" >&2 + exit 1 + } + if [ "$TREEHOUSE_PROJECT_LOCK_HELD" != 1 ] \ + || [ "$TREEHOUSE_PROJECT_LOCK" != "$EXPECTED_TREEHOUSE_PROJECT_LOCK" ]; then + echo "REFUSED: task $ID's Treehouse project identity changed while teardown acquired its locks; nothing was changed" >&2 + exit 1 + fi +elif [ "$TREEHOUSE_SLOT_LOCK_REQUIRED" = 1 ]; then + echo "REFUSED: task $ID stopped naming a live Treehouse slot while teardown acquired its locks; nothing was changed" >&2 + exit 1 +fi MODE=$(grep '^mode=' "$META" | cut -d= -f2- || true) [ -n "$MODE" ] || MODE=no-mistakes + +# A record accepted as a legacy incarnation (no spawn_gen, --legacy-record +# given) may be torn down only when its recorded endpoint is confidently gone +# or agent-less; only the recovery-grade classifier's dead and missing license +# that, and every ambiguous, unreadable, or unverified endpoint state refuses +# while the record is still intact. Acceptance resolves the incarnation token +# here; the record itself is stamped only once every landed-work refusal has +# passed, immediately before the close marker binds to it, so any refusal +# leaves the record byte-identical. +if [ "$TEARDOWN_LEGACY_PENDING" = 1 ]; then + TEARDOWN_LEGACY_ENDPOINT=$(fm_backend_agent_state "$BACKEND" "$T") + case "$TEARDOWN_LEGACY_ENDPOINT" in + dead|missing) ;; + *) + echo "REFUSED: task $ID's record predates spawn_gen and its recorded endpoint reads '$TEARDOWN_LEGACY_ENDPOINT', not confidently dead or agent-less; --legacy-record teardown is refused while an agent may still be bound to it. Nothing was changed." >&2 + echo "Reconcile the endpoint first (bin/fm-crew-state.sh $ID), or relaunch the task to publish an unambiguous incarnation, then retry teardown." >&2 + exit 1 + ;; + esac + if [ -n "$TEARDOWN_LEGACY_RETAINED_STAMP" ]; then + TEARDOWN_META_SPAWN_GEN=$TEARDOWN_LEGACY_RETAINED_STAMP + else + TEARDOWN_META_SPAWN_GEN="legacy-$(date -u +%Y%m%dT%H%M%SZ)-$$" + fi + TEARDOWN_LEGACY_ACCEPTED=1 +fi + PUBLIC_FOLLOWUP_HOME=$FM_HOME PUBLIC_FOLLOWUP_STATE=$STATE PUBLIC_FOLLOWUP_WORK_HOME=main @@ -1224,10 +1407,12 @@ backlog_done_args() { # invariant). This prints what already happened, so the follow-up wording stays # only where a human still owes the edit. backlog_refresh_reminder() { - local backlog_display + local backlog_display root [ "$KIND" = secondmate ] && return 0 [ "$CLEANUP_RECOVERY" = orca ] && return 0 - if backlog_display=$(fm_backlog_file "$DATA"); then + if root=$(fm_backlog_root "$DATA") && [ "$(fm_tasks_axi_backend "$root")" != markdown ]; then + backlog_display="this home's configured tasks-axi backend (data directory $DATA)" + elif backlog_display=$(fm_backlog_file "$DATA"); then : else backlog_display="${DATA%/}/backlog.md" @@ -1891,6 +2076,92 @@ require_orca_worktree_path_match_if_present() { require_orca_worktree_path_match "$worktree_id" "$inspected" } +# The task's own live slot, canonicalized, or empty when this record has no slot +# to release (a secondmate home, a record with no worktree=, or a path that is +# already gone). Every slot-ownership check below is scoped to that value, so a +# record with nothing live to return skips them rather than refusing. +teardown_live_slot_path() { + [ "$KIND" != secondmate ] || return 1 + is_treehouse_pool_slot "$PROJ" "$WT" || return 1 + canonical_existing_dir "$WT" +} + +collect_local_firstmate_states() { + local record_state=$1 root home reg line child known existing i=0 + local -a homes + TREEHOUSE_OWNER_STATES=("$record_state") + root=$(fm_firstmate_root_home "$FM_HOME") || { + echo "REFUSED: cannot resolve the root Firstmate home; nothing was changed" >&2 + return 1 + } + homes=("$root") + while [ "$i" -lt "${#homes[@]}" ]; do + home=${homes[$i]} + i=$((i + 1)) + known=0 + for existing in "${TREEHOUSE_OWNER_STATES[@]}"; do + [ "$existing" != "$home/state" ] || known=1 + done + [ "$known" = 1 ] || TREEHOUSE_OWNER_STATES+=("$home/state") + reg="$home/data/secondmates.md" + [ ! -e "$reg" ] && [ ! -L "$reg" ] && continue + [ -f "$reg" ] && [ ! -L "$reg" ] || { + echo "REFUSED: local Firstmate registry is unsafe at $reg; nothing was changed" >&2 + return 1 + } + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + "- "*) + secondmate_registry_parse_line "$line" || { + echo "REFUSED: malformed local Firstmate registry entry in $reg; nothing was changed" >&2 + return 1 + } + [ "$SECONDMATE_REGISTRY_REMOTE" -eq 0 ] || continue + child=$(canonical_existing_dir "$SECONDMATE_REGISTRY_HOME") || { + echo "REFUSED: registered local Firstmate home is unavailable: $SECONDMATE_REGISTRY_HOME; nothing was changed" >&2 + return 1 + } + known=0 + for existing in "${homes[@]}"; do + [ "$existing" != "$child" ] || known=1 + done + [ "$known" = 1 ] || homes+=("$child") + ;; + esac + done < "$reg" + done +} + +require_exclusive_worktree_slot_record() { + local record_meta=$1 record_id=$2 record_state=$3 worktree=$4 + local slot state_dir other other_id field other_path other_slot + slot=$(canonical_existing_dir "$worktree") || return 0 + collect_local_firstmate_states "$record_state" || return 1 + for state_dir in "${TREEHOUSE_OWNER_STATES[@]}"; do + for other in "$state_dir"/*.meta; do + [ -f "$other" ] && [ ! -L "$other" ] || continue + [ "$other" != "$record_meta" ] || continue + other_id=$(basename "$other" .meta) + for field in worktree home; do + other_path=$(fm_meta_get "$other" "$field") + [ -n "$other_path" ] || continue + other_slot=$(canonical_existing_dir "$other_path") || continue + [ "$other_slot" = "$slot" ] || continue + echo "REFUSED: task $record_id's recorded worktree $slot is also task $other_id's recorded $field." >&2 + echo "Returning that pool slot would kill $other_id's processes and reset its copy, so nothing was changed - not even with --force." >&2 + echo "Reconcile whichever record is wrong (bin/fm-crew-state.sh $record_id; bin/fm-crew-state.sh $other_id), then re-run teardown." >&2 + return 1 + done + done + done +} + +require_exclusive_task_worktree_slot() { + local slot + slot=$(teardown_live_slot_path) || return 0 + require_exclusive_worktree_slot_record "$META" "$ID" "$STATE" "$slot" +} + firstmate_home_has_treehouse_slot() { local home=$1 worktree_registered_for_project "$FM_ROOT" "$home" @@ -2307,6 +2578,7 @@ preflight_descendant_task_locks() { DESCENDANT_TASK_IDS=() DESCENDANT_TASK_KINDS=() DESCENDANT_TASK_HOMES=() + DESCENDANT_TREEHOUSE_LOCK_PATHS=() collect_descendant_task_locks "$home" || return 1 # Acquisition order, which every other holder of these locks must match so # they cannot cycle: each home's task-set lock first (parent home before child @@ -2356,6 +2628,61 @@ preflight_descendant_task_locks() { done } +preflight_descendant_treehouse_slots() { + local i state task_id meta kind backend target worktree project lock_path held + for ((i=0; i < ${#DESCENDANT_TASK_IDS[@]}; i++)); do + state=${DESCENDANT_TASK_STATES[$i]} + task_id=${DESCENDANT_TASK_IDS[$i]} + meta="$state/$task_id.meta" + kind=$(meta_value "$meta" kind) + [ -n "$kind" ] || kind=ship + backend=$(fm_backend_of_meta "$meta") + worktree=$(meta_value "$meta" worktree) + project=$(meta_value "$meta" project) + if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then + continue + fi + if ! is_treehouse_pool_slot "$project" "$worktree"; then + continue + fi + lock_path=$(fm_treehouse_project_lock_path "$project") || { + echo "REFUSED: cannot resolve the shared Treehouse project lock for child $task_id; forced teardown changed nothing" >&2 + return 1 + } + held=0 + [ "$TREEHOUSE_PROJECT_LOCK_HELD" != 1 ] || [ "$TREEHOUSE_PROJECT_LOCK" != "$lock_path" ] || held=1 + for target in "${DESCENDANT_TREEHOUSE_LOCK_PATHS[@]}"; do + [ "$target" != "$lock_path" ] || held=1 + done + if [ "$held" = 0 ]; then + fm_lock_try_acquire "$lock_path" || { + echo "REFUSED: another Treehouse slot allocation or return is in progress for child $task_id; forced teardown changed nothing" >&2 + return 1 + } + DESCENDANT_TREEHOUSE_LOCK_PATHS+=("$lock_path") + DESCENDANT_LOCK_PATHS+=("$lock_path") + fi + done + for ((i=0; i < ${#DESCENDANT_TASK_IDS[@]}; i++)); do + state=${DESCENDANT_TASK_STATES[$i]} + task_id=${DESCENDANT_TASK_IDS[$i]} + meta="$state/$task_id.meta" + kind=$(meta_value "$meta" kind) + [ -n "$kind" ] || kind=ship + backend=$(fm_backend_of_meta "$meta") + worktree=$(meta_value "$meta" worktree) + project=$(meta_value "$meta" project) + if [ "$kind" = secondmate ] || [ "$backend" = orca ]; then + continue + fi + if ! is_treehouse_pool_slot "$project" "$worktree"; then + continue + fi + fm_backend_validate_task_endpoint "$meta" "$task_id" || return 1 + require_exclusive_worktree_slot_record "$meta" "$task_id" "$state" "$worktree" || return 1 + done +} + validate_firstmate_home_children_removal() { local home=$1 sub_state child_meta child_id child_wt child_proj child_kind child_home child_backend child_orca_worktree_id sub_state="$home/state" @@ -2648,7 +2975,7 @@ cleanup_firstmate_home_children() { status_retire_presentation_task "$sub_state" "$child_id" || return 1 fm_backlog_atomic_transition remove "$sub_state/$child_id.meta" "task record" "$sub_state" || return 1 rm -f "$sub_state/$child_id.turn-ended" \ - "$sub_state/$child_id.pi-ext.ts" \ + "$sub_state/$child_id.pi-ext.ts" "$sub_state/$child_id.omp-ext.ts" \ "$sub_state/$child_id.grok-turnend-token" "$sub_state/$child_id.kimi-turnend-token" \ "$sub_state/$child_id.muse-session" "$sub_state/$child_id.muse-session-current" \ "$sub_state/$child_id.cursor-session" "$sub_state/$child_id.reconcile-nudged" \ @@ -2671,6 +2998,8 @@ remove_secondmate_registry_entry() { return "$rc" } +require_exclusive_task_worktree_slot || exit 1 + validate_pr_poll_cleanup "$STATE" "$ID" || exit 1 if [ "$KIND" = secondmate ]; then @@ -2686,6 +3015,7 @@ if [ "$KIND" = secondmate ]; then validate_firstmate_home_children_removal "$HOME_PATH" || exit 1 preflight_descendant_task_locks "$HOME_PATH" || exit 1 validate_firstmate_home_children_removal "$HOME_PATH" || exit 1 + preflight_descendant_treehouse_slots || exit 1 if [ "$BACKEND" = herdr ]; then teardown_herdr_preflight_target "$T" "$ID" || exit 1 fi @@ -2967,11 +3297,55 @@ fi # A projected Herdr close can still defer action-free above, so record the # transition only after that boundary. Captain-held rows carry the same retain # flag upstream uses for ordinary teardown. +teardown_legacy_stamp_rollback() { + [ "$TEARDOWN_LEGACY_PRESTAMP_SIZE" -gt 0 ] 2>/dev/null || return 1 + perl -e 'truncate($ARGV[0], $ARGV[1]) or exit 1' -- \ + "$META" "$TEARDOWN_LEGACY_PRESTAMP_SIZE" || return 1 + [ "$(wc -c < "$META" | tr -d ' ')" = "$TEARDOWN_LEGACY_PRESTAMP_SIZE" ] +} + if [ "$BACKLOG_CLOSED" = 1 ]; then - fm_backlog_close_marker_write "$STATE" "$ID" "$DATA" "$META_SPAWN_GEN" \ - "${BACKLOG_TRANSITION_FLAGS[@]+"${BACKLOG_TRANSITION_FLAGS[@]}"}" \ - "${BACKLOG_DONE_ARGS[@]+"${BACKLOG_DONE_ARGS[@]}"}" \ - || { echo "error: the pending backlog $BACKLOG_TRANSITION for $ID could not be recorded ($FM_BACKLOG_TRANSITION_ERROR); retaining every durable task record" >&2; exit 1; } + # Stamp only after a projected close can no longer defer; a failed marker + # publication must restore the exact legacy record before a retry. + if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ] && [ -z "$TEARDOWN_LEGACY_RETAINED_STAMP" ]; then + TEARDOWN_LEGACY_PRESTAMP_SIZE=$(wc -c < "$META" | tr -d ' ') + TEARDOWN_LEGACY_STAMP_FAILED= + if [ -s "$META" ] && [ -n "$(tail -c 1 -- "$META" 2>/dev/null)" ]; then + printf '\n' >> "$META" || TEARDOWN_LEGACY_STAMP_FAILED=newline + fi + if [ -z "$TEARDOWN_LEGACY_STAMP_FAILED" ]; then + printf 'spawn_gen=%s\n' "$TEARDOWN_META_SPAWN_GEN" >> "$META" \ + || TEARDOWN_LEGACY_STAMP_FAILED=append + fi + if [ -z "$TEARDOWN_LEGACY_STAMP_FAILED" ] \ + && ! fm_backlog_meta_spawn_gen "$META" "$STATE"; then + TEARDOWN_LEGACY_STAMP_FAILED=validate + fi + if [ -n "$TEARDOWN_LEGACY_STAMP_FAILED" ]; then + teardown_legacy_stamp_rollback \ + || echo "error: the legacy incarnation stamp on $ID's record could not be rolled back; re-run teardown with --legacy-record after reconciling its endpoint" >&2 + if [ "$TEARDOWN_LEGACY_STAMP_FAILED" = validate ]; then + echo "error: the stamped legacy incarnation does not validate for $ID ($FM_BACKLOG_TRANSITION_ERROR); refusing destructive teardown" >&2 + else + echo "error: could not stamp the accepted legacy incarnation into task $ID's record; refusing destructive teardown" >&2 + fi + exit 1 + fi + fi + if ! fm_backlog_close_marker_write "$STATE" "$ID" "$DATA" "$META_SPAWN_GEN" \ + "${BACKLOG_TRANSITION_FLAGS[@]+"${BACKLOG_TRANSITION_FLAGS[@]}"}" \ + "${BACKLOG_DONE_ARGS[@]+"${BACKLOG_DONE_ARGS[@]}"}"; then + if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ] && [ -z "$TEARDOWN_LEGACY_RETAINED_STAMP" ] \ + && teardown_legacy_stamp_rollback; then + echo "error: the pending backlog $BACKLOG_TRANSITION for $ID could not be recorded ($FM_BACKLOG_TRANSITION_ERROR); the accepted legacy incarnation was rolled back, retaining every durable task record" >&2 + else + echo "error: the pending backlog $BACKLOG_TRANSITION for $ID could not be recorded ($FM_BACKLOG_TRANSITION_ERROR); retaining every durable task record" >&2 + if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ] && [ -z "$TEARDOWN_LEGACY_RETAINED_STAMP" ]; then + echo "error: the legacy incarnation stamp on $ID's record could not be rolled back; re-run teardown with --legacy-record after reconciling its endpoint" >&2 + fi + fi + exit 1 + fi fi # Best-effort: drop the local task branch so the shared repo does not accumulate refs. @@ -3093,7 +3467,7 @@ remove_pr_poll_artifacts "$STATE" "$ID" || exit 1 retire_busy_state "$STATE" "$ID" "$BUSY_GEN" || exit 1 status_retire_presentation_task "$STATE" "$ID" || exit 1 rm -f "$STATE/$ID.turn-ended" \ - "$STATE/$ID.pi-ext.ts" "$STATE/$ID.grok-turnend-token" \ + "$STATE/$ID.pi-ext.ts" "$STATE/$ID.omp-ext.ts" "$STATE/$ID.grok-turnend-token" \ "$STATE/$ID.kimi-turnend-token" "$STATE/$ID.muse-session" \ "$STATE/$ID.muse-session-current" "$STATE/$ID.cursor-session" \ "$STATE/$ID.copilot-prompt-submitted" \ @@ -3146,5 +3520,9 @@ fi if [ -d "$STATE" ]; then "$SCRIPT_DIR/fm-home-summary-refresh.sh" --best-effort || true fi -echo "teardown $ID complete (window $T, worktree $WT)" +if [ "$TEARDOWN_LEGACY_ACCEPTED" = 1 ]; then + echo "teardown $ID complete (window $T, worktree $WT, legacy record accepted without spawn_gen: endpoint $TEARDOWN_LEGACY_ENDPOINT, incarnation $TEARDOWN_META_SPAWN_GEN)" +else + echo "teardown $ID complete (window $T, worktree $WT)" +fi backlog_refresh_reminder diff --git a/bin/fm-test-isolation-proof.sh b/bin/fm-test-isolation-proof.sh index e9ecd53d32d..4ed6175ae38 100755 --- a/bin/fm-test-isolation-proof.sh +++ b/bin/fm-test-isolation-proof.sh @@ -216,8 +216,8 @@ EOF dir_mode() { local path=$1 - if stat -f %Lp "$path" >/dev/null 2>&1; then - stat -f %Lp "$path" + if /usr/bin/stat -f %Lp "$path" >/dev/null 2>&1; then + /usr/bin/stat -f %Lp "$path" else stat -c %a "$path" fi diff --git a/bin/fm-test-run.sh b/bin/fm-test-run.sh index e567ee6f88d..8dbd1516798 100755 --- a/bin/fm-test-run.sh +++ b/bin/fm-test-run.sh @@ -28,7 +28,11 @@ # fm-test-run.sh --aggregate-json <out.json> <lane.json> [more lane.json...] # # Options: -# --json <path> write a deterministic timing artifact after the run +# --json <path> write a deterministic timing artifact after the run. Each +# script record carries its family, expected gate-skip class, +# exit, duration, whether it gate-skipped, and the reason it +# gave (empty when it ran), so a lane can say which harness or +# tool this host could not exercise. # --list print selected script paths (one per line) and exit 0 # --list-scheduled # print selected paths longest-hint-first and exit 0 @@ -86,11 +90,26 @@ # FM_TEST_SLOWEST rank=<k> script=<path> duration_ms=<n> # FM_TEST_BUDGET max_wall_ms=<n> duration_ms=<n> (only with --max-wall-ms) # +# Placement refusal: +# A task worker is assigned an isolated worktree, and that placement is +# checked only when its task starts. When FM_TASK_ID marks such a worker and +# this runner resolves to the repository's PRIMARY checkout, every executing +# mode refuses before selecting a suite: the suite creates and switches +# branches, and the primary is the checkout every linked worktree resolves +# against. Inspection modes execute nothing and stay available, and a run with +# no FM_TASK_ID set is unchanged. +# # Exit status is non-zero if any selected script exits non-zero, a configured # --fail-on-gate-skip token appears, the measured duration exceeds # --max-wall-ms, timing-artifact finalization fails, or a concurrent worker # violates its isolation check. Other gate skips (first meaningful line -# matching ^skip:) remain successful and are counted as skipped_gate. +# matching ^skip:) remain successful and are counted as skipped_gate; each one +# is logged with its reason and recorded in the timing artifact. +# +# expected_gate_skip classes name why a family is allowed to skip: herdr (the +# pinned real-Herdr lane), optional-binary (a backend whose binary is optional), +# live-capability (a live-harness guard governed by fm_live_gate, which records +# unavailable tools and explicit policy skips; see tests/lib.sh), or none. # # Family labels, the changed-file map, and production portable-shard composition # live in this script only (one owner). The proven-isolated candidate set remains @@ -102,7 +121,7 @@ # share a machine. This script owns <n>: a lane whose <n> disagrees with the # configured shard count is refused, so a CI matrix cannot silently drop a shard. # --changed is conservative: it over-selects related families rather than -# under-selecting, and never expands to the complete suite unless --all. The one +# under-selecting; their union can cover the complete inventory. The one # place it is deliberately narrow is a bin/ path with no curated family: a test # that names it is selected as that SCRIPT, because the reference is per-script # evidence. Consumer bin/ scripts still resolve through the curated map, so @@ -239,12 +258,12 @@ worker_directory_private() { local dir=$1 mode [ -d "$dir" ] && [ ! -L "$dir" ] || return 1 case "$(uname -s 2>/dev/null)" in - MSYS*|MINGW*|CYGWIN*) + MSYS*|MINGW*) native_windows_private_directory_valid "$dir" return ;; esac - mode=$(stat -c %a "$dir" 2>/dev/null || stat -f %Lp "$dir" 2>/dev/null) || return 1 + mode=$(stat -c %a "$dir" 2>/dev/null || /usr/bin/stat -f %Lp "$dir" 2>/dev/null) || return 1 case "$mode" in 700|0700) return 0 ;; *) return 1 ;; @@ -255,6 +274,29 @@ now_iso() { date -u +%Y-%m-%dT%H:%M:%SZ } +# Enforce the placement refusal described in this script's header. +# +# The primary checkout is the working tree whose own git dir IS the repository's +# common git dir; every linked worktree has a git dir under it instead. That is +# the same predicate bin/fm-spawn.sh uses to keep a launch out of the primary, +# and unlike comparing top-level paths it still holds when the primary is +# reached through a different path. When git resolves neither directory - a +# non-repository fixture, a detached copy - nothing proves this is the primary, +# so the run proceeds. +refuse_primary_checkout_for_task() { + local task_id git_dir common_dir top + task_id=${FM_TASK_ID:-} + [ -n "$task_id" ] || return 0 + git_dir=$(git -C "$ROOT" rev-parse --absolute-git-dir 2>/dev/null) \ + && git_dir=$(cd "$git_dir" 2>/dev/null && pwd -P) || git_dir= + common_dir=$(git -C "$ROOT" rev-parse --path-format=absolute --git-common-dir 2>/dev/null) \ + && common_dir=$(cd "$common_dir" 2>/dev/null && pwd -P) || common_dir= + [ -n "$git_dir" ] && [ -n "$common_dir" ] || return 0 + [ "$git_dir" = "$common_dir" ] || return 0 + top=$(cd "$ROOT" && pwd -P) + die "refusing to run in the repository primary checkout $top while FM_TASK_ID=$task_id is set; run from the assigned task worktree instead" +} + cpu_count() { local n n=$(getconf _NPROCESSORS_ONLN 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1) @@ -290,9 +332,9 @@ family_for_basename() { fm-classify-decision-key.test.sh|\ fm-composer-ghost.test.sh|fm-composer-lib.test.sh|\ fm-crew-state.test.sh|fm-captain-hold-lifecycle.test.sh|\ - fm-copilot-harness.test.sh|fm-documentation-audiences.test.sh|\ + fm-copilot-harness.test.sh|fm-reconcile-validation.test.sh|fm-documentation-audiences.test.sh|\ fm-ensure-agents-md.test.sh|fm-grok-harness.test.sh|\ - fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ + fm-kimi-harness.test.sh|fm-muse-harness.test.sh|fm-rovo-harness.test.sh|fm-omp-harness.test.sh|fm-herdr-lab.test.sh|fm-lint.test.sh|\ fm-lint-workflows.test.sh|\ fm-operational-input.test.sh|fm-pi-primary-types.test.sh|\ fm-harness-adapter-references.test.sh|\ @@ -344,6 +386,7 @@ family_for_basename() { printf '%s\n' session-bootstrap ;; fm-afk-pi-herdr-return-e2e.test.sh|\ + fm-bearings-board-lavish-live-e2e.test.sh|\ fm-claude-stop-autoarm-live-e2e.test.sh|\ fm-cmux-claude-composer-live-e2e.test.sh|\ fm-composer-matrix-live-e2e.test.sh|\ @@ -352,11 +395,11 @@ family_for_basename() { fm-cursor-primary-live-e2e.test.sh|\ fm-grok-stop-live-e2e.test.sh|fm-harness-adapter-instructions-live-e2e.test.sh|\ fm-harness-liveness-drift-live-e2e.test.sh|\ - fm-muse-signals-live-e2e.test.sh|\ + fm-muse-signals-live-e2e.test.sh|fm-rovo-signals-live-e2e.test.sh|\ fm-herdr-version-floor-live-e2e.test.sh|\ fm-opencode-primary-live-e2e.test.sh|fm-pi-branch-live-e2e.test.sh|\ fm-pi-branch-responsiveness-live-e2e.test.sh|\ - fm-pi-primary-live-e2e.test.sh|\ + fm-pi-primary-live-e2e.test.sh|fm-omp-primary-live-e2e.test.sh|\ fm-sessionstart-hook-live-e2e.test.sh|fm-sessionstart-instruction-refresh-live-e2e.test.sh|\ fm-quota-array-dispatch-live-e2e.test.sh|fm-send-secondmate-marker-herdr-e2e.test.sh|\ fm-send-inbox-doorbell-live-e2e.test.sh|\ @@ -400,6 +443,7 @@ family_for_basename() { fm-no-mistakes-required.test.sh|fm-peek-remote.test.sh|\ fm-pending-reply.test.sh|fm-pi-branch-extension.test.sh|\ fm-procevent-quota.test.sh|fm-procevent-when.test.sh|fm-procevent.test.sh|\ + fm-live-gate.test.sh|\ fm-project-origin.test.sh|fm-public-followup.test.sh|fm-quota-choose.test.sh|\ fm-remote-entrypoint.test.sh|fm-remote-secondmate-parent-binding.test.sh|\ fm-send-remote-delivery.test.sh|fm-spawn-pool-base-freshen.test.sh|\ @@ -417,7 +461,7 @@ family_for_basename() { expected_gate_skip_for_family() { case "$1" in real-herdr-gated) printf '%s\n' herdr ;; - live-harness-optin) printf '%s\n' optin-env ;; + live-harness-optin) printf '%s\n' live-capability ;; cmux|zellij|orca) printf '%s\n' optional-binary ;; snapshot-bearings) printf '%s\n' optional-binary ;; *) printf '%s\n' none ;; @@ -628,8 +672,8 @@ is_proven_isolated_script() { # The portable serial remainder: every tests/*.test.sh that is neither # proven-isolated nor real-herdr-gated. Watcher, lock, AFK, real tmux, daemon, -# secondmate lifecycle, bootstrap, live-harness opt-in, GUI-backend, and other -# unproven work stays here. Derived rather than enumerated so a newly added test +# secondmate lifecycle, bootstrap, the live-harness-optin family, GUI-backend, +# and other unproven work stays here. Derived rather than enumerated so a newly added test # lands here by default instead of falling out of every lane. list_portable_serial() { local s base fam @@ -712,6 +756,7 @@ tests/fm-home-summary-refresh.test.sh 34793 tests/fm-inactive-reconcile.test.sh 41826 tests/fm-kimi-harness.test.sh 18015 tests/fm-lint-workflows.test.sh 855 +tests/fm-live-gate.test.sh 6000 tests/fm-muse-harness.test.sh 55572 tests/fm-muse-signals-live-e2e.test.sh 23 tests/fm-on.test.sh 11692 @@ -724,10 +769,12 @@ tests/fm-pi-branch-live-e2e.test.sh 56 tests/fm-pi-branch-responsiveness-live-e2e.test.sh 21 tests/fm-pi-primary-live-e2e.test.sh 20 tests/fm-pi-watch-extension.test.sh 42970 +tests/fm-pi-windows-shell-invocation.test.sh 5121 tests/fm-pr-check-security.test.sh 160475 tests/fm-procevent-quota.test.sh 1949 tests/fm-procevent-when.test.sh 17392 tests/fm-procevent.test.sh 69715 +tests/fm-reconcile-validation.test.sh 27000 tests/fm-project-origin.test.sh 137 tests/fm-public-followup.test.sh 196745 tests/fm-quota-array-dispatch-live-e2e.test.sh 21 @@ -1536,6 +1583,9 @@ families_for_unmapped_bin() { families_for_changed_path() { local path=$1 fixture_ref case "$path" in + skills/reconcile-firstmate-upstream/scripts/*|tests/fm-reconcile-validation.test.mjs) + printf '%s\n' __script__:fm-reconcile-validation.test.sh + ;; tests/fm-backend-herdr-eventwait.test.py) printf '%s\n' real-herdr-gated printf '%s\n' backend-dispatch @@ -1638,6 +1688,7 @@ families_for_changed_path() { .pi/extensions/lib/fm-operational-input.ts) # The same rule for the operational-input library, whose reach is wider: # every Pi extension that classifies or encodes operational text. + printf '%s\n' __script__:fm-pi-windows-shell-invocation.test.sh printf '%s\n' __script__:fm-pi-branch-extension.test.sh printf '%s\n' __script__:fm-pi-watch-extension.test.sh printf '%s\n' __script__:fm-calm-pi-extension.test.sh @@ -1651,6 +1702,7 @@ families_for_changed_path() { .pi/extensions/fm-primary-turnend-guard.ts) # The run tier's two harness-supplied facts (source vocabulary and # context-reset stdout injection) only show up against a real harness. + printf '%s\n' __script__:fm-pi-windows-shell-invocation.test.sh printf '%s\n' session-bootstrap printf '%s\n' live-harness-optin ;; @@ -1755,7 +1807,11 @@ families_for_changed_path() { docs/configuration.md|docs/supervision-protocols/*) printf '%s\n' pure-contract-unit ;; - .gitattributes|.gitignore) + .gitattributes) + printf '%s\n' pure-contract-unit + printf '%s\n' __script__:fm-gitignore-config.test.sh + ;; + .gitignore) printf '%s\n' __script__:fm-gitignore-config.test.sh ;; .opencode/plugins/*|.pi/extensions/*) @@ -1793,8 +1849,15 @@ families_for_changed_path() { README.md|LICENSE|assets/*|docs/*) ;; *) - families_for_test_reference "$path" \ - || emit_unmapped_changed_path "$path" + if [ -e "$path" ]; then + families_for_test_reference "$path" \ + || emit_unmapped_changed_path "$path" + else + # A retired source path with no remaining test consumer cannot select + # a runnable suite. Known source paths above retain their mappings, + # and a still-referenced removal is found by the same reference scan. + families_for_test_reference "$path" || true + fi ;; esac } @@ -1911,6 +1974,17 @@ detect_gate_skip() { esac } +# Echo the reason a gate skip gave, i.e. the first meaningful output line with +# its leading "skip:" removed. Tabs and stray whitespace are folded so the +# reason stays one field of the tab-separated record the JSON artifact is built +# from. Callers only use this once detect_gate_skip has already said yes. +gate_skip_reason() { + local file=$1 first + first=$(awk 'NF { print; exit }' "$file" 2>/dev/null || true) + first=${first#skip:} + printf '%s\n' "$first" | tr '\t' ' ' | sed -e 's/^ *//' -e 's/ *$//' +} + # True when any output line contains "skip: <token>" (token may contain spaces). detect_gate_skip_token() { local file=$1 token=$2 @@ -1964,7 +2038,7 @@ with open(records_file, encoding="utf-8") as fh: line = line.rstrip("\n") if not line: continue - path, family, expected, exit_s, dur_s, gate = line.split("\t") + path, family, expected, exit_s, dur_s, gate, reason = line.split("\t") scripts.append({ "path": path, "family": family, @@ -1972,6 +2046,7 @@ with open(records_file, encoding="utf-8") as fh: "duration_ms": int(dur_s), "exit": int(exit_s), "gate_skip": gate == "true", + "gate_skip_reason": reason, }) families = [] @@ -2232,6 +2307,16 @@ case "$PER_SCRIPT_TIMEOUT_SECS" in ''|*[!0-9]*) die "--per-script-timeout-secs requires a whole number of seconds (0 disables)" ;; esac +# Refuse before any suite is selected or run. The inspection modes execute +# nothing: --list-families, --list-concurrent-safe-families, --list-lanes, +# --check-coverage, --concurrent-safe-family-jobs-max and --aggregate-json have +# already exited above, and --list/--list-scheduled print their selection and +# exit below. An unset MODE still falls through to the usage error, so a caller +# who named no selection mode is told that rather than this. +if [ -n "${MODE:-}" ] && [ "$LIST_ONLY" -eq 0 ] && [ "$LIST_SCHEDULED" -eq 0 ]; then + refuse_primary_checkout_for_task +fi + case "${MODE:-}" in all) select_all @@ -2482,7 +2567,7 @@ family_bump() { record_script_result() { local script=$1 rc=$2 duration=$3 out=$4 end_iso=$5 - local base family expected gate_skip fail_delta + local base family expected gate_skip gate_reason fail_delta base=${script##*/} family=$(family_for_basename "$base") expected=$(expected_gate_skip_for_family "$family") @@ -2493,9 +2578,14 @@ record_script_result() { fi gate_skip=false + gate_reason= if [ "$rc" -eq 0 ] && detect_gate_skip "$out"; then gate_skip=true + gate_reason=$(gate_skip_reason "$out") SKIPPED_GATE=$((SKIPPED_GATE + 1)) + # A capability skip is the runner's only record of what this host could not + # exercise, so name it rather than leaving a silent green. + log "gate skip: $script: ${gate_reason:-<no reason given>}" fi printf 'FM_TEST_END %s %s exit=%s duration_ms=%s gate_skip=%s\n' \ @@ -2508,8 +2598,8 @@ record_script_result() { AGG_RC=1 fi - printf '%s\t%s\t%s\t%s\t%s\t%s\n' \ - "$script" "$family" "$expected" "$rc" "$duration" "$gate_skip" >>"$RECORDS" + printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \ + "$script" "$family" "$expected" "$rc" "$duration" "$gate_skip" "$gate_reason" >>"$RECORDS" family_bump "$family" "$duration" "$fail_delta" TOTAL=$((TOTAL + 1)) } @@ -2588,8 +2678,10 @@ if [ "$JOBS" -eq 1 ]; then done else # Bounded concurrent execution for admitted scripts. Each worker gets a - # private mode-0700 TMPDIR so mktemp roots cannot collide. Retries are never - # used as a green strategy. + # private mode-0700 TMPDIR so mktemp roots cannot collide. Native Windows + # Bash layers report synthetic POSIX modes, so retain chmod there but enforce + # its observed mode only where the host reports real POSIX permissions. + # Retries are never used as a green strategy. worker_n=0 active_workers=0 @@ -2615,7 +2707,7 @@ else cat "$out" fi if ! worker_directory_private "$work"; then - mode=$(stat -c %a "$work" 2>/dev/null || stat -f %Lp "$work" 2>/dev/null || echo unknown) + mode=$(stat -c %a "$work" 2>/dev/null || /usr/bin/stat -f %Lp "$work" 2>/dev/null || echo unknown) log "isolation failure: worker root is not private (reported mode $mode; $work)" rc=1 fi diff --git a/bin/fm-turnend-guard.sh b/bin/fm-turnend-guard.sh index a10211f5aad..76680ee24b1 100755 --- a/bin/fm-turnend-guard.sh +++ b/bin/fm-turnend-guard.sh @@ -35,9 +35,17 @@ # Away mode (state/.afk): the away-mode daemon owns supervision and runs the # watcher one-shot, restarting it after every wake, so the watch lock is # regularly unheld at a turn boundary with nothing wrong. A live -# identity-matched daemon holding this home, plus the unchanged fresh-beacon -# test, is what proves supervision there - see fm_afk_daemon_owns_supervision in -# bin/fm-wake-lib.sh. The strict watcher predicate is unchanged everywhere else. +# identity-matched daemon holding this home, plus a fresh beacon, is what +# proves supervision there - see fm_afk_daemon_owns_supervision in +# bin/fm-wake-lib.sh. The beacon freshness test there uses AFK_GRACE +# (fm_poll_derived_grace, docs/turnend-guard.md "Guard grace and the poll +# cadence"), not the flat $GRACE every other check on this page uses: the +# daemon starts a fresh one-shot watcher only after it finishes handling the +# previous wake, and that handling can legitimately run past a flat 300s +# window under load (a slow registered check, a busy supervisor pane) with the +# daemon perfectly healthy throughout. The strict watcher predicate and $GRACE +# are unchanged everywhere else, including for a dead daemon pid or a beacon +# older than AFK_GRACE, which still block. # # Loop-guard, codex/Grok (default) mode: never block twice in the same turn. # Codex uses stop_hook_active and Grok uses stopHookActive; typed camel-case @@ -194,10 +202,15 @@ fi # hand-off, when no watcher process holds the lock and nothing is wrong, so # requiring one here alarmed on healthy away-mode supervision. A live # identity-matched daemon holding this home is the right owner to test for. -# The beacon half of the predicate is deliberately unchanged: a daemon that -# stops restarting its watcher still blocks once the beacon passes grace, and -# a home with no daemon and no watcher blocks exactly as before. -if [ "$FM_SUP_WATCHER_FRESH" = true ] && fm_afk_daemon_owns_supervision "$STATE"; then +# The beacon half of the predicate still applies: a daemon that stops +# restarting its watcher still blocks once the beacon passes grace, and a home +# with no daemon and no watcher blocks exactly as before. It uses AFK_GRACE +# (poll-cadence-derived, see the comment above) instead of the flat $GRACE +# every other check on this page uses, so a daemon that is genuinely still +# cycling - just slower than a fixed 300s window - is not misread as down. +AFK_GRACE=${FM_GUARD_GRACE:-$(fm_poll_derived_grace)} +if [ "$(fm_path_age "$STATE/.last-watcher-beat")" -lt "$AFK_GRACE" ] \ + && fm_afk_daemon_owns_supervision "$STATE"; then allow_supervised_stop fi @@ -217,6 +230,8 @@ block_stop() { printf '● %s task(s) in flight, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_IN_FLIGHT" "$FM_SUP_BEACON_DESC" elif [ "$FM_SUP_SOURCES" -gt 0 ]; then printf '● %s process-event source(s) registered, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_SOURCES" "$FM_SUP_BEACON_DESC" + elif [ "$FM_SUP_CHECKS" -gt 0 ]; then + printf '● %s registered custom check(s), but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_CHECKS" "$FM_SUP_BEACON_DESC" else printf '● X-mode relay polling needs supervision, but no live watcher holds this home lock (last beat: %s).\n' "$FM_SUP_BEACON_DESC" fi @@ -452,6 +467,8 @@ if [ "$terminal_status" -eq 0 ]; then NEED_DESC="$FM_SUP_IN_FLIGHT task(s) in flight" elif [ "$FM_SUP_SOURCES" -gt 0 ]; then NEED_DESC="$FM_SUP_SOURCES process-event source(s) registered" + elif [ "$FM_SUP_CHECKS" -gt 0 ]; then + NEED_DESC="$FM_SUP_CHECKS registered custom check(s)" else NEED_DESC="X-mode relay polling active" fi diff --git a/bin/fm-wake-drain.sh b/bin/fm-wake-drain.sh index 73bbd30d8c6..8268bb917fa 100755 --- a/bin/fm-wake-drain.sh +++ b/bin/fm-wake-drain.sh @@ -1,5 +1,6 @@ #!/usr/bin/env bash -# Present durable watcher wake records, optionally acknowledge handled records, +# Present durable watcher wake records, retire rows no actor could ever consume, +# optionally acknowledge handled records, # annotate every unread line for validated signal status keys, surface unread # informational status lines, latest captain-facing statuses not covered by a # newer branch outcome, OPEN DECISIONS, and captain-call record divergence, @@ -67,38 +68,71 @@ ELIGIBLE_ROWS_FILE="$STATE/.branch-eligible-rows" ELIGIBLE_OWNER_FILE="$STATE/.branch-eligible-owner" MAIN_ROWS_FILE="$STATE/.main-eligible-rows" -rows_file_valid() { - [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" -} - -branch_grant_live_locked() { - local version pid identity generation current - [ -f "$ELIGIBLE_OWNER_FILE" ] && [ ! -L "$ELIGIBLE_OWNER_FILE" ] || return 1 - exec 8< "$ELIGIBLE_OWNER_FILE" || return 1 - IFS= read -r version <&8 || { exec 8<&-; return 1; } - IFS= read -r pid <&8 || { exec 8<&-; return 1; } - IFS= read -r identity <&8 || { exec 8<&-; return 1; } - IFS= read -r generation <&8 || { exec 8<&-; return 1; } - if IFS= read -r _extra <&8; then exec 8<&-; return 1; fi - exec 8<&- - [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 - case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac - case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac - current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 - [ -n "$current" ] && [ "$current" = "$identity" ] -} +rows_file_valid() { fm_wake_grant_rows_valid "$1"; } reclaim_stale_branch_grant_locked() { [ -e "$ELIGIBLE_ROWS_FILE" ] || [ -L "$ELIGIBLE_ROWS_FILE" ] || return 0 - if ! rows_file_valid "$ELIGIBLE_ROWS_FILE" || ! branch_grant_live_locked; then + if ! fm_wake_branch_grant_live "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE"; then rm -f -- "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE" fi } +# Retire rows no actor can ever consume. A claim, a presentation, and an +# acknowledgement all require the five appended fields and a numeric sequence, +# so a truncated or corrupted row is counted as queued while it can never be +# presented and can never be named by an --ack-through cutoff: left alone it +# wedges the queue for good. Main owns that repair - a branch grant can only +# name sequences that were structurally valid when it was published - and it +# runs under the queue lock, so no concurrent append is observed half-written. +# A repair that cannot be written (state/ full, unwritable, unreadable) is +# reported and never fatal: the usable rows are still presentable and +# acknowledgeable, and failing the whole drain would strand them too. +retire_unconsumable_rows_locked() { + local retired unusable queued kept + [ -f "$FM_WAKE_QUEUE" ] || return 0 + if DRAIN_TMP=$(mktemp "$STATE/.wake-queue.retire.XXXXXX") \ + && chmod 0600 "$DRAIN_TMP" \ + && unusable=$(awk -F '\t' -v keep="$DRAIN_TMP" ' + NF >= 5 && $2 ~ /^[0-9]+$/ { print > keep; next } + { shown++; if (shown <= 20) printf "wake drain: %s\n", $0 } + END { if (shown > 20) printf "wake drain: ... %d further unusable row(s) not shown\n", shown - 20 } + ' "$FM_WAKE_QUEUE"); then + queued=$(awk 'END { print NR }' "$FM_WAKE_QUEUE") + kept=$(awk 'END { print NR }' "$DRAIN_TMP") + retired=$(( queued - kept )) + if [ "$retired" -eq 0 ]; then + rm -f -- "$DRAIN_TMP" + DRAIN_TMP= + return 0 + fi + if _fm_atomic_replace "$DRAIN_TMP" "$FM_WAKE_QUEUE"; then + DRAIN_TMP= + printf 'wake drain: retired %s unusable queue row(s) that carried no sequence to present or acknowledge:\n%s\n' \ + "$retired" "$unusable" >&2 + return 0 + fi + fi + printf 'wake drain: unusable queue row(s) could not be retired (check that %s is readable and %s is writable); continuing with the rows that remain usable\n' \ + "$FM_WAKE_QUEUE" "$STATE" >&2 +} + +# One bounded line naming the rows a live branch grant is holding, so a main +# drain with nothing of its own never looks like a silently swallowed wake. +print_branch_held_notice() { + local held seqs + held=$(fm_wake_actor_pending_count branch "$ELIGIBLE_ROWS_FILE" "$ELIGIBLE_OWNER_FILE") || return 0 + [ "$held" -gt 0 ] || return 0 + seqs=$(fm_wake_grant_rows_valid "$ELIGIBLE_ROWS_FILE" \ + && awk 'NR <= 20 { printf "%s%s", (NR > 1 ? "," : ""), $1 } END { if (NR > 20) printf ",..." }' \ + "$ELIGIBLE_ROWS_FILE") + printf 'WAKE ROWS HELD BY SUPERVISION BRANCH: %s queued row(s) (%s) are granted to the live supervision branch, which presents and acknowledges them.\n' \ + "$held" "${seqs:-unknown}" +} + write_rows_file_locked() { # <target> <source> local target=$1 source=$2 if [ ! -s "$source" ]; then - rm -f -- "$target" + rm -f -- "$target" "$source" return fi chmod 0600 "$source" || return 1 @@ -606,6 +640,7 @@ else fi DRAIN_LOCK_HELD=true reclaim_stale_branch_grant_locked || exit 1 +[ "$ACTOR" != main ] || retire_unconsumable_rows_locked [ "$ACTOR" != branch ] || require_branch_eligible_rows || exit 1 if [ -n "$ACK_THROUGH" ]; then @@ -756,6 +791,11 @@ if [ "$ACTOR" = main ]; then fi claim_main_rows_locked || exit 1 if [ ! -s "$MAIN_ROWS_FILE" ]; then + # Every remaining row is reserved by the live branch grant, which presents + # and acknowledges them itself. Say so rather than exiting silently: a + # drain that prints nothing while the queue is visibly non-empty reads as a + # lost wake, and leaves the caller with no idea who owns what is queued. + print_branch_held_notice fm_lock_release "$FM_WAKE_QUEUE_LOCK" DRAIN_LOCK_HELD=false (print_status_presentation) || true diff --git a/bin/fm-wake-grant.sh b/bin/fm-wake-grant.sh index 2cc604f5f2c..bdff2fd4ead 100755 --- a/bin/fm-wake-grant.sh +++ b/bin/fm-wake-grant.sh @@ -22,27 +22,11 @@ trap cleanup EXIT trap 'exit 130' INT trap 'exit 143' TERM -rows_valid() { - [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" -} +# fm-wake-lib.sh owns both the grant row-list shape and the owner-record read. +rows_valid() { fm_wake_grant_rows_valid "$1"; } -owner_matches() { - local expected_pid=${1:-} expected_generation=${2:-} version pid identity generation current - [ -f "$BRANCH_OWNER" ] && [ ! -L "$BRANCH_OWNER" ] || return 1 - exec 8< "$BRANCH_OWNER" || return 1 - IFS= read -r version <&8 || { exec 8<&-; return 1; } - IFS= read -r pid <&8 || { exec 8<&-; return 1; } - IFS= read -r identity <&8 || { exec 8<&-; return 1; } - IFS= read -r generation <&8 || { exec 8<&-; return 1; } - if IFS= read -r _extra <&8; then exec 8<&-; return 1; fi - exec 8<&- - [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 - case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac - case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac - [ -z "$expected_pid" ] || [ "$pid" = "$expected_pid" ] || return 1 - [ -z "$expected_generation" ] || [ "$generation" = "$expected_generation" ] || return 1 - current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 - [ -n "$current" ] && [ "$current" = "$identity" ] +owner_matches() { # [<pid>] [<generation>] + fm_wake_branch_owner_matches "$BRANCH_OWNER" "${1:-}" "${2:-}" } case "${1:-}" in diff --git a/bin/fm-wake-lib.sh b/bin/fm-wake-lib.sh index 4753920fe3a..cc40180c1a6 100755 --- a/bin/fm-wake-lib.sh +++ b/bin/fm-wake-lib.sh @@ -34,8 +34,17 @@ _fm_wake_require_timeout() { . "$FM_WAKE_LIB_DIR/fm-timeout-lib.sh" } -fm_current_pid() { - printf '%s\n' "${BASHPID:-$$}" +# Pass a variable name to capture this frame's pid without forking it in $(). +# On Bash 3.2, exec a child shell so its PPID identifies this frame, unlike $$. +fm_current_pid() { # [output-variable] + local fm_pid + fm_pid=${BASHPID:-$(exec sh -c 'printf "%s\n" "$PPID"')} || return 1 + case "$fm_pid" in ''|*[!0-9]*|0) return 1 ;; esac + if [ "$#" -gt 0 ]; then + printf -v "$1" '%s' "$fm_pid" + else + printf '%s\n' "$fm_pid" + fi } fm_pid_alive() { @@ -85,7 +94,7 @@ fm_pid_identity() { fm_path_mtime() { if [ "$_FM_UNAME" = Darwin ]; then - stat -f %m "$1" 2>/dev/null + /usr/bin/stat -f %m "$1" 2>/dev/null else stat -c %Y "$1" 2>/dev/null fi @@ -101,6 +110,25 @@ fm_path_age() { echo $(( now - m )) } +# fm_poll_derived_grace [poll-seconds] +# Default guard-grace derivation: max(300, poll + 60). A watcher touches its +# liveness beacon once per poll cycle, so a fixed 300s grace stops correctly +# bounding staleness once the poll cadence reaches or exceeds it; growing the +# default with the cadence while keeping the historical 300s floor for the +# common short-poll case fixes that without a caller-specific constant. +# Defaults to $FM_POLL (fm-watch.sh's own poll env var) when no argument is +# given, so a caller with no independent notion of the poll cadence still +# derives the same default fm-watch.sh itself would use. +# docs/turnend-guard.md "Guard grace and the poll cadence" is the single owner +# of the rationale; every FM_GUARD_GRACE default should derive from this. +fm_poll_derived_grace() { + local poll=${1:-${FM_POLL:-15}} margin=60 derived + case "$poll" in ''|*[!0-9]*) poll=15 ;; esac + derived=$((poll + margin)) + [ "$derived" -ge 300 ] || derived=300 + printf '%s\n' "$derived" +} + # fm_watcher_lock_unheld <state> # True when the watcher lock or its symlinked owner directory is absent, or when # the existing lock records no pid at all. Any non-empty pid remains held here; @@ -186,7 +214,7 @@ fm_supervision_model() { harness=$("$FM_WAKE_LIB_DIR/fm-harness.sh" 2>/dev/null || printf unknown) case "$harness" in claude|copilot|cursor) printf 'autoarm\n' ;; - pi|pi-signed) printf 'extension\n' ;; + pi|pi-signed|omp) printf 'extension\n' ;; *) printf 'persistent\n' ;; esac } @@ -235,14 +263,39 @@ fm_pi_extension_loaded() { # backstop that catches a cycle the watch extension failed to restore, so a home # missing it has no benign hand-off to tolerate. fm_pi_extension_owns_supervision() { - local state=$1 root=$2 lock session_pid pair source marker version - lock="$state/.lock" - for pair in \ + fm_extension_pair_owns_supervision "$1" "$2/.pi/extensions" \ "fm-primary-pi-watch.ts:.pi-watch-extension-loaded" \ - "fm-primary-turnend-guard.ts:.pi-turnend-extension-loaded"; do + "fm-primary-turnend-guard.ts:.pi-turnend-extension-loaded" +} + +# fm_omp_extension_owns_supervision <state> <root> +# The omp (Oh My Pi) primary's proof, keyed on its own two tracked extensions +# under .omp/extensions/ and their own state markers. It is a separate proof on +# purpose: omp must never inherit the Pi tolerance by accident, and a Pi home +# never satisfies the omp markers. Both proofs bind to the pid in state/.lock, +# so a session on one harness cannot vouch for a home held by the other. +fm_omp_extension_owns_supervision() { + fm_extension_pair_owns_supervision "$1" "$2/.omp/extensions" \ + "fm-primary-omp-watch.ts:.omp-watch-extension-loaded" \ + "fm-primary-turnend-guard.ts:.omp-turnend-extension-loaded" +} + +# fm_extension_owns_supervision <state> <root> +# The extension-model proof the verdict below consults: whichever extension +# family's markers the lock-owning session recorded. Exactly one family can +# match because both bind to the same lock pid. +fm_extension_owns_supervision() { + fm_pi_extension_owns_supervision "$1" "$2" || fm_omp_extension_owns_supervision "$1" "$2" +} + +fm_extension_pair_owns_supervision() { # <state> <extension-dir> <source:marker>... + local state=$1 dir=$2 lock session_pid pair source marker version + shift 2 + lock="$state/.lock" + for pair in "$@"; do source=${pair%%:*} marker=${pair#*:} - version=$(fm_pi_extension_version "$root/.pi/extensions/$source") || return 1 + version=$(fm_pi_extension_version "$dir/$source") || return 1 fm_pi_extension_loaded "$state/$marker" "$version" "$lock" || return 1 done session_pid=$(sed -n '1p' "$lock" 2>/dev/null) @@ -292,7 +345,8 @@ fm_afk_daemon_owns_supervision() { # because the watcher only runs between turns; only a stale beacon is a lapse. # extension: a live identity-matched watcher is the ordinary healthy state, but a # genuinely unheld lock is also healthy while the beacon is fresh AND a live Pi -# session provably owns continuity (fm_pi_extension_owns_supervision) - that is the +# session provably owns continuity (fm_extension_owns_supervision: the Pi or the +# omp extension pair, whichever the lock-owning session recorded) - that is the # extension's own tear-down-and-respawn hand-off, which it retries and escalates # itself. A lock with any recorded pid remains down if the strict health check fails. # Without ownership proof an unheld lock is down exactly as before, so an unloaded, @@ -326,7 +380,7 @@ fm_watcher_supervision_verdict() { FM_WATCHER_VERDICT_OK=true elif [ "$fresh" = true ]; then if [ "$model" = extension ] && fm_watcher_lock_unheld "$state" \ - && fm_pi_extension_owns_supervision "$state" "$root"; then + && fm_extension_owns_supervision "$state" "$root"; then # shellcheck disable=SC2034 # Read by callers after the function returns. FM_WATCHER_VERDICT_OK=true else @@ -354,7 +408,7 @@ fm_lock_set_role() { autoarm|terminal-check) : ;; *) return 1 ;; esac - current=${BASHPID:-$$} + fm_current_pid current || return 1 pid=$(cat "$lockdir/pid" 2>/dev/null || true) [ "$pid" = "$current" ] || return 1 printf '%s\n' "$role" > "$lockdir/role" 2>/dev/null || return 1 @@ -394,7 +448,7 @@ fm_lock_owner_dir() { fm_lock_prepare_owner() { local ownerdir=$1 mypid back - mypid=${BASHPID:-$$} + fm_current_pid mypid || return 1 printf '%s\n' "$mypid" > "$ownerdir/pid" 2>/dev/null || return 1 back=$(cat "$ownerdir/pid" 2>/dev/null || true) [ "$back" = "$mypid" ] @@ -457,7 +511,7 @@ fm_lock_claim_blocked_by_steal() { fm_lock_claim() { local lockdir=$1 ownerdir=$2 allowed_steal_owner=${3:-} mypid back - mypid=${BASHPID:-$$} + fm_current_pid mypid || return 1 if ! { printf '%s\n' "$mypid" > "$ownerdir/pid"; } 2>/dev/null; then fm_lock_discard_owner "$ownerdir" return 1 @@ -965,7 +1019,7 @@ fm_lock_try_acquire_recovery() { } fm_lock_try_acquire() { - local lockdir=$1 pid steal cur rc steal_owner primary_owner + local lockdir=$1 pid steal cur rc steal_owner primary_owner current FM_LOCK_HELD_PID= FM_LOCK_OWNER_DIR= FM_LOCK_RECOVERED_PID= @@ -974,10 +1028,9 @@ fm_lock_try_acquire() { return 0 fi - # Compare against ${BASHPID:-$$} inline, never via a command substitution: - # $() forks a subshell whose BASHPID is not this frame's pid. + fm_current_pid current || return 1 pid=$(cat "$lockdir/pid" 2>/dev/null || true) - if [ -n "$pid" ] && [ "$pid" = "${BASHPID:-$$}" ]; then + if [ -n "$pid" ] && [ "$pid" = "$current" ]; then # The recorded holder is THIS very process. Single-threaded bash can only # observe that when an interrupting trap abandoned the frame that held the # lock mid-critical-section (e.g. TERM inside a recovery-marker section, @@ -1099,7 +1152,7 @@ _fm_lock_acquire_wait_handoff() { # <lockdir> <caller-pid> else ownerdir=$lockdir fi - current=${BASHPID:-$$} + fm_current_pid current || { fm_lock_release "$lockdir"; return 1; } back=$(cat "$ownerdir/pid" 2>/dev/null || true) if [ "$back" != "$current" ] \ || ! printf '%s\n' "$caller_pid" > "$ownerdir/pid" 2>/dev/null \ @@ -1127,7 +1180,7 @@ fm_lock_acquire_wait_bounded() { return 0 fi - caller_pid=${BASHPID:-$$} + fm_current_pid caller_pid || return 1 # shellcheck disable=SC2016 # Positional parameters expand in the child shell. if fm_run_timed "$seconds" env \ "FM_STATE_OVERRIDE=$STATE" \ @@ -1172,7 +1225,7 @@ fm_lock_acquire_wait_bounded() { fm_lock_release() { local lockdir=$1 pid current ownerdir - current=${BASHPID:-$$} + fm_current_pid current || return 1 if [ -L "$lockdir" ]; then ownerdir=$(fm_lock_link_owner "$lockdir" 2>/dev/null || true) [ -n "$ownerdir" ] || return 0 @@ -1225,6 +1278,75 @@ fm_task_set_lock_path() { # <state-dir> printf '%s/.task-set.lock\n' "$state" } +# The top-most firstmate home reachable from this one on THIS machine, used as +# the single anchor every local home agrees on for machine-local shared state. +# +# A local parent binding is followed upward. A remote parent binding terminates +# the walk at the current home, which is the correct answer rather than an +# error: the parent lives on another machine, so its filesystem can neither hold +# nor be observed by a lock taken here, and a remote-seeded home is itself the +# top of the local tree that bin/fm-teardown.sh's collect_local_firstmate_states +# enumerates (that walk already skips remote registry entries for the same +# reason). Refusing a remote binding instead made every operation anchored here +# fail closed inside a remote secondmate home and its local descendants. +# +# Everything else still fails closed: an unreadable or malformed binding, an +# unreachable local parent, a cycle, and a chain deeper than the bound. +fm_firstmate_root_home() { + local home=${1:-$FM_HOME} marker parent seen="|" depth=0 + home=$(CDPATH='' cd -- "$home" 2>/dev/null && pwd -P) || return 1 + while [ -e "$home/.fm-secondmate-parent" ] || [ -L "$home/.fm-secondmate-parent" ]; do + marker="$home/.fm-secondmate-parent" + if ! command -v fm_secondmate_parent_record_parse >/dev/null 2>&1; then + # shellcheck source=bin/fm-secondmate-parent-lib.sh + . "$FM_WAKE_LIB_DIR/fm-secondmate-parent-lib.sh" + fi + fm_secondmate_parent_record_parse "$marker" || return 1 + case "$FM_SECONDMATE_PARENT_ROUTE" in + local) ;; + remote) break ;; + *) return 1 ;; + esac + parent=$(CDPATH='' cd -- "$FM_SECONDMATE_PARENT_HOME" 2>/dev/null && pwd -P) || return 1 + case "$seen" in *"|$parent|"*) return 1 ;; esac + seen="$seen$home|" + home=$parent + depth=$((depth + 1)) + [ "$depth" -le 64 ] || return 1 + done + printf '%s\n' "$home" +} + +# The one lock serializing Treehouse slot allocation and return for a project. +# +# It is anchored in the local root home's state directory so that every home on +# this machine that can reach the same pool - the root, and each secondmate home +# below it, including a remote-seeded home and its own local descendants - +# derives the identical path. Its identity is the project's resolved origin, so +# separate clones of one origin share a single lock; an origin-less local-only +# project falls back to its own worktree top instead of failing to resolve. +fm_treehouse_project_lock_path() { # <project-dir> + local project=$1 root origin identity hash top + [ -d "$project" ] || return 1 + root=$(fm_firstmate_root_home "$FM_HOME") || return 1 + origin=$(git -C "$project" remote get-url origin 2>/dev/null || true) + if [ -n "$origin" ]; then + case "$origin" in + /*) [ ! -d "$origin" ] || origin=$(CDPATH='' cd -- "$origin" 2>/dev/null && pwd -P) || return 1 ;; + *://*|*:* ) ;; + *) [ ! -d "$project/$origin" ] || origin=$(CDPATH='' cd -- "$project/$origin" 2>/dev/null && pwd -P) || return 1 ;; + esac + identity=$origin + else + top=$(git -C "$project" rev-parse --show-toplevel 2>/dev/null) || return 1 + top=$(CDPATH='' cd -- "$top" 2>/dev/null && pwd -P) || return 1 + identity=$top + fi + hash=$(printf '%s' "$identity" | git hash-object --stdin 2>/dev/null) || return 1 + [ -d "$root/state" ] || return 1 + printf '%s/.treehouse-project-%s.lock\n' "$root/state" "$hash" +} + fm_failure_episode_reset() { local state=$1 mode=${2:-acquire} lock current pid acquired=0 path lock="$state/.turnend-claude-blocks.lock" @@ -1234,7 +1356,7 @@ fm_failure_episode_reset() { acquired=1 ;; held) - current=${BASHPID:-$$} + fm_current_pid current || return 1 pid=$(cat "$lock/pid" 2>/dev/null || true) [ "$pid" = "$current" ] || return 1 ;; @@ -1691,6 +1813,23 @@ fm_wake_queued_keys_locked() { "$FM_WAKE_QUEUE" 2>/dev/null || true } +fm_wake_secondmate_progress_marker_write() { # <task> <observed-at> <oldest-row-key> + local task=$1 observed_at=$2 oldest_row_key=$3 marker tmp + case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + case "$observed_at" in ''|*[!0-9]*) return 1 ;; esac + case "$oldest_row_key" in ''|*[!0-9-]*) return 1 ;; esac + marker="$STATE/.secondmate-wake-progress-$task" + if [ -e "$marker" ] || [ -L "$marker" ]; then + [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + fi + tmp=$(mktemp "$STATE/.secondmate-wake-progress.XXXXXX") || return 1 + if ! printf '%s\t%s\n' "$observed_at" "$oldest_row_key" > "$tmp" || ! chmod 0600 "$tmp" \ + || ! _fm_atomic_replace "$tmp" "$marker"; then + rm -f -- "$tmp" + return 1 + fi +} + fm_wake_secondmate_stall_marker_write() { # <task> <row-key> local task=$1 row_key=$2 marker tmp case "$task" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac @@ -1788,6 +1927,86 @@ fm_wake_print_deduped() { ' "$file" } +# --- branch grant evidence and per-actor pending rows ------------------------ +# +# docs/watcher-continuity.md "Per-actor acknowledgement" owns the contract these +# helpers read; this is its single implementation, shared by the drain (which +# repairs and consumes a grant under the queue lock), the grant publisher, and +# the guard (which only counts, and never takes the lock). + +# 0 when <rows-file> is a non-empty list of distinct sequence numbers. +fm_wake_grant_rows_valid() { # <rows-file> + [ -s "$1" ] && awk 'BEGIN { ok=1 } !/^[0-9]+$/ || seen[$0]++ { ok=0 } END { exit !ok }' "$1" +} + +# 0 when <owner-file> holds the supported record, names a live process whose +# identity still matches what was recorded, and matches any expected pid and +# generation the caller pins. An unreadable, malformed, or superseded record is +# not a match, so uncertainty reads as "no live owner". +fm_wake_branch_owner_matches() { # <owner-file> [<pid>] [<generation>] + local file=$1 expected_pid=${2:-} expected_generation=${3:-} + local version pid identity generation current extra + [ -f "$file" ] && [ ! -L "$file" ] || return 1 + exec 8< "$file" || return 1 + IFS= read -r version <&8 || { exec 8<&-; return 1; } + IFS= read -r pid <&8 || { exec 8<&-; return 1; } + IFS= read -r identity <&8 || { exec 8<&-; return 1; } + IFS= read -r generation <&8 || { exec 8<&-; return 1; } + if IFS= read -r extra <&8; then exec 8<&-; return 1; fi + exec 8<&- + [ "$version" = fm-branch-eligible-owner-v1 ] || return 1 + case "$pid" in ''|*[!0-9]*|1) return 1 ;; esac + case "$generation" in ''|*[!A-Za-z0-9._-]*) return 1 ;; esac + [ -z "$expected_pid" ] || [ "$pid" = "$expected_pid" ] || return 1 + [ -z "$expected_generation" ] || [ "$generation" = "$expected_generation" ] || return 1 + current=$(fm_pid_identity "$pid" 2>/dev/null) || return 1 + [ -n "$current" ] && [ "$current" = "$identity" ] +} + +# 0 when a branch grant is currently reserving rows: a valid row snapshot whose +# recorded owner is still live. Anything else means no row is reserved. +fm_wake_branch_grant_live() { # <rows-file> <owner-file> + fm_wake_grant_rows_valid "$1" && fm_wake_branch_owner_matches "$2" +} + +# How many queued rows <actor> can act on right now - exactly the rows a drain +# by that actor would present or retire, and therefore the only rows worth +# telling that actor to drain. Main owns every structurally valid row a live +# branch grant does not reserve, plus every structurally invalid row. The branch +# owns exactly the rows its live grant names. Read without the queue lock: a +# torn read can only mis-count one poll, and the drain re-derives the set under +# the lock before it presents or mutates anything. +fm_wake_actor_pending_count() { # <actor> [<rows-file> <owner-file>] + local actor=${1:-main} rows=${2:-$STATE/.branch-eligible-rows} + local owner=${3:-$STATE/.branch-eligible-owner} grant='' count='' + [ -f "$FM_WAKE_QUEUE" ] || { printf '0\n'; return 0; } + if fm_wake_branch_grant_live "$rows" "$owner"; then + grant=$rows + fi + if [ "$actor" = branch ]; then + [ -n "$grant" ] || { printf '0\n'; return 0; } + count=$(awk -F '\t' -v seqs="$grant" ' + BEGIN { while ((getline line < seqs) > 0) keep[line] = 1 } + NF >= 5 && $2 ~ /^[0-9]+$/ && ($2 in keep) { n++ } + END { print n + 0 } + ' "$FM_WAKE_QUEUE") || count='' + else + count=$(awk -F '\t' -v seqs="$grant" ' + BEGIN { if (seqs != "") while ((getline line < seqs) > 0) reserved[line] = 1 } + NF < 5 || $2 !~ /^[0-9]+$/ { n++; next } + !($2 in reserved) { n++ } + END { print n + 0 } + ' "$FM_WAKE_QUEUE") || count='' + fi + # A queue that exists but cannot be counted (unreadable file, unreadable + # state/) is not evidence of an empty queue: report a pending row so callers + # still raise the alarm on a queue nobody can prove is drained. A failed count + # is decided by awk's exit status, not by what it printed, because an awk that + # reaches END after failing to open the queue would otherwise report 0 rows. + case "$count" in ''|*[!0-9]*) count=1 ;; esac + printf '%s\n' "$count" +} + # --- signal announcement signatures ----------------------------------------- # # The watcher's per-file signal scan (bin/fm-watch.sh scan_signals) detects a @@ -1805,7 +2024,7 @@ fm_wake_signal_sig() { # <file> -> reported-state signature status_observed_signature "$1" ;; *) - if [ "$_FM_UNAME" = Darwin ]; then stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi + if [ "$_FM_UNAME" = Darwin ]; then /usr/bin/stat -f '%z:%Fm' "$1" 2>/dev/null; else stat -c '%s:%Y' "$1" 2>/dev/null; fi ;; esac } diff --git a/bin/fm-watch-checkpoint.sh b/bin/fm-watch-checkpoint.sh index 1fb2b118b2a..35280f1f6f4 100755 --- a/bin/fm-watch-checkpoint.sh +++ b/bin/fm-watch-checkpoint.sh @@ -63,12 +63,19 @@ run_with_perl_timeout() { } local $SIG{ALRM} = sub { kill "TERM", -$pid; - select undef, undef, undef, 0.2; - kill "KILL", -$pid; + my $grace = $ENV{FM_SIGNAL_GRACE} || 5; + local $SIG{ALRM} = sub { + kill "KILL", -$pid; + waitpid $pid, 0; + exit 124; + }; + alarm $grace; + waitpid $pid, 0; exit 124; }; alarm $seconds; waitpid $pid, 0; + alarm 0; exit($? >> 8); ' "$SECONDS_ARG" "$SCRIPT_DIR/fm-watch.sh" } diff --git a/bin/fm-watch.sh b/bin/fm-watch.sh index 51b7de4a67f..82d1c1f6c8c 100755 --- a/bin/fm-watch.sh +++ b/bin/fm-watch.sh @@ -79,11 +79,14 @@ # check: inactive-outcome bounded poll-loop reconciliation found a suspicious # inactive terminal outcome that still lacks its durable # upstream receipt -# check: secondmate wake-loop stalled: mate=<id> row=<seq> age=<seconds>s -# the oldest valid row in an endpoint-recorded local -# secondmate home's durable wake queue exceeded -# FM_SECONDMATE_WAKE_STALL_SECS; observation is read-only -# and one parent receipt suppresses repeats for that row +# check: secondmate wake-loop stalled: mate=<id> row=<seq> idle=<seconds>s +# an actionable row in an endpoint-recorded local +# secondmate home's durable wake queue did not advance +# between observations for FM_SECONDMATE_WAKE_STALL_SECS +# while the mate was not in an active turn; declared +# external-wait pause rows do not feed this escalation, +# observation is read-only, and one parent notification +# covers each no-progress episode # For normal supervision, resume the session-start primary-harness protocol # after each printed reason. Direct duplicate invocations of this script still # no-op through the watcher singleton lock. @@ -129,15 +132,15 @@ mkdir -p "$STATE" # shellcheck source=bin/fm-busy-lib.sh . "$SCRIPT_DIR/fm-busy-lib.sh" # Steering-inbox loss detection: bin/fm-task-inbox-lib.sh owns the record, -# doorbell, and re-ring ladder contracts; this watcher only supplies the busy -# gate and the wake emission (inbox_steer_check below). +# doorbell, re-ring ladder, and unavailable-endpoint contracts; this watcher +# supplies their live endpoint and busy checks plus wake emission +# (inbox_steer_check below). # shellcheck source=bin/fm-task-inbox-lib.sh . "$SCRIPT_DIR/fm-task-inbox-lib.sh" WATCH_LOCK="$STATE/.watch.lock" WATCH_PATH="$SCRIPT_DIR/fm-watch.sh" WATCHER_DOWNTIME_MARKER="$STATE/.watcher-down" -WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-300}} # The singleton-lock acquisition, EXIT trap, and the blocking supervision loop # all live below the source guard at the very bottom of this file (see "Main # entry"). Sourcing this file for unit tests therefore loads the functions - @@ -152,8 +155,10 @@ WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-300}} # appended to that garbage. Arithmetic under `set -u` then aborts on the stray # token (e.g. the word "File" read as an unset variable), which silently kills the # watcher mid-cycle. Detect the platform once and pick the right form. +# On Darwin, call /usr/bin/stat rather than PATH-resolved stat so GNU coreutils +# cannot shadow the BSD `-f` syntax. if [ "$(uname)" = Darwin ]; then - stat_mtime() { stat -f %m "$1" 2>/dev/null; } # epoch seconds of mtime + stat_mtime() { /usr/bin/stat -f %m "$1" 2>/dev/null; } # epoch seconds of mtime else stat_mtime() { stat -c %Y "$1" 2>/dev/null; } fi @@ -162,6 +167,15 @@ fi # turn-ended signature, annotation staleness checks, and guarded bookkeeping writes. POLL=${FM_POLL:-15} # seconds between cycles +# The liveness beacon is touched once per cycle, immediately before the +# terminal wait below (event_wait_or_sleep) as well as at the top of the next +# one, so a healthy cycle's beacon can legitimately age up to POLL seconds +# between touches. fm_poll_derived_grace (bin/fm-wake-lib.sh, already sourced +# transitively above) is the single owner of the max(300, poll+60) +# derivation - see docs/turnend-guard.md "Guard grace and the poll cadence". +# This recomputes the library default above now that the real configured +# POLL is known. +WATCHER_STALE_GRACE=${FM_WATCHER_STALE_GRACE:-${FM_GUARD_GRACE:-$(fm_poll_derived_grace "$POLL")}} HEARTBEAT=${FM_HEARTBEAT:-600} # base seconds between heartbeat scans HEARTBEAT_MAX=${FM_HEARTBEAT_MAX:-7200} # heartbeat backoff cap CHECK_INTERVAL=${FM_CHECK_INTERVAL:-300} # seconds between *.check.sh sweeps @@ -215,8 +229,12 @@ STALE_ESCALATE_SECS=${FM_STALE_ESCALATE_SECS:-240} # idle secs before a provabl # between completed turns, including long tool calls, builds, or test runs. BUSY_TURN_MAX_SECS=${FM_BUSY_TURN_MAX_SECS:-3600} # A local secondmate's foreign queue is checked on every poll, but only after this -# bounded age can it produce a parent notification. -SECONDMATE_WAKE_STALL_SECS=${FM_SECONDMATE_WAKE_STALL_SECS:-60} +# bounded interval with no drain progress can it produce a parent notification. +# A healthy mate drains its queue between turns, not inside one, so this default +# sits above a real turn; it is only the backstop behind the active-turn gate in +# secondmate_wake_stall_tick, never a substitute for it. +SECONDMATE_WAKE_STALL_SECS=${FM_SECONDMATE_WAKE_STALL_SECS:-} +case "$SECONDMATE_WAKE_STALL_SECS" in ''|*[!0-9]*|0) SECONDMATE_WAKE_STALL_SECS=180 ;; esac # A crew that declared a pause is idling on a known external wait, so its stale # pane is absorbed rather than wedge-escalated. # A captain-held or paused crew whose agent has confidently exited uses the same @@ -324,13 +342,30 @@ window_key() { # <window> printf '%s' "${key//./_}" } +inbox_steer_escalate_unavailable() { # <window> <task> <record> + local w=$1 task=$2 rec=$3 reason + reason="stale: $w (unread firstmate instruction: $rec is unhandled and the worker's agent has exited or its endpoint is missing, so the doorbell was not typed; recover the worker)" + if [ ! -d "${rec%/*}" ] || [ ! -f "$rec" ]; then + fm_task_inbox_due_action "$STATE" "$task" >/dev/null || true + return 0 + fi + fm_wake_append stale "$w" "$reason" || exit 1 + if ! fm_task_inbox_record_escalated "$STATE" "$task" "$rec"; then + echo "error: stale wake was queued for $task but its inbox escalation marker could not be written" >&2 + exit 1 + fi + wake "$reason" +} + # Steering-inbox loss detection, one cheap check per recorded window per poll. # Quiet when healthy: an absent, empty, or handled inbox costs one directory # glob and produces nothing. When the ladder (fm_task_inbox_due_action, the # policy owner) reports a due action, a busy pane just waits - the record is # durable and the worker will reach a turn boundary - an idle pane gets one # delivery attempt, and a spent attempt budget surfaces as an ordinary stale -# wake for stuck-crewmate-recovery. If the attempt's ladder write fails while +# wake for stuck-crewmate-recovery, and a pane whose agent is positively dead +# or missing skips the ladder altogether: it is never typed into and surfaces +# as that same stale wake exactly once. If the attempt's ladder write fails while # its record remains unhandled, that unwritable state surfaces through the same # stale path instead of silently re-ringing forever; acknowledgement or teardown # still makes the race quiet. The attempt is data-plane typing or a @@ -339,7 +374,7 @@ window_key() { # <window> # too: their pane-staleness exemption is about quiet panes being healthy, # while an unacknowledged instruction past the ladder is a stuck steer. inbox_steer_check() { # <window> <task> - local w=$1 task=$2 action verb rec count tail40 reason ring_rc + local w=$1 task=$2 action verb rec count tail40 reason ring_rc backend agent_state action=$(fm_task_inbox_due_action "$STATE" "$task") || return 0 verb=${action%% *} [ "$verb" != quiet ] || return 0 @@ -351,14 +386,26 @@ inbox_steer_check() { # <window> <task> rec=${rec% *} ;; esac - tail40=$(fm_backend_capture "$(window_backend "$w")" "$w" 40 "$(window_label "$w")" 2>/dev/null) || tail40= + backend=$(window_backend "$w") + agent_state=$(fm_backend_agent_state "$backend" "$w" 2>/dev/null || true) + case "$agent_state" in + dead|missing) + inbox_steer_escalate_unavailable "$w" "$task" "$rec" + return 0 + ;; + esac + tail40=$(fm_backend_capture "$backend" "$w" 40 "$(window_label "$w")" 2>/dev/null) || tail40= if window_is_busy "$w" "$tail40"; then return 0 fi case "$verb" in ring) ring_rc=0 - fm_task_inbox_ring "$(window_backend "$w")" "$w" "$rec" "$(window_label "$w")" || ring_rc=$? + fm_task_inbox_ring "$backend" "$w" "$rec" "$(window_label "$w")" || ring_rc=$? + if [ "$ring_rc" -eq 3 ]; then + inbox_steer_escalate_unavailable "$w" "$task" "$rec" + return 0 + fi if ! fm_task_inbox_record_ring "$STATE" "$task" "$rec"; then if [ ! -f "$rec" ]; then fm_task_inbox_due_action "$STATE" "$task" >/dev/null || true @@ -602,14 +649,23 @@ recorded_windows() { done } -# Print the oldest structurally valid row in a local secondmate's foreign queue. -# This is a read-only observation: the receiving home owns acknowledgement and -# this parent never changes the row or the foreign queue. +# Print the oldest structurally valid ACTIONABLE row in a local secondmate's +# foreign queue. A stale recheck that explicitly identifies itself as a declared +# external-wait pause is not evidence that the mate's wake loop is stuck: the +# pause cadence already owns that bounded visibility, and blocked waits remain +# actionable because they do not carry this declaration. This is a read-only +# observation: the receiving home owns acknowledgement and this parent never +# changes the row or the foreign queue. secondmate_oldest_queue_row() { # <queue-path> local queue=$1 [ -f "$queue" ] && [ ! -L "$queue" ] || return 0 awk -F '\t' ' - NF >= 5 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ { + function declared_external_pause(kind, payload) { + return kind == "stale" \ + && payload ~ /^stale: .*\(paused [0-9]+s, awaiting external - declared (pause,|paused\))/ + } + NF >= 5 && $1 ~ /^[0-9]+$/ && $2 ~ /^[0-9]+$/ \ + && !declared_external_pause($3, $5) { if (!found || $2 < seq) { found = 1 seq = $2 @@ -620,14 +676,39 @@ secondmate_oldest_queue_row() { # <queue-path> ' "$queue" 2>/dev/null || true } -# Surface one durable parent check for one unchanged foreign row after its -# bounded age. The primary marker and queued-key check make repeated watcher -# cycles converge without a notification storm, while an empty queue removes -# only this home's marker so a later row can be observed. +# 0 iff <task> is demonstrably inside an active turn, through the watcher's own +# busy-state knowledge: an exact busy verdict from the semantic contract, bounded +# by the same BUSY_TURN_MAX_SECS that stops a busy pane from proving liveness +# forever. A mate mid-turn has not stopped draining its queue - it simply drains +# between turns - so this gate, not the elapsed interval, is what separates a +# healthy mate from a frozen wake loop. Any absence of proof (no window, a failed +# capture, an idle or unknown verdict, a busy pane past the bound) is NOT an +# active turn, so a frozen queue still escalates. +secondmate_in_active_turn() { # <task> <window> + local task=$1 w=$2 tail40 + [ -n "$w" ] || return 1 + ! busy_turn_over_age "$task" || return 1 + tail40=$(fm_backend_capture "$(window_backend "$w")" "$w" 40 "$(window_label "$w")" 2>/dev/null) || return 1 + window_is_busy "$w" "$tail40" +} + +# Surface one durable parent check when the foreign queue's drain position has +# not moved for the bounded interval. The progress marker records that position +# as the same epoch-sequence row identity the stall receipts use, so the timer +# restarts whenever a different row becomes the oldest actionable one - as the +# mate drains, and as a queue reprovisioned under the same task id starts its +# own generation of rows at whatever sequence it restarts, and neither is a +# continued no-progress episode; row creation time belongs to that identity but +# never to the interval. A moved position ends an alerted episode and starts a +# new observation interval, so a newly-oldest row cannot alert immediately while +# a later genuine freeze remains visible. A mate demonstrably inside an active +# turn never escalates, so the interval is only the backstop behind that gate. +# Receipts close the append-before-marker crash window without changing the +# foreign queue. secondmate_wake_stall_tick() { local now=$(( $(date +%s) )) threshold=$SECONDMATE_WAKE_STALL_SECS - local meta task kind remote_host home queue row epoch seq row_key marker receipt receipt_dir notify_key queued age reason - case "$threshold" in ''|*[!0-9]*|0) threshold=60 ;; esac + local meta task kind remote_host home queue row epoch seq row_key marker progress_marker progress observed_at observed_key + local receipt receipt_dir notify_key queued idle reason episode_alerted # Endpoint metadata admits this queue-loop check; secondmate-liveness owns registered mates whose endpoint is missing or dead. for meta in "$STATE"/*.meta; do [ -e "$meta" ] || continue @@ -645,9 +726,10 @@ secondmate_wake_stall_tick() { queue="$home/state/.wake-queue" row=$(secondmate_oldest_queue_row "$queue") marker="$STATE/.secondmate-wake-stall-$task" + progress_marker="$STATE/.secondmate-wake-progress-$task" receipt_dir="$STATE/.secondmate-wake-stall-receipts/$task" if [ -z "$row" ]; then - rm -f "$marker" + rm -f "$marker" "$progress_marker" if [ -e "$receipt_dir" ] || [ -L "$receipt_dir" ]; then [ -d "$receipt_dir" ] && [ ! -L "$receipt_dir" ] || return 1 rm -rf -- "$receipt_dir" || return 1 @@ -659,17 +741,39 @@ $row EOF case "$epoch" in ''|*[!0-9]*) continue ;; esac case "$seq" in ''|*[!0-9]*) continue ;; esac - age=$((now - epoch)) - [ "$age" -ge "$threshold" ] || continue row_key="$epoch-$seq" - receipt="$receipt_dir/$row_key" + episode_alerted=0 if [ -e "$marker" ] || [ -L "$marker" ]; then [ -f "$marker" ] && [ ! -L "$marker" ] || return 1 + episode_alerted=1 + fi + progress=$(cat "$progress_marker" 2>/dev/null || true) + observed_at=${progress%%[[:space:]]*} + observed_key=${progress#*[[:space:]]} + if [ "$observed_at" = "$progress" ]; then + observed_key= + else + observed_key=${observed_key%%[[:space:]]*} + fi + case "$observed_at" in ''|*[!0-9]*) observed_at= ;; esac + case "$observed_key" in ''|*[!0-9-]*) observed_key= ;; esac + if [ -z "$observed_at" ] || [ -z "$observed_key" ] \ + || [ "$now" -lt "$observed_at" ] || [ "$row_key" != "$observed_key" ]; then + fm_wake_secondmate_progress_marker_write "$task" "$now" "$row_key" || return 1 + [ "$episode_alerted" -eq 0 ] || rm -f "$marker" || return 1 + continue + fi + [ "$episode_alerted" -eq 0 ] || continue + idle=$((now - observed_at)) + [ "$idle" -ge "$threshold" ] || continue + ! secondmate_in_active_turn "$task" "$(fm_backend_target_of_meta "$meta")" || continue + receipt="$receipt_dir/$row_key" + if [ "$(cat "$receipt" 2>/dev/null || true)" = "$row_key" ]; then + fm_wake_secondmate_stall_marker_write "$task" "$row_key" || return 1 + continue fi - [ "$(cat "$marker" 2>/dev/null || true)" = "$row_key" ] && continue - [ "$(cat "$receipt" 2>/dev/null || true)" = "$row_key" ] && continue notify_key="secondmate-wake-loop-$task-$row_key" - reason="check: secondmate wake-loop stalled: mate=$task row=$seq age=${age}s" + reason="check: secondmate wake-loop stalled: mate=$task row=$seq idle=${idle}s" queued=$(fm_wake_queued_keys check) if ! printf '%s\n' "$queued" | grep -Fx "$notify_key" >/dev/null 2>&1; then fm_wake_append check "$notify_key" "$reason" || return 1 @@ -972,9 +1076,9 @@ pause_state_class() { # <window> <task> # ordinary crew whose agent the gate above confirmed dead, so no live decision gate # is being silenced, or a secondmate, whose endpoint liveness is deliberately never # read and so cannot supply that confirmation. Without the mate case a mate's - # captain hold - which has no current-state mapping and so arrives as `none` - - # would be silenced by every caller rather than taking the bounded re-surface - # cadence, and a forgotten hold would rot invisibly. + # status-declared `captain-held` transfer - which has no current-state mapping + # and so arrives as `none` - would be silenced by every caller rather than taking + # the bounded re-surface cadence, and a forgotten declaration would rot invisibly. [ "$class" = none ] && class=paused case "$class" in paused) date +%s > "$recheck_file" ;; @@ -983,6 +1087,99 @@ pause_state_class() { # <window> <task> printf '%s' "$class" } +# The two records of one ordinary crew wait, and why its stale alarm reads both. +# +# status_is_paused_or_captain_held reads the status LINE a worker wrote, which is +# the only record when the worker itself is waiting. It is not the only record +# there is: once firstmate hands work to the captain, the wait is written into the +# BACKLOG by bin/fm-captain-hold.sh, and the worker's last line stays whatever it +# was - routinely `done: PR ...` after a delivery, which no line predicate can +# read as a wait. An alarm bounded only by the line therefore re-fires for the +# captain's whole thinking time, on exactly the work they already have in hand. +# +# `open` is that record's own read-only predicate and owns its semantics: exit 0 +# still an open captain call, 1 not, 2 could not be established. Only a 0 bounds +# an alarm here, so an unreadable backlog, an incompatible or absent tasks-axi, +# and a row this home does not carry all keep alarming exactly as they do today - +# a wait this watcher cannot prove is not a wait. +# +# The read costs one subprocess and runs only where the watcher is about to +# alarm, so at most once per distinct stale hash per window, beside the crew-state +# read the same paths already pay. The secondmate stale gate deliberately runs +# before this bound and admits only status-declared waits: a backlog-only hold +# whose mate still says `working:` or `done:` does not reach this read. Reaching +# it would put backlog reads into windows deliberately skipped on ordinary polls. +STALE_WAIT_DECLARATION= + +CAPTAIN_CALL_IDENTITY= + +task_captain_call_open() { # <task> + local task=$1 + CAPTAIN_CALL_IDENTITY= + [ -n "$task" ] || return 1 + CAPTAIN_CALL_IDENTITY=$(FM_HOME="$FM_HOME" "$SCRIPT_DIR/fm-captain-hold.sh" \ + open "$task" --identity 2>/dev/null) || return 1 + return 0 +} + +# The identity a re-surface throttle is bound to: the task's whole status-log +# signature. Any new status event - a replacement wait, a fresh delivery, a +# blocker - changes it and so starts its own window instead of inheriting the +# silence of the one before it. +stale_wait_declaration() { # <task> + printf 'declared:%s' "$(fm_wake_signal_sig "$STATE/$1.status" || true)" +} + +# The same scope for a captain call, carrying the CALL's own lifecycle identity +# beside the status signature. The status log is not enough on its own: a task +# can be answered with `--release` and held again as a genuinely different call +# without any status append, and binding the throttle to the signature alone let +# the second call inherit the first one's silence and absorbed its first sight. +# That first sight is the one alarm this bound must never swallow - a decision +# waiting on the captain that is never surfaced is invisible, where a delivery +# announced twice is merely noise. +captain_call_declaration() { # <task> <call-identity> + printf 'captain-hold:%s:%s' "$2" "$(fm_wake_signal_sig "$STATE/$1.status" || true)" +} + +# 0 when <declaration> has already been alarmed for this window inside the +# current PAUSE_RESURFACE_SECS. A pure read: recording an alarm is the caller's, +# so the throttle is never advanced by a sighting it just absorbed. +stale_wait_throttled() { # <window-key> <declaration> + local throttle="$STATE/.paused-resurfaced-$1" + [ "$(cat "$throttle" 2>/dev/null || true)" = "$2" ] \ + && [ "$(age_of "$throttle")" -lt "$PAUSE_RESURFACE_SECS" ] +} + +# The same bound, for a stale window whose last line IS captain-relevant. That +# line is real and its first sight must still reach the captain, but a delivery +# they are already holding has nothing new to say on the next pane tick. +# Sets STALE_WAIT_DECLARATION to the scope this sighting is bound to, and leaves +# it EMPTY when no open captain call bounds it, so an unheld delivery, a blocker, +# and a failure alarm exactly as they do today. +# Returns 0 to absorb this sighting; 1 to alarm, after which the caller records +# the throttle through stale_wait_record once its own wake append has succeeded. +# Record a fired wake against the bounded cadence, and ONLY after that wake was +# durably appended. A marker written ahead of the append outlives a failed one: +# the watcher exits with no wake queued, and the next sighting reads the fresh +# marker and absorbs the retry, which is the single way this bound could swallow +# an alarm outright rather than delay it. +stale_wait_record() { # <window-key> + [ -n "$STALE_WAIT_DECLARATION" ] || return 0 + printf '%s' "$STALE_WAIT_DECLARATION" > "$STATE/.paused-resurfaced-$1" +} + +# Bound a due stale alarm for an ordinary crew task held for the captain. +# Backlog-only secondmate holds are outside this guard because the earlier gate +# preserves their no-backlog-read hot path. +captain_call_stale_bound() { # <window-key> <task> + local key=$1 task=$2 + STALE_WAIT_DECLARATION= + task_captain_call_open "$task" || return 1 + STALE_WAIT_DECLARATION=$(captain_call_declaration "$task" "$CAPTAIN_CALL_IDENTITY") + stale_wait_throttled "$key" "$STALE_WAIT_DECLARATION" +} + # Surface a stale pane no classifier could resolve, so firstmate inspects it: it # may have finished through an interactive menu that wrote no status, be waiting on # a decision, or be wedged. pause_state_class deliberately answers `none` for a @@ -990,30 +1187,38 @@ pause_state_class() { # <window> <task> # decision is never silenced - which routes every parked-but-live worker here, on # first sight of each distinct stale hash. # -# So a declared wait bounds this path to the same once-per-PAUSE_RESURFACE_SECS +# So a legitimate wait bounds this path to the same once-per-PAUSE_RESURFACE_SECS # cadence resurface_absorbed owns for the absorbed paths, throttled by this # window's own .paused-resurfaced-<key> marker: an idle parked pane still churns # its hash (a clock, a token counter), and each new hash re-enters this path, so -# without that bound one declared wait re-alarms firstmate for its whole duration. +# without that bound one wait re-alarms firstmate for its whole duration. # The FIRST sight still wakes, keeping the inspect-an-inconclusive-state intent, # and the throttle is read BEFORE anything is queued and advanced only by a wake # that really fires - a throttle written by the wake it should have prevented, or # read after that wake was already appended, bounds nothing. +# Both records of an ordinary crew wait bound it (see task_captain_call_open +# above): the status line the worker declared, and the backlog hold firstmate +# recorded once the captain took the work in hand. surface_nonterminal_stale() { # <window> <hash> - local win=$1 h=$2 key task last declaration='' declared=1 throttled=1 + local win=$1 h=$2 key task last declared=1 bounded=1 throttled=1 key=$(window_key "$win") task=$(window_to_task "$win" "$STATE") last=$(last_status_line "$STATE/$task.status") + STALE_WAIT_DECLARATION= if status_is_paused_or_captain_held "$last"; then declared=0 - declaration="declared:$(fm_wake_signal_sig "$STATE/$task.status" || true)" - if [ "$(cat "$STATE/.paused-resurfaced-$key" 2>/dev/null || true)" = "$declaration" ] \ - && [ "$(age_of "$STATE/.paused-resurfaced-$key")" -lt "$PAUSE_RESURFACE_SECS" ]; then - throttled=0 - fi + bounded=0 + STALE_WAIT_DECLARATION=$(stale_wait_declaration "$task") + stale_wait_throttled "$key" "$STALE_WAIT_DECLARATION" && throttled=0 + elif captain_call_stale_bound "$key" "$task"; then + bounded=0 + throttled=0 + elif [ -n "$STALE_WAIT_DECLARATION" ]; then + bounded=0 fi if [ "$throttled" -ne 0 ]; then fm_wake_append stale "$win" "stale: $win" || exit 1 + stale_wait_record "$key" fi printf '%s' "$h" > "$STATE/.stale-$key" rm -f "$STATE/.stale-since-$key" @@ -1021,12 +1226,19 @@ surface_nonterminal_stale() { # <window> <hash> if [ "$declared" -eq 0 ]; then : > "$STATE/.paused-$key" date +%s > "$STATE/.paused-rechecked-$key" - [ "$throttled" -eq 0 ] || printf '%s' "$declaration" > "$STATE/.paused-resurfaced-$key" + elif [ "$bounded" -eq 0 ]; then + # A backlog hold is NOT a declared pause, and must not be dressed up as one: + # the loop-top reconciliation and pause_state_class both read the status LINE, + # so a .paused-* flag this line does not support would be cleared on the next + # poll - taking the throttle with it - and would hand the mate and dead-agent + # cadences a declaration they were never given. Only the shared re-surface + # marker is kept, which is the whole of what this bound needs. + rm -f "$STATE/.paused-$key" "$STATE/.paused-rechecked-$key" else clear_pause_state "$key" fi if [ "$throttled" -eq 0 ]; then - triage_log "absorbed non-terminal stale (declared wait already re-surfaced this window): $win" + triage_log "absorbed non-terminal stale (declared wait or open captain call already re-surfaced this window): $win" return 0 fi wake "stale: $win" @@ -1211,18 +1423,26 @@ run_check_capture() { # hiding the `needs-decision`, `blocked`, `failed`, or `done` event that arrived # just before it: the .seen-* marker advances either way, so an event absorbed # here is never re-read. Non-.status arguments (.turn-ended markers, which carry -# no verb) are skipped. A 1 here is NOT "benign" on its own: a no-verb signal -# still needs the authoritative working proof or the eligible opt-in bare -# turn-end pane-churn proof before it is benign. +# no verb) are skipped. A 1 here is NOT "benign" on its own: a no-verb signal, +# including a newly declared captain hold, still needs the authoritative working +# proof or the eligible opt-in bare turn-end pane-churn proof before it is benign. +# Also populates FM_SIGNAL_NEEDS_DECISION_FILES (space-separated status-file +# paths) with exactly the files whose newly classified span carries one of the +# decision-owned classes defined by the status-span contract, so the caller can +# route those - and only those - signal rows as main-only +# (docs/pi-supervision-branch.md). Stale and heartbeat rows retain their existing +# eligibility rules. signal_files_actionable() { # <status-file> ... - local f task record rest endpoint ident rc found=1 + local f task record rest endpoint ident needs_decision rc found=1 FM_SIGNAL_SURFACE_ENDPOINTS='' + FM_SIGNAL_NEEDS_DECISION_FILES='' for f in "$@"; do case "$f" in *.status) ;; *) continue ;; esac [ -e "$f" ] || [ -L "$f" ] || continue task=$(basename "$f"); task="${task%.status}" - record=$(status_span_first_actionable_record "$f" \ - "$(fm_wake_signal_seen_size "$STATE" "$f")") + record=''; needs_decision=0 + status_span_first_actionable_record "$f" \ + "$(fm_wake_signal_seen_size "$STATE" "$f")" record needs_decision rc=$? [ "$rc" -eq 1 ] && [ -z "$record" ] && continue if [ "$rc" -eq 2 ]; then @@ -1235,7 +1455,12 @@ signal_files_actionable() { # <status-file> ... fi endpoint=${record%%$'\t'*}; rest=${record#*$'\t'}; ident=${rest%%$'\t'*} FM_SIGNAL_SURFACE_ENDPOINTS="${FM_SIGNAL_SURFACE_ENDPOINTS}${f}"$'\t'"${endpoint}"$'\t'"${ident}"$'\n' - [ "$rc" -eq 0 ] && found=0 + if [ "$needs_decision" -eq 1 ]; then + FM_SIGNAL_NEEDS_DECISION_FILES="${FM_SIGNAL_NEEDS_DECISION_FILES} ${f}" + fi + if [ "$rc" -eq 0 ] || [ "$needs_decision" -eq 1 ]; then + found=0 + fi done return "$found" } @@ -1737,15 +1962,27 @@ EOF # ordering evaluates them ONLY for a non-afk signal with no captain-relevant # status span, and the capture only once the authoritative verdict comes up short. FM_SIGNAL_SURFACE_ENDPOINTS='' + FM_SIGNAL_NEEDS_DECISION_FILES='' # shellcheck disable=SC2086 # $files is a space-separated status-path list (ids carry no spaces) signal_files_actionable $files signal_actionable=$? + # A decision-owned file's queued row payload is marked "needs-decision:" + # instead of the ordinary "signal:" below (other files in the same batch + # keep the ordinary payload). The wake reason line itself, and every + # harness-arm consumer that pattern-matches it, stays byte-identical - + # only the per-row payload changes, which is what + # docs/pi-supervision-branch.md's Pi-only branch dispatcher reads to keep a + # decision-owned row off the supervision branch (fm-branch-dispatch.ts, + # fm-primary-pi-watch.ts). Every other harness and script keeps seeing the + # exact same "signal:$files" wake it always has. # shellcheck disable=SC2086 # same space-separated status-path list if afk_present || [ "$signal_actionable" -eq 0 ] \ || { ! signal_crew_provably_working $files && ! signal_turnend_panes_churned $files; }; then while IFS=$(printf '\t') read -r sf sig f; do [ -n "$sf" ] || continue - fm_wake_append signal "$(basename "$f")" "$reason" || exit 1 + file_reason="$reason" + case " $FM_SIGNAL_NEEDS_DECISION_FILES " in *" $f "*) file_reason="needs-decision:$files" ;; esac + fm_wake_append signal "$(basename "$f")" "$file_reason" || exit 1 done <<EOF $pending EOF @@ -1819,12 +2056,12 @@ EOF clear_pause_tracking "$key" fi # An idle secondmate endpoint is healthy by design, so a mate is admitted to - # the pane-stale path ONLY to serve a declared wait's bounded re-surface - - # the same declarations pause_state_class reconciles below, which is why this - # gate reads the shared predicate rather than the pause verb alone. Narrowing - # it to `paused` would leave a mate's captain hold rotting invisibly: the - # clear above already spares its pause tracking, but nothing would ever - # re-surface it. + # the pane-stale path ONLY to serve a status-declared wait's bounded + # re-surface. This gate reads the shared predicate rather than the pause verb + # alone so it includes a declared `captain-held` status. A hold recorded only + # in the backlog while the mate still says `working:` or `done:` is outside + # this guard: reaching it would require backlog reads for windows this gate + # deliberately skips, putting that read on the ordinary poll hot path. if [ "$kind" = secondmate ] && ! status_is_paused_or_captain_held "$last"; then continue fi @@ -1882,8 +2119,21 @@ EOF date +%s > "$ssf" clear_write_tracking "$key" triage_log "absorbed stale (provably working, overriding a stale captain-relevant status): $w" + elif captain_call_stale_bound "$key" "$task"; then + # The line is captain-relevant and stays so, but the backlog says + # the captain already holds this work: further NEW pane hashes with + # the same status-log state have nothing to add while they are + # deciding. Only that new-hash repetition is bounded - the first + # sight already alarmed, a new hash inside the window is absorbed, + # and a new hash after it alarms again. A stable hash stays as inert + # here as it already was after a first terminal alarm. + printf '%s' "$h" > "$sf" + rm -f "$ssf" + clear_write_tracking "$key" + triage_log "absorbed stale (open captain call already surfaced for this status): $w" else fm_wake_append stale "$w" "stale: $w" || exit 1 + stale_wait_record "$key" printf '%s' "$h" > "$sf" rm -f "$ssf" clear_write_tracking "$key" diff --git a/bin/fm-x-lib.sh b/bin/fm-x-lib.sh index d3a2774b1dc..85449c19d87 100644 --- a/bin/fm-x-lib.sh +++ b/bin/fm-x-lib.sh @@ -90,8 +90,8 @@ fmx_single_link_file_valid() { local file=$1 expected_device=${2-} links device [ -f "$file" ] && [ ! -L "$file" ] || return 1 if [ "$(uname)" = Darwin ]; then - links=$(stat -f %l "$file" 2>/dev/null) || return 1 - device=$(stat -f %d "$file" 2>/dev/null) || return 1 + links=$(/usr/bin/stat -f %l "$file" 2>/dev/null) || return 1 + device=$(/usr/bin/stat -f %d "$file" 2>/dev/null) || return 1 else links=$(stat -c %h "$file" 2>/dev/null) || return 1 device=$(stat -c %d "$file" 2>/dev/null) || return 1 @@ -241,7 +241,7 @@ fmx_single_link_file_mode_valid() { return fi if [ "$_FM_X_UNAME" = Darwin ]; then - mode=$(stat -f %Lp "$file" 2>/dev/null) || return 1 + mode=$(/usr/bin/stat -f %Lp "$file" 2>/dev/null) || return 1 else mode=$(stat -c %a "$file" 2>/dev/null) || return 1 fi @@ -252,8 +252,8 @@ fmx_private_artifact_dir_device() { local dir=$1 mode device [ -d "$dir" ] && [ ! -L "$dir" ] || return 1 if [ "$_FM_X_UNAME" = Darwin ]; then - mode=$(stat -f %Lp "$dir" 2>/dev/null) || return 1 - device=$(stat -f %d "$dir" 2>/dev/null) || return 1 + mode=$(/usr/bin/stat -f %Lp "$dir" 2>/dev/null) || return 1 + device=$(/usr/bin/stat -f %d "$dir" 2>/dev/null) || return 1 else mode=$(stat -c %a "$dir" 2>/dev/null) || return 1 device=$(stat -c %d "$dir" 2>/dev/null) || return 1 @@ -554,7 +554,7 @@ fmx_request_relay_context() { fmx_context_registry_mtime() { local file=$1 mtime - mtime=$(stat -f '%m' "$file" 2>/dev/null) || mtime=$(stat -c '%Y' "$file" 2>/dev/null) || return 1 + mtime=$(/usr/bin/stat -f '%m' "$file" 2>/dev/null) || mtime=$(stat -c '%Y' "$file" 2>/dev/null) || return 1 case "$mtime" in ''|*[!0-9]*) return 1 ;; esac diff --git a/docs/agent-control.md b/docs/agent-control.md index bd5a01f89ea..fd1f24e7b68 100644 --- a/docs/agent-control.md +++ b/docs/agent-control.md @@ -48,7 +48,7 @@ The clear is refused before anything is sent when the recorded backend cannot de Removing a worktree, closing an endpoint, or discarding work stays with [`bin/fm-teardown.sh`](../bin/fm-teardown.sh), which owns the landed-work test. **`resume` is not a verb.** -It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, copilot, pi, pi-signed, and kimi have no verified pane-resume contract. +It is not deterministic across the verified adapters: codex, grok, and gemini resume only from a session id printed at exit, opencode continues the most recent session for the cwd, and claude, copilot, pi, pi-signed, omp, and kimi have no verified pane-resume contract. `relaunch` covers the same need on every adapter, because the brief on disk - not a harness-private session - is the durable instruction. ## Transactional relaunch diff --git a/docs/architecture.md b/docs/architecture.md index 2ed4a6764f9..984a76931c2 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -4,12 +4,18 @@ How firstmate works, in depth. The [README](../README.md) carries the high-level diagram and a short synopsis. This document expands every part of it. -firstmate's always-loaded operating contract and routing index for conditional procedures is [`AGENTS.md`](../AGENTS.md); this is the human-facing companion. +firstmate's supervisor contract and routing index for conditional procedures is [`AGENTS.md`](../AGENTS.md); this is the human-facing companion. ## Event-driven supervision A zero-token bash watcher (`bin/fm-watch.sh`) sleeps on the fleet, classifies detected wakes in bash, and wakes the first mate only when something is actionable. Actionable wakes include captain-relevant status signals, no-verb signals without positive evidence that their crew is still executing, authenticated check output such as PR merge polling or a Relay mention, stale panes whose crew is not provably working whether their status log looks terminal or non-terminal, provably-working stale panes that persist past `FM_STALE_ESCALATE_SECS` without their own task worktree being written, declared external waits and verified captain-held transfers that remain declared past `FM_PAUSE_RESURFACE_SECS`, and heartbeat backstop hits. +For an ordinary crew task, a wait is read from both of its records: the status line a worker declared, and the backlog hold `bin/fm-captain-hold.sh` recorded once firstmate handed the work to the captain. +So a delivered ordinary crew task whose last line stays `done: PR ...` bounds repeated alarms from new pane hashes to the `FM_PAUSE_RESURFACE_SECS` cadence for the length of the captain's decision. +The first hash still alarms, each new hash inside that window is absorbed, and a new hash after the window re-surfaces the hold; a terminal pane hash that never changes stays inert after its first alarm exactly as it did before this bound. +The throttle is scoped to both the current captain-call lifecycle and the status-log state, so releasing and re-holding the same task without a status append starts a fresh window whose first new hash alarms. +A secondmate reaches the stale path only for a wait declared in its status line, so a hold recorded only in the backlog while its last line is `working:` or `done:` is outside this guard. +Reaching that case would require consulting the backlog for windows the secondmate gate deliberately skips, putting backlog reads on the ordinary poll hot path this design preserves. Repeated provably-working stale escalations on the same unchanged pane add an escalation count to the wake reason and, at `FM_WEDGE_DEMAND_INSPECT_COUNT`, a `demand-deep-inspection` marker. A pane holding a file newer than the start of its own quiet window, anywhere in the worktree recorded for that task, is deferred instead of escalated, because a crew writing source, then tests, then documentation behind a static pane is liveness that neither pane quietness nor the run step can show. That deferral re-surfaces on the same `FM_PAUSE_RESURFACE_SECS` cadence as a declared wait, with a reason naming the write evidence rather than a wedge, and it is bounded to one pruned, depth-bounded, wall-clock-bounded walk (`FM_WORKTREE_WRITE_PRUNE`, `FM_WORKTREE_WRITE_MAXDEPTH`, `FM_WORKTREE_WRITE_TIMEOUT`) taken only in the branch that was about to escalate, never on every poll. @@ -21,10 +27,11 @@ Lifting the declaration restores the unchanged busy-pane wedge path, while a pan While away mode is active, a busy pane that crosses the bound under a declared wait is handed to the daemon as the plain wake identity instead of taking that recheck in the watcher, because the daemon owns triage there and a wake already decorated as a possible wedge would override the daemon's own declared-wait verdict; an undeclared busy pane past the bound still takes the wedge escalation in away mode. That handoff is keyed on the declaration itself (the status log's signature) rather than on the pane capture, so a harness footer that ticks on every poll wakes the daemon once per declaration instead of once per poll, and it clears the wedge timer, escalation count, and worktree-write deferral exactly as the normal-mode absorber does, so an undeclared busy phase's timer does not resume when the declaration lifts. Those actionable wakes are written to a durable local queue (`state/.wake-queue`) only after generation-bound recovery evidence is published, so an interrupted watcher or handling turn can be recovered without losing the queue record. -Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. -Once that row reaches `FM_SECONDMATE_WAKE_STALL_SECS`, the primary appends one keyed `check` wake naming the mate, row sequence, and observed age; parent receipts and queued-key deduplication suppress repeats for the same row across watcher and handling crashes, while empty and younger queues remain silent. +Agent endpoint liveness and queue-consumption liveness are separate: on each poll, the primary watcher reads the oldest valid actionable row from every endpoint-recorded local secondmate home's durable wake queue without locking, consuming, or rewriting that foreign queue. +A queue that is draining is not stalled, so the primary times the interval since that oldest actionable row last changed rather than the age of the row itself, and rows that declare themselves a bounded external wait (`awaiting external - declared pause`) are not actionable evidence at all. +Once that no-progress interval reaches `FM_SECONDMATE_WAKE_STALL_SECS` and the mate is not provably inside an active turn (an exact busy verdict, bounded by `FM_BUSY_TURN_MAX_SECS`), the primary appends one keyed `check` wake naming the mate, row sequence, and observed idle interval; parent receipts and queued-key deduplication suppress repeats across watcher and handling crashes, one notification covers a whole no-progress episode, and any move of that position - drain progress, or the fresh rows of a queue reprovisioned under the same task id, at whatever sequence it restarts - ends that episode and starts a fresh observation interval, while empty, advancing, and declared-wait queues remain silent. Endpointless registered mates remain outside this scan because startup secondmate-liveness owns dead or missing endpoint recovery, and remote homes retain their host-local supervision boundary. -`tests/fm-wake-queue.test.sh` pins the notification, idempotence, quiet-queue, and byte-for-byte foreign-row preservation guarantees. +`tests/fm-wake-queue.test.sh` pins the no-progress notification, drain-progress reset, declared-pause exclusion, active-turn deferral, idempotence, quiet-queue, and byte-for-byte foreign-row preservation guarantees. When a canonical validated PR poll returns exactly `merged`, the watcher routes it through the shared merge-outcome emitter before retiring the poll. [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns role routing, PR-specific wake identity, marker-locked normal deduplication, and the at-least-once ordering that prefers a rare duplicate over silence. After successful outcome publication, the watcher immediately delivers the emitter's local actionable poll row and publishes a private retirement receipt bound to the poll's registration, bytes, file identities, metadata, provider, URL, and task ID. @@ -45,7 +52,7 @@ A crew that declares `paused:` for a known external wait, or carries a verified For an ordinary crew that has stopped, the normal-mode watcher first surfaces one stale wake, then applies that same cadence to an unchanged `paused:` or durable `captain-held` endpoint; the pause classification itself is recovered only when the backend confidently reports its agent dead. Live or inconclusive liveness remains fail-open at that initial surface, so a worker genuinely waiting on a decision is never silenced. Its later sights are still held to that same bounded cadence rather than re-alarming on every pane-hash change, because the throttle is keyed to the declaration and not to the pane an idle parked worker keeps ticking. -A secondmate's endpoint liveness is still never read at all; a mate is admitted to that same cadence only to serve a declared wait's bounded re-surface, so a forgotten pause or captain hold on a mate cannot rot invisibly. +A secondmate's endpoint liveness is still never read at all; a mate is admitted to that same cadence only to serve a status-declared wait's bounded re-surface, so a forgotten `paused:` or `captain-held` declaration on a mate cannot rot invisibly. Its initial normal-mode status signal still surfaces through the no-verb path, while away mode self-handles that routine signal and owns the later recheck. Fresh stale panes use the same current-state read before trusting the status log, so an active run or a proven busy worker outranks an old captain-relevant status-log line left behind before validation. No-change heartbeats are also benign. @@ -68,11 +75,14 @@ The explicit resolution is written by the actor that answers, not the busy worke This home's answerer close, pending-reply escalation close, and captain-held transfer use the provenance-guarded append owned by `bin/fm-wake-lib.sh`, so they advance the watcher marker only across their own bytes when all earlier bytes were already announced; pending or interleaved foreign bytes fail toward an ordinary wake. A turn-ended-only queue row omits its historical status annotation when that status file exactly matches the same seen marker. Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. -`bin/fm-crew-state.sh <id>` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed. +`bin/fm-crew-state.sh <id>` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed, except that a `blocked:` event reporting a refused or missing daemon socket outranks a potentially stale active run record. +For other daemon, timeout, or unreachability claims, a running or fixing run with recent pipeline-reported activity supersedes the event and names reattachment as the recovery instead of surfacing a false block. [`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. A run head the task copy cannot resolve locally is attributed only when the pipeline's own runs ledger proves it is an active continuation of the submitted head, so a pipeline fix round never reads as an older failed run. During no-mistakes' `ci` monitor phase, it also reads the ci step log tail because `axi status` reports both "still waiting on checks" and "checks green, waiting on merge" as `ci,running`. The most recent recognized ci log marker wins, so checks-green monitoring reports done while a later re-arm, failed-check, or issue marker returns the crew to working. +A terminal failed run whose only failure is the ci monitor step, after every substantive step completed and the same marker reads checks green, also reports done with the run's PR URL, because a monitor whose only remaining job is to observe a human merge decision must not convert the absence of that decision into a failure verdict. +In the coarse runs-ledger fallback, which has no steps table and no ci log, a terminal failed record whose daemon an explicit `daemon status` probe proves down reports unknown as unverified instead: an instrument failure must never read as work failure. Only when no matching run exists does it consult semantic busy state; exact busy reports working, exact idle permits fallback to a status-log event whose verb maps to a recognized run-state, and unknown or a dead pane stays unknown instead of trusting a stale log. Decision-only events such as `resolved` never become current state or leak their prose into the current-state detail. In that status-log fallback, a declared external wait reports the distinct `paused` state with its reason. @@ -103,24 +113,24 @@ Live GitHub enrichment exists only behind the bearings `--include-prs` opt-in. Optional Relay integrates with the watcher only after explicit opt-in; [configuration.md](configuration.md#relay-env) owns its generated-artifact and dispatch mechanics. At session start, `bin/fm-session-start.sh` emits exactly one primary-harness supervision block rendered by `bin/fm-supervision-instructions.sh` from `docs/supervision-protocols/`. -That block owns the live wait shape for the running primary harness: Claude's Stop `asyncRewake` hook owns tokenless re-arm cycles, Copilot uses a tracked asynchronous watcher with shell-completion notification re-entry, Cursor's stop hook parks on the watcher, Grok uses background-notify cycles, Codex uses bounded foreground checkpoints, Pi and pi-signed use the same two tracked primary extensions, and OpenCode uses its TUI plugin. +That block owns the live wait shape for the running primary harness: Claude's Stop `asyncRewake` hook owns tokenless re-arm cycles, Copilot uses a tracked asynchronous watcher with shell-completion notification re-entry, Cursor's stop hook parks on the watcher, Grok uses background-notify cycles, Codex uses bounded foreground checkpoints, Pi and pi-signed use the same two tracked primary extensions, omp uses its own two tracked `.omp/extensions/` files, and OpenCode uses its TUI plugin. `bin/fm-watch-arm.sh` remains the verified arm wrapper for protocols that call it; it forks the watcher as a tracked child, verifies it is genuinely alive with a fresh liveness beacon, and prints an honest `started`, `attached`, or nonzero `FAILED` status. [`watcher-continuity.md`](watcher-continuity.md#arm-layer-cycle-contract) owns the arm layer's successor, terminal-delivery, re-arm recovery, and typed clean-close failure contract. The arm layer records one bounded lifecycle row per observed cycle in `state/.watch-cycle-exits.log`; `state/.watch-triage.log` remains exclusively the absorbed-wake debug log. -Pi and OpenCode verify session-lock ownership and launch one singleton successor from their child-close handlers before delivering an actionable wake prompt, with bounded exponential retry for failed restoration. +Pi, omp, and OpenCode verify session-lock ownership and launch one singleton successor from their child-close handlers before delivering an actionable wake prompt, with bounded exponential retry for failed restoration. Claude's `bin/fm-claude-stop-autoarm.sh` hook fires on every Stop and, when the home is eligible and still needs supervision, claims one home-scoped cycle, foregrounds the arm wrapper, and translates actionable closes into exit-2 rewakes. It suppresses failed-looking closes when the same identity-matched watcher is healthy, retries genuine failures within a bound, and coordinates exhausted failure episodes with the Claude turn-end guard as documented in [`turnend-guard.md`](turnend-guard.md). [`watcher-continuity.md`](watcher-continuity.md) owns Claude's residual active-turn coverage and watcher-status command-gating boundary. Cursor's `bin/fm-turnend-guard-cursor.sh` hook is the same between-turns shape in one synchronous step: it parks the awaited `stop` hook on the arm wrapper and translates an actionable close into one `followup_message`, with a generation baton that makes an older park still running after the next `stop` claim stand down instead of leaking a stale duplicate wake. Copilot starts the same watcher owner through its shell tool's native asynchronous mode, using `bin/fm-watch-arm.ps1` as the native Windows bridge. Its tracked shell-completion notification hook converts an actionable close into typed Firstmate input, while `bin/fm-copilot-stop.sh` keeps a session-scoped seven-continuation repair ledger below Copilot CLI's hard eight-block override without launching or waiting for the watcher. -The existing turn-end guard remains the final backstop for every harness-engine protocol, with pi-signed sharing Pi's protocol, the `--claude` mode cooperating with the auto-arm claim, Cursor's `--cursor` mode rendering a bounded follow-up, and Copilot using its dedicated nonblocking stop backstop. +The existing turn-end guard remains the final backstop for every harness-engine protocol, with pi-signed sharing Pi's protocol, omp's blocking `session_stop` hook compelling one continuation per turn, the `--claude` mode cooperating with the auto-arm claim, Cursor's `--cursor` mode rendering a bounded follow-up because its `stop` step cannot be blocked, and Copilot using its dedicated nonblocking stop backstop. Its `--restart` mode signals only the watcher recorded in the current home's `state/.watch.lock`, so restarting one home cannot kill sibling secondmate watchers. -A pull-based guard (`bin/fm-guard.sh`) warns through supervision tool output if the primary checkout is tangled or if work, process-event sources, or Relay polling has an unhealthy model-aware supervision verdict; on main it also warns when queued wakes are waiting to be drained. +A pull-based guard (`bin/fm-guard.sh`) warns through supervision tool output if the primary checkout is tangled or if work, process-event sources, registered custom checks, or Relay polling has an unhealthy model-aware supervision verdict; on main it also warns when queued wakes are waiting for main itself to drain. The drain script calls that guard after presenting the queue; records remain durable until the exact generation-bound acknowledgement printed by the drain succeeds after handling, and main may keep the queued-wakes warning visible until then. -The Pi supervision branch's deliberate queued-wake warning exception is owned by [`pi-supervision-branch.md`](pi-supervision-branch.md#components-and-their-owners). +The Pi supervision branch's deliberate queued-wake warning exception is owned by [`pi-supervision-branch.md`](pi-supervision-branch.md#components-and-their-owners), while [`watcher-continuity.md`](watcher-continuity.md#per-actor-acknowledgement) owns the guard's per-actor counting, the advisory main gets for rows a live branch grant holds, and main's retirement of queue rows no actor could ever present or acknowledge. It leads with a prominent bordered tangle banner, while `bin/fm-guard.sh` owns the watcher-down banner and reminder policy so repeated guarded commands stay noisy without reprinting the full banner in the same episode. -On every verified primary harness, tracked hook integration gives the primary session a push-based backstop: when work, a process-event source, or Relay polling needs supervision and no supervision owner provably holds this home with a fresh beacon, blocking-capable Stop hooks block and nonblocking turn-end integrations force one bounded follow-up. +On every verified primary harness, tracked hook integration gives the primary session a push-based backstop: when work, a process-event source, a registered custom check, or Relay polling needs supervision and no supervision owner provably holds this home with a fresh beacon, blocking-capable Stop hooks block and nonblocking turn-end integrations force one bounded follow-up. The guard covers the main primary and genuinely marked secondmate homes, exempts child crewmate/scout worktrees, is loop-safe per harness, and is documented in [turnend-guard.md](turnend-guard.md). A presence-gated sub-supervisor (`bin/fm-supervise-daemon.sh`) extends this for walk-away supervision: the `/afk` skill starts it through the tracked foreground helper `bin/fm-afk-start.sh`, after which the watcher reverts to daemon-managed one-shot mode and the daemon self-handles routine wakes in bash. @@ -147,6 +157,7 @@ Unsupported supervisor backends refuse at daemon startup. Stalled escalation delivery writes `state/.subsuper-inject-wedged` and attempts a configured backend-independent active alert after `FM_MAX_DEFER_SECS` instead of silently deferring forever. On an unmarked return, `bin/fm-afk-return.sh` owns ordered shutdown, durable catch-up evidence, and the fail-closed gate that keeps ordinary work behind every live firstmate-actionable blocker. `fm-send.sh` delivers every remote text steer and ordinary local text steer as a durable steering-inbox record plus a best-effort constant doorbell line (`bin/fm-task-inbox-lib.sh`). +The doorbell line is a shell no-op and is never typed into an endpoint classified as dead or missing; that record surfaces once for recovery instead of walking the re-ring ladder (`bin/fm-task-inbox-lib.sh` header). Its local-only typed plane - harness-native invocations and explicit backend targets - selects a pre-Enter popup-settle for slash commands and for codex `$...` skill invocations using metadata-routed target `harness=` values, then adds its own `FM_SEND_SETTLE` pause after successful typed sends so immediate peeks catch the receiving turn starting; the sub-supervisor uses only the shared submit core and does not pay that post-submit pause. Text for a worker to read and commands that drive a worker's process are separate planes. @@ -159,7 +170,7 @@ Text for a worker to read and commands that drive a worker's process are separat `bin/fm-busy-lib.sh` is the single owner of what "this worker is busy" means, and `bin/fm-busy-event.sh` is the only writer of the per-task records it reads. Every classification returns a verdict of busy, idle, unknown, or dead together with the source that produced it, so a consumer or a diagnostic can never confuse semantic state with a fallback. -Each converted adapter reports its own turn lifecycle through a machine-readable contract the vendor already exposes, rather than through rendered footer text: Pi and pi-signed through the Firstmate-owned extension's `agent_start` and `agent_settled` confirmed by `ctx.isIdle()`, OpenCode through its plugin's semantic `session.status`, Claude through owned `UserPromptSubmit`, `Stop`, `StopFailure`, and `SessionEnd` hooks, Muse through its session log, and Cursor through its conversation transcript. +Each converted adapter reports its own turn lifecycle through a machine-readable contract the vendor already exposes, rather than through rendered footer text: Pi and pi-signed through the Firstmate-owned extension's `agent_start` and `agent_settled` confirmed by `ctx.isIdle()`, omp through its extension's `agent_start` and `agent_end` without `willContinue`, OpenCode through its plugin's semantic `session.status`, Claude through owned `UserPromptSubmit`, `Stop`, `StopFailure`, and `SessionEnd` hooks, Muse through its session log, and Cursor through its conversation transcript. Kimi behind Pi inherits Pi's lifecycle. Codex and standalone Kimi classify unknown behind explicit probes until a semantic source is live-verified for them, and Grok keeps one clearly isolated rendered-tail fallback that can only ever classify a Grok task. @@ -176,7 +187,7 @@ All are harness-scoped rather than a global pattern union, and none is a recorde The runtime backend is the session-provider layer below firstmate's scripts. It owns task endpoint creation, bounded capture, text/key sends, current-path reads for spawn-time worktree discovery when the backend does not create the worktree itself, live-window fallback lookup, agent-process liveness probes where verified, and endpoint teardown. -`bin/fm-backend.sh` centralizes backend selection, `state/<id>.meta` helpers, metadata-only cleanup identity validation, selector resolution, and operation dispatch; `bin/backends/tmux.sh` is the verified reference adapter ([`docs/tmux-backend.md`](tmux-backend.md)), and `bin/backends/herdr.sh` (P2), `bin/backends/zellij.sh` (P3), `bin/backends/orca.sh` (P4), and `bin/backends/cmux.sh` (P5) are experimental task-spawn adapters. +`bin/fm-backend.sh` centralizes backend selection, `state/<id>.meta` helpers, metadata-only cleanup identity validation, selector resolution, and operation dispatch; `bin/backends/tmux.sh` is the verified reference adapter ([`docs/tmux-backend.md`](tmux-backend.md)), `bin/backends/herdr.sh` (P2) has its own required CI lane ([`docs/herdr-backend.md`](herdr-backend.md)), and `bin/backends/zellij.sh` (P3), `bin/backends/orca.sh` (P4), and `bin/backends/cmux.sh` (P5) remain experimental task-spawn adapters with no dedicated real-backend CI lane. [`configuration.md`](configuration.md#runtime-backend-configbackend--fm_backend) owns new-spawn backend selection precedence and authorization. Runtime auto-detection is innermost-first: `$TMUX` wins over `HERDR_ENV=1`, which wins over cmux's primary `CMUX_WORKSPACE_ID` marker and documented fallback signals; auto-detected herdr or cmux prints a one-time opt-out notice, auto-detected tmux stays silent, and zellij and orca are never auto-detected (only explicit selection). Unknown backend names fail loudly. @@ -187,7 +198,7 @@ tmux, zellij, orca, and cmux expose no native busy primitive at all, so a task o That poll loop is still the default event source for backends with no native push events, so this stays an extraction of the abstraction rather than a watcher rewrite. For capable Herdr sessions, the same watcher replaces its terminal sleep with a bounded native event wait that immediately surfaces `blocked`; [Push events and polling fallback](herdr-backend.md#push-events-and-polling-fallback) owns the current mechanism and capability gates, while [runtime backend verification](verification/runtime-backends.md#native-blocked-event) owns the active evidence. The deeper session-start agent-process liveness probe is separate from that busy-state poll: tmux and Herdr have verified classifiers for secondmate recovery, Zellij remains unverified, and Orca and cmux do not support secondmate spawns. -Herdr is experimental and can be selected explicitly or by runtime auto-detection: Treehouse remains its worktree provider, [`herdr-backend.md`](herdr-backend.md) owns current setup and safety limits, and [`verification/runtime-backends.md`](verification/runtime-backends.md#herdr) owns active empirical evidence. +Herdr can be selected explicitly or by runtime auto-detection: Treehouse remains its worktree provider, [`herdr-backend.md`](herdr-backend.md) owns current setup, CI coverage, and safety limits, and [`verification/runtime-backends.md`](verification/runtime-backends.md#herdr) owns active empirical evidence. Herdr uses one tab per task; [Watching and task containers](herdr-backend.md#watching-and-task-containers) owns launcher-bound workspace placement, the label-only fallback, and recovery scope. Its default-on presentation projection may place one clean new task in a disposable workspace without changing endpoint authority or lifecycle ownership; [Presentation spaces](herdr-backend.md#presentation-spaces) owns that conditional design, the Herdr version floor its unconfigured default is gated behind, and its narrow home-local restored-shell cleanup at locked session start. Zellij is experimental and selected only explicitly: Treehouse remains its worktree provider, [`zellij-backend.md`](zellij-backend.md) owns current setup and limits, and [`verification/runtime-backends.md`](verification/runtime-backends.md#zellij) owns active empirical evidence. @@ -213,6 +224,7 @@ Only a named non-default branch checked out in `FM_ROOT` is a worktree tangle. `fm-guard.sh` prints the repair command on the next mutable fleet action, while `bin/fm-session-start.sh` reports the same condition through bootstrap as a `TANGLE:` line at session start. If another live session holds the fleet lock, both surfaces keep the alarm but switch to read-only wording with no repair command. Ship briefs also tell the crewmate to verify `pwd -P` and `git rev-parse --show-toplevel` before creating `fm/<id>`, then stop with a blocked status if it landed in the primary checkout. +Placement is proven only at launch, so `bin/fm-spawn.sh` also exports the task id as `FM_TASK_ID` into every ship and scout pane, and `bin/fm-test-run.sh` refuses to execute the behavior suite from the primary checkout while that marker is set; the runner's header owns the predicate and [`tests/fm-test-run.test.sh`](../tests/fm-test-run.test.sh) pins it. ## No-mistakes gate authority boundary @@ -236,7 +248,7 @@ The session-start bootstrap step keeps valid dispatch configuration silent unles When the file exists, `fm-spawn.sh` refuses crewmate and scout launches without an explicit harness, so `config/crew-harness` is only automatic when no dispatch profile file is active. Secondmate launches are exempt because they resolve the secondmate harness and any optional secondmate model or effort tokens instead. Unsupported effort values are still recorded in task meta when passed to `fm-spawn.sh`, but the launch template omits any effort flag that the selected harness does not accept. -That keeps spawn launch compatible across claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, cursor, gemini, and muse while preserving the requested profile for later audit. +That keeps spawn launch compatible across claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, cursor, gemini, muse, rovo, and omp while preserving the requested profile for later audit. ## Optional secondmates @@ -277,7 +289,7 @@ For a local route, an explicit per-spawn harness or raw launch command does not Remote routes accept verified harness adapters only and reject raw launch commands. `config/crew-harness` remains the crewmate harness and is inherited into secondmate homes. `config/crew-dispatch.json` is inherited too; secondmates use the same natural-language dispatch profiles when spawning their own crewmates. -The [`secondmate-provisioning` skill](../.agents/skills/secondmate-provisioning/SKILL.md) owns the complete inherited-local-material allowlist and propagation contract. +The [`secondmate-provisioning` skill](../.agents/skills/secondmate-provisioning/SKILL.md) owns the inherited-local-material propagation contract and points to the implementation's item declaration. The `data/secondmates.md` line contract is owned by the [`secondmate-provisioning` skill](../.agents/skills/secondmate-provisioning/SKILL.md#routing-table), and the secondmate environment variables are documented in [configuration.md](configuration.md). @@ -308,8 +320,10 @@ Every GitHub refusal states what it could not observe as plainly as what it did, A confirmed merge leaves a durable role-routed outcome instead of living only in the merging agent's memory, and [`bin/fm-merge-outcome-lib.sh`](../bin/fm-merge-outcome-lib.sh)'s header owns its destination, shape, identity, normal-case deduplication, and at-least-once recovery. The same emitter handles a merge firstmate performed and one its poll detected, while the watcher immediately delivers the emitter's local actionable poll row. Teardown is fail-closed for ship worktrees: dirty worktrees refuse, and committed work must be landed before the worktree is returned. +A pool worktree is only returned after teardown passes the slot-ownership proof: a contradictory task record or supported live endpoint refuses without touching either task, and no discard authority relaxes that. +Allocation and return serialize on one project lock per machine-local Firstmate tree: every home reachable through local parent links shares that lock, and a home seeded from another machine anchors its own, because a lock taken on this filesystem is neither held nor observable across that boundary. Before the worktree is returned, teardown concludes the task's own no-mistakes run when it is parked at a gate, including a run whose head the task copy cannot resolve - the shared runs-ledger continuation proof is the only recognition for that case, so cleanup never orphans a parked run the pipeline advanced past the submitted head. -[`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, PR-discovery fallback, pre-teardown run conclusion, and stale-lock recovery procedure. +[`bin/fm-teardown.sh`](../bin/fm-teardown.sh)'s header owns the landed-work proofs, slot-ownership proof, PR-discovery fallback, pre-teardown run conclusion, and stale-lock recovery procedure; [`tests/fm-teardown-endpoint-safety.test.sh`](../tests/fm-teardown-endpoint-safety.test.sh) and [`tests/fm-secondmate-safety.test.sh`](../tests/fm-secondmate-safety.test.sh) pin the slot-collision boundary. ## Optional Relay diff --git a/docs/arm-pretool-check.md b/docs/arm-pretool-check.md index 5a5e8c8a2b8..835646c9fdd 100644 --- a/docs/arm-pretool-check.md +++ b/docs/arm-pretool-check.md @@ -25,7 +25,7 @@ It tokenizes the bytes and classifies lexical execution positions only. - Stdin JSON at `.tool_input.command` for Claude and Codex. - Stdin JSON at `.toolInput.command` for Grok. - Stdin JSON at `.tool_input.command` for GitHub Copilot CLI's Bash and Windows PowerShell tools, rendered through `--copilot` as a native `permissionDecision: "deny"` response. -- `--command <exact string>` for OpenCode, Pi, and pi-signed. +- `--command <exact string>` for OpenCode, Pi, pi-signed, and omp. - `--background` as a compatibility-only field that never changes the decision. - `--claude` to preserve Claude's stderr-only deny requirement. @@ -154,7 +154,7 @@ Prose may improve without changing adapter behavior. - `--claude` suppresses stdout completely because Claude ignores a PreToolUse deny when stdout is nonempty. - Codex blocks on exit 2 and displays stderr. - OpenCode throws only when the checker exits 2. -- Pi and pi-signed return `{block: true}` only when the checker exits 2. +- Pi, pi-signed, and omp return `{block: true}` only when the checker exits 2. ## Harness wiring @@ -166,6 +166,7 @@ Prose may improve without changing adapter behavior. | Grok | `.toolInput.command` | `.grok/hooks/fm-primary-pretool-check.json` forwards stdin and Grok consumes the stdout `decision=deny` object. | | OpenCode | `output.args.command` | `.opencode/plugins/fm-primary-pretool-check.js` passes one `--command` argument and throws only for exit 2. | | Pi / pi-signed | `event.input.command` | `.pi/extensions/fm-primary-turnend-guard.ts` passes one `--command` argument and returns `{block: true}` only for exit 2. | +| omp | `event.input.command` | `.omp/extensions/fm-primary-turnend-guard.ts` passes one `--command` argument and returns `{block: true, reason}` only for exit 2; omp surfaces the reason verbatim to the model (verified 18.1.2). | | Cursor | `.tool_input.command` | `.cursor/hooks.json` matches `tool_name` `Shell` and forwards stdin with `--cursor`. Cursor reads the RETURNED object rather than the exit status, so `--cursor` prints `{"permission":"deny","user_message":"[code] reason"}` on stdout and exits 0; only that rendering is verified to block the command and surface the reason. | Cursor also loads `<project>/.claude/settings.json`, so the tracked Claude entry receives the same event. Without `--cursor` a Cursor-delivered payload is that duplicate and allows without re-classifying, decided from the payload's own `cursor_version` by `bin/fm-hook-host-lib.sh`; [`turnend-guard.md`](turnend-guard.md#harness-integrations) owns why that predicate reads the payload rather than the environment. diff --git a/docs/calm-mode-feasibility.md b/docs/calm-mode-feasibility.md index 989e56254ce..288803e8f00 100644 --- a/docs/calm-mode-feasibility.md +++ b/docs/calm-mode-feasibility.md @@ -17,6 +17,7 @@ Pi 0.81.1 was installed when Calm was first built, and Pi 0.82.0 was the later r The inspected Pi CHANGELOG shows no relevant presentation API introduced at either version, so those versions remain verification evidence rather than compatibility bounds. The exported classes used by the adapters (`AssistantMessageComponent` and `InteractiveMode`) are undocumented internals with no stated version guarantee. `tests/fm-calm-pi-extension.test.sh` records the installed Pi version as evidence without gating on it and covers both newer synthetic versions and an unavailable adapter seam. +This host tracks Pi latest, so the version the evidence is pinned to moves; the [2026-09-07 record](#2026-09-07-pi-0851-renderer-and-export-dom-verification) owns the currently pinned version and the renderer comparison behind it. ### Built-in tool override constraints @@ -237,12 +238,12 @@ The test fixture enumerates every class below through the centralized policy, an | `system-notice` | `showStatus`, `showError`, compaction, retry, and startup warning rows | Unsupported boundary; remains visible. | | `cache-notice` | Non-persisted cache-miss `Text` row | Unsupported boundary; remains visible. | | `project-trust-warning` | Non-persisted startup `Text` row | Unsupported boundary; remains visible. | -| `synthetic-user` | Firstmate extension `sendUserMessage`, terminal-injected input, Firstmate-generated Pi positional brief, or the already non-displayed session-start nudge | Canonically classified text-only operational user messages stay ordinary semantic user messages but render through the zero-height adapter (verified on Pi 0.81.1 through 0.82.0) under Calm; legacy entries stay gaplessly controllable, and the session-start nudge retains its existing non-displayed custom-message path. | +| `synthetic-user` | Firstmate extension `sendUserMessage`, terminal-injected input, Firstmate-generated Pi positional brief, or the already non-displayed session-start nudge | Canonically classified text-only operational user messages stay ordinary semantic user messages but render through the zero-height adapter under Calm; legacy entries stay gaplessly controllable, and the session-start nudge retains its existing non-displayed custom-message path. | | `synthetic-assistant` | No authoritative Firstmate source found | Policy-hidden, but Pi exposes no generic assistant-role renderer. | | `unknown` | Future or unclassified transcript component | Policy-hidden, but no generic renderer exists; never claimed as covered. | The installed extension API has no supported global transcript filter, user-message renderer, assistant-message renderer, chat-container API, or generic custom-tool wrapper. -Pi 0.81.1 through 0.82.0 and Pi 0.84.4 export `AssistantMessageComponent` and `InteractiveMode`, so Calm uses separate idempotent, API-probed adapters for assistant thinking layout and the complete operational-user transcript row while leaving all message data and non-Calm rendering unchanged; see the [compatibility contract](calm.md#pi-compatibility) for how a future Pi lacking one of those exports is handled. +Pi 0.81.1 through 0.82.0, Pi 0.84.4, and Pi 0.85.1 export `AssistantMessageComponent` and `InteractiveMode`, so Calm uses separate idempotent, API-probed adapters for assistant thinking layout and the complete operational-user transcript row while leaving all message data and non-Calm rendering unchanged; see the [compatibility contract](calm.md#pi-compatibility) for how a future Pi lacking one of those exports is handled. General component replacement, ANSI cursor erasure, provider-context mutation, and installed-file patching remain rejected as unsupported or preservation-breaking workarounds. ## Cross-harness verification record @@ -286,7 +287,7 @@ The operational provider path covers Calm loaded on, loaded off, default prefere It asserts one persisted and rendered captain answer, exact user-role operational envelopes in order, no replacement custom messages, one processing result, zero operational transcript rows, and the two-row neighboring-assistant geometry for live, adjacent, and restart paths. Quoted current markers, ASCII-only labels, ordinary text before a marker, unrelated U+2063 placement, and image-bearing input remain visible in component and native transcript checks. `tests/fm-pi-primary-live-e2e.test.sh` also proves the working ship replaces the built-in `Working...` row while Calm is active on the credentialed provider path, and that it clears when the run settles, before continuing its ordinary watcher lifecycle. -`tests/fm-pi-primary-types.test.sh` performs strict no-emit TypeScript checking against the installed Pi declarations, currently package version 0.84.4. +`tests/fm-pi-primary-types.test.sh` performs strict no-emit TypeScript checking against whichever Pi declarations are installed, without pinning a version of its own. The relevant commands are: @@ -540,3 +541,71 @@ FM_TEST_END 2026-08-29T01:01:30Z tests/fm-pi-branch-extension.test.sh exit=0 dur ``` The real renderer comparison exercised twelve outcome lines and reported collapsed and expanded parity with Pi stock, zero visible rows under Calm, restored stock parity after toggling Calm off, and delegated stock HTML export fallback. + +## 2026-09-07 Pi 0.85.1 renderer and export-DOM verification + +This host tracks Pi latest, so the version this contract's evidence is pinned to moves. +The renderer and lifecycle evidence below was taken against installed `@earendil-works/pi-coding-agent` 0.85.1 with `@earendil-works/pi-server` 0.85.0 also installed globally. + +Calm's rendered rows are unchanged across 0.84.4, 0.85.0, and 0.85.1. +`FM_PI_PACKAGE_DIR` points `tests/fm-calm-pi-extension.test.sh` at an isolated install, so each comparison ran against its own temporary dependency tree and never mutated the globally installed packages. + +```text +$ pi --version +0.85.1 + +$ npm ls -g --depth 0 @earendil-works/pi-coding-agent @earendil-works/pi-server +├── @earendil-works/pi-coding-agent@0.85.1 +└── @earendil-works/pi-server@0.85.0 +``` + +```text +$ FM_PI_PACKAGE_DIR=<pi 0.84.4> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +$ FM_PI_PACKAGE_DIR=<pi 0.85.0> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +$ FM_PI_PACKAGE_DIR=<pi 0.85.1> tests/fm-calm-pi-extension.test.sh +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +``` + +Reaching that parity on 0.85 took one contract adaptation, landed earlier in 85ad5e7. +Pi 0.84 and older silently substituted a built-in's stock definition when a `ToolExecutionComponent` was constructed without one, so the calm-off equivalence baseline could be built definition-less and still read as stock. +Pi 0.85 removed that substitution, so the definition-less baseline renders Pi's generic text fallback instead - which is what produced `read collapsed rendering changed while calm mode was off`. +The renderer change was real, and it was the contract's baseline that had to adapt, not Calm's wrappers: the wrapped rows matched Pi stock before and after. +`tests/fm-calm-pi-extension.test.sh` now builds each baseline from the real stock tool-definition factories that `dist/core/tools/index.js` exports, calling the built-in's own factory with `process.cwd()`, which reads as stock on 0.84.4 and on 0.85.x alike and no longer depends on the removed substitution. + +Pi 0.85.0 alone requires a package it does not declare. +Its `dist/experimental/server.js` statically imports `@earendil-works/pi-server`, which is absent from 0.85.0's `dependencies`, `peerDependencies`, and `optionalDependencies`, so a clean install of 0.85.0 on its own cannot load Pi's interactive mode at all: + +```text +Error [ERR_MODULE_NOT_FOUND]: Cannot find package '@earendil-works/pi-server' imported from + .../node_modules/@earendil-works/pi-coding-agent/dist/experimental/server.js +``` + +Installing `@earendil-works/pi-server@0.85.0` beside it restores the identical Calm rendering, and 0.85.1 no longer reaches that import. +That packaging gap is a separate installation defect, not the renderer change above: it stops Pi from loading at all rather than altering any rendered row. + +The `could not render calm-mode HTML export DOM` failure was a headless-Chrome start-up flake, not a change in Pi's export shape. +It appeared in exactly one of the thirteen most recent CI runs, and that run installed the same Pi 0.85.1 as the runs immediately before and after it, which both passed. +The render step is a vendor-tool step: the assertions that follow it are what protect the Calm conversation boundary. +It now retries a bounded number of Chrome start-ups on a fresh profile and, when every attempt fails, reports the Chrome binary, its version, the installed Pi version, each attempt's exit status, whether that attempt was timed out, and Chrome's own stderr, so the next occurrence is diagnosable from the CI log alone. +`test_export_dom_render_guard` in the same script pins that behavior with real processes and no browser. + +The complete Calm suite against installed Pi 0.85.1, with `FM_CHROME_BIN` naming the Chrome the render step used: + +```text +$ FM_CHROME_BIN=<chrome> tests/fm-calm-pi-extension.test.sh +ok - Pi calm resolves its persistent home independently of Pi's launch directory +ok - Pi calm compatibility evidence never rejects a Pi version for being newer than 0.82.0, and still fails closed on a missing or malformed version +ok - a missing collapsed-thinking presentation API degrades only that Calm adapter with a clear skip reason, while the rest of Calm still registers +ok - missing Pi presentation class exports reach the independent adapter degradation path +ok - Calm registers none of its 7 built-in tool wrappers at load while config/calm is off, and all 7 synchronously at load while config/calm is on +ok - Calm's first same-session /calm activation claims every uncontested built-in, leaves a foreign bash tool fully intact and callable, warns prominently and logs the contested name, and only rows constructed before that activation - the documented bound - fail to retroactively collapse +ok - Pi calm centralizes transcript visibility, preserves execution/export data, keeps Pi's stock working row visible while no run is active, and persists its choice across session starts +ok - Pi calm on collapses mid-turn assistant working notes to zero height while Calm off keeps them, leaves streaming, truncated-final, and genuine final replies untouched, never mutates the messages, ignores every /calm argument, and restores a legacy persisted max as ordinary Calm on +ok - Pi operational follow-up E2E processes exact user-role notifications once while Calm hides current and adjacent rows, Calm off and absent render them, and restart preserves semantics +ok - Pi Calm native /skill:ahoy geometry keeps every collapsed thinking and tool block at zero height while preserving expansion, history, restart, and Calm-off rendering +ok - Pi Calm working ship moves on a slow independent cadence over faster fixed-cell blue water, paints the complete boat standard yellow with balanced resets, keeps ANSI-stripped width exact, flips the directional sail on the exact bounce at both edges and every width, clamps visible and hidden resizes, falls back deterministically when narrow, freezes and resumes column/direction across settle/start without hidden-time jumps or duplicate timers, resets only on a fresh session, and installs and removes one scheduler-owning widget across starts, settle, abort, failure, shutdown, reload, replacement, and Calm toggles while leaving Calm-off visibility untouched +ok - the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails +ok - Pi calm native E2E replaces the stock working row with a moving, resize-clamped working ship that freezes and resumes across two working periods in one Pi session, clears on abort, keeps captain turns visible, hides exact operational user rows without changing persistence, restores stock rendering Calm-off, survives restart, and preserves export plus Ctrl+O behavior +``` diff --git a/docs/captain-hold-lifecycle.md b/docs/captain-hold-lifecycle.md index 2959341f600..842846ea1f3 100644 --- a/docs/captain-hold-lifecycle.md +++ b/docs/captain-hold-lifecycle.md @@ -37,7 +37,7 @@ The policy prefers holding the very work item a question gates, so the backlog r `bin/fm-teardown.sh` therefore asks the read-only `open` subcommand before its automatic close: exit 0 means the row is still an open captain call (not Done, `hold_kind: captain`), 1 means it is not, and 2 means the answer could not be established, which teardown treats as a refusal before any destructive step rather than as permission to close. On 0 only the close changes: after cleanup and still under the task's own lock, teardown records one `Deliverable of the finished work: ...` line at the end of the task body and runs `tasks-axi reopen`, so the row returns to Queued with its hold intact and remains on the appropriate Captain's Call or Charted Next decision surface instead of reading as work still under way. The pending-close record teardown already stages before destructive cleanup carries that intent as a `mode=retain` line, so an interrupted cleanup replays the retention at the next session start through the same record, validator, and lock as an ordinary close and never closes the row; an answer that closed the row first simply retires the record. -`--force` does not lift the deferral, because it authorizes discarding unlanded work, never the captain's question, and `answer` remains the only act that closes the call. +`--force` does not lift the deferral, because it authorizes discarding unlanded work, never the captain's question; only `answer` with the captain's words or evidence-backed `reconcile close` closes the call. `bin/fm-backlog-transition-lib.sh` owns the transition and its record, and `bin/fm-captain-hold.sh --help` owns the predicate's contract. ## Answer-time closure @@ -56,6 +56,63 @@ Two channels feed that one intake today, and both are ordinary callers rather th Trusted external process-event adapters intentionally expose no answer operation and cannot feed this authority-bearing intake; [`extension-bindings.md`](extension-bindings.md#trust-boundary) owns that boundary. `bin/fm-procevent-lavish.sh answers` is one such adapter command; it reads only rows tagged `choice`, relays a card's declared close mode, and can never let freeform captain prose forge a task id or a mode. +## Reconcile: re-check reality, never a blind close + +A captain call can stop being a question without the captain ever answering it because the subject lands, the premise turns out to be false, or the choice becomes a matter of fact rather than the captain's to make. +`reconcile` is the standing third option for that case, and its whole point is that it is NOT an answer. +It means "go verify the latest state", and it resolves in exactly one of two ways once that verification has actually been done: close the call with the evidence that made it moot, or leave it open with a note recording that it is genuinely still active. + +The value remains reserved at the shared keyed-answer intake, which visibly refuses it from every channel and never passes it to `answer`. +A reconcile value delivered through chat or any ordinary keyed-answer caller therefore cannot complete a task, lift a hold, write a resolution record, or create a reconcile request. + +Board request creation uses a separate captured-source seam. +The board emits `fm-bearings-answer.v1` context with the slug-shaped selected option and freeform note in separate fields, so annotating Reconcile cannot turn it into an ordinary answer value. +`bin/fm-procevent-lavish.sh answers` emits an exact non-reconcile selection, or a bare note when no option was selected, while `reconciles` emits only task ids whose structured selection is Reconcile and carries their notes as request provenance. +Current rows require the versioned shape and the `choice` tag; a time-limited rollout branch accepts ordinary answers from the old question/answer shape but refuses its bare and separator-annotated reconcile values from both intakes because those rows do not separate the selected option from its note. +Every other structurally uncertain capture feeds neither intake, remains announced, and cannot forge a task id from freeform prose. +The adapter-agnostic runner pipes reconcile rows into `reconcile-requests` only for a bound source, and that intake verifies the named binding again before it creates anything. +Failures remain best-effort and never acknowledge or suppress the captured result. +What this captured-source intake records is a durable reconcile request under `state/reconcile-requests/`, one private record per task, carrying the requesting provenance and a UTC timestamp. +The record exists so the obligation to re-check cannot be lost between the wake that carried the answer and the turn that acts on it. +It is idempotent per task: repeating a reconcile keeps one request and its original timestamp. +The supported creator is the runner carrying the captain's board selection; the binding-checked `reconcile-requests` command is that internal intake rather than an operator reconciliation outcome. + +Verification retires a request through one of two outcomes, and each one requires both the pending board-created request and the operator input that supports its claim: + +- `reconcile close <task-id> --evidence-file <path>` is the moot outcome. + It writes a resolution record whose mode is `reconciled` and whose body is the supplied EVIDENCE under a `Reconciliation evidence:` label, then closes the task. + The distinct mode and label are what keep the record honest: it says the call dissolved against verified evidence, and it never claims the captain answered. +- `reconcile note <task-id> --note-file <path>` is the still-active outcome. + It appends one dated `Captain hold reconciled:` note to the task body, leaves the hold in place, and retires the request. + The call stays the captain's, now carrying what the re-check found; a marker bound to the request timestamp, provenance, and note digest lets a matching retry finish retirement without appending again while a later request with the same finding still receives its own dated note. + +`reconcile list` is the read-only enumeration of pending requests filed by board answers. +A successful normal answer also retires any pending request, because an answered call has no remaining re-check obligation. +Every retirement is checked: if request removal fails after an answer, close, or note is already durable, the durable outcome stands but the command fails and leaves the pending request visible for retry. +No path here closes a captain call without either the captain's words through `answer` or the evidence through `reconcile close`. + +## Card hygiene: a landed subject is not a live call + +`bin/fm-bearings-board.sh build` cross-checks every `decision` card before it publishes and drops stale subjects rather than trusting the composed inventory alone. + +Three checks run, all on exact identity and none on prose: + +- The card's key is the captain-held task id, so `bin/fm-captain-hold.sh open --distinguish-absent` is asked whether that task is still an open captain call. + Exit 1 - present but closed, or no longer held for the captain - drops the card. + Exit 2 means the answer could not be established and exit 3 means the task is absent from the main backlog; both keep the card, because a card wrongly shown is recoverable and a call wrongly hidden is not. +- The payload's own `landed` rows are the recently-landed artifacts. + A decision card whose task id or `pr_url` appears among them has already shipped its subject, so it drops. +- A version decision can carry a structured `subject` with an artifact and numeric three-part version. + A landed row carrying the same artifact at that version or a newer one supersedes the card without parsing prose. + +Dropped cards are named on stderr as `dropped-landed-card:` lines so a rebuild states what it removed rather than quietly shrinking Captain's Call. +The landing procedure requires one immediate board rebuild to remove already-stale merged-PR and superseded-version cards without a committed migration or change-worktree state mutation. +A subject whose state cannot be established is kept, because a wrongly shown card is safer than a wrongly hidden call. +The validator's reservation scope must equal the adapter's reconcile-classification scope, which is all card types because the captured payload carries no card type. +Owner-aware routing for remote-secondmate decision cards is tracked separately: that follow-up must query landedness and route reconciliation in the authoritative secondmate home while honoring the remote and local consistency principle. +Until then, an absent main-home task passes through this hygiene check unchanged, and its Reconcile selection remains announced but cannot create a main-home request because the main intake refuses an absent task. +For a main-home call, the reconcile option is the recovery path for whatever still slips through. + ## Structured read surfaces `bin/fm-fleet-snapshot.sh` parses canonical tasks-axi `(hold: ...)`, `(hold-kind: ...)`, and `(hold-until: ...)` metadata alongside existing backlog fields. @@ -110,17 +167,28 @@ If tasks-axi is unavailable or its listing cannot be parsed, the guard cannot re Older installs created derived `<origin>-decision-<key>` identities through the retired `bin/fm-decision-hold.sh`. Those rows are already plain task ids, so they render, answer, verify, and close through the collapsed surfaces with no data migration. Three legacy inputs are resolved in place: a `decision_keys=` metadata entry that names no task resolves through `<origin>-decision-<entry>`; a channel key that names no task resolves the same way when the source's binding carries a concrete legacy origin; and resolution records written by the old script are recognized wherever a record is read. +On the Beads backend, an attested legacy markdown id that resolves to no task is accepted through the row the markdown-to-beads hold migration produced, found by the authoritative evidence first: a row whose notes carry the marker line `migrated from data/backlog.md id <legacy id>`, either alone or followed by ` on <date>` as fm-hold-migration wrote it on 2026-09-04. +Only when no row carries that marker line is the legacy id tried under the configured beads prefix, and that name-only guess is accepted solely for a single row still held for the captain - two such rows refuse rather than attest. +Because that acceptance rests on a name rather than on evidence, `complete` names the resolved row beside each prefix-attested legacy id in its completion line, so the guess is auditable after the fact. +A markdown home keeps its legacy rows verbatim, so its resolution is unchanged. The shim recognizes an exact replay of a pre-collapse routed resolution by its historical answer digest and routed ids, then finishes any still-recorded dependency-edge cleanup without rewriting the old decision text. `bin/fm-decision-hold.sh` itself remains for one release as a thin command-mapping shim over `bin/fm-captain-hold.sh`, so in-flight work briefed before the collapse keeps working; its header owns the exact mapping. ## Verification record -Verification date: 2026-09-03. - The focused end-to-end regression suite is `tests/fm-captain-hold-lifecycle.test.sh`, using only synthetic `sample` identities and decision text. It proves: cleanup of a finished task whose own row is the captain call leaves that call open, queued, held, carrying its deliverable, and visible in Bearings' Captain's Call, leaves no pending record behind, survives a `--force` cleanup, and closes only when `answer` records the captain's words, while an ordinary finished task in the same home still closes with its report link; an interrupted cleanup leaves the row In flight and untouched with its pending record, and the next session start retains it as queued and held with the deliverable recorded; a relocated data directory keeps the retention in its one configured backlog; a ship row whose captain hold cannot be read refuses cleanup before any destructive step and surfaces the read failure; the reconstructed silent-divergence case is signalled - a status resolution over a still-open captain-held task reaches both `diverged` and the drain's `RECORD DIVERGENCE` section, under the collapsed and the legacy identity alike, while the backlog task, its hold, and the status log all survive the report unchanged and the printed hint names both reconciliation directions; the false-signal boundary holds - a captain call with no routed work item, a verified `captain-held` transfer, a still-open status decision, an already answered call, and an ordinary task whose keyed question was answered all stay silent; a report-only unresolved captain call refuses `--none` completion before teardown can erase the source; non-forced scout teardown always requires the durable inventory verification; the recorded-answer guard (a bare `tasks-axi done` close fails `verify` until `answer` records the captain's word, and an ordinary finished task cannot be dressed up as an answered call); answer-time closure through a bound channel with task-id keys, including the `release` close mode, mode-matched replay idempotence, and the refusal of drifted, mode-mismatched, absent, unheld, and already-closed keys; the chat channel reaching the same intake; hold-set stamping that precedes visible hold state, preserves an active lifecycle's timestamp, and resets after release; interrupted answer closure retaining the stamp until close and restoring resolution-first ordering on retry; deferral through `--until` leaving `captain_actionable` false until due; and every legacy path (composed identities through the shim, pre-collapse `decision_keys=` metadata, routed-resolution replay, and a concrete-origin binding). +The markdown-to-beads migration family runs the same suite's beads fixture (bd-driven scratch graph, self-skipping on markdown-only tasks-axi installs) and proves: `verify` and `complete` resolve an attested legacy id through a migrated row's marker note, through the configured prefix when no row carries a note - naming the resolved row in the completion line - and through the marker note of a pre-collapse derived identity; a marker-noted row wins over an unrelated captain-held row occupying the bare prefix namesake; an unresolvable id is refused once naming the id (never an empty name); and the attested id stays in `decision_keys=` for idempotent re-verification. +One case in that family needs no beads install and always runs: a stubbed tasks-axi that fails any markdown file override proves the captain-hold hold, answer, and close mutations reach a beads-configured home without one. + +The reconcile path is pinned in the same suite: a reconcile answer arriving through the keyed-answer intake, in the default close mode and in the `release` mode a captain-gated work card declares, is refused and leaves both tasks held with no resolution record or request; only the separately bound captured-source intake records one durable request per task idempotently across a replay. +It also proves the two verification outcomes - an evidence-backed `reconciled` close that records the evidence under its own label and never as the captain's words, and a note that leaves the call queued, held, and dated - while both outcomes refuse without a pending board request, each durable mutation applies only once across close, probe, and request-retirement failures, a later distinct request with the same note still appends its own dated record, every failed retirement is surfaced with its pending request retained, incompatible resolution modes cannot replay as captain answers, and normal close, release, and replay paths retire pending requests. +The captured-source coverage proves Lavish deduplicates each card before separating versioned structured selections from notes, bare and annotated Reconcile choices never reach keyed answers, genuine current and legacy choices still close normally, legacy bare and separator-annotated reconcile values feed neither intake, mixed repeated selections preserve every other card's final value, the generic runner creates a request only through a verified bound source, chat reconcile text creates none, and the resulting board request authorizes evidence-backed closure. +The board's half is pinned in `tests/fm-bearings-board.test.sh`: every published decision card carries exactly one reconcile option, authored options reserve that value across every card type, recommendations name authored options, a decision card whose structured subject appears in the payload's landed rows is dropped while a genuinely open one is kept even when an unrelated landed id contains its key after a newline, a build requires a fresh authoritative listed-open result before binding or arming, a reopen retires the pre-reopen source generation and waits for a fresh live listener, and a rebuild of an already-armed board with no live listener starts one. +That suite drives its Lavish session through a protocol-shaped stub, and `tests/fm-bearings-board-lavish-live-e2e.test.sh` is the default-on capability guard for the installed provider; [`verification/process-event-sources.md`](verification/process-event-sources.md) owns the version-scoped evidence. +[`verification/process-event-sources.md`](verification/process-event-sources.md) owns the process-event ownership and reclamation evidence exercised by `tests/fm-procevent.test.sh`. `tests/fm-classify-decision-key.test.sh` pins `status_key_closing_verb` itself: it separates a resolution from the durable-transfer close and from a still-open key, reports the last real transition across re-openings and both key positions, and treats a prose mention as no transition. Projection regressions live in `tests/fm-fleet-snapshot-view.test.sh` (the total structured-only bucket classifier, hold-until parsing, kind-independent captain actionability, undated-hold aging, and title stripping) and `tests/fm-bearings-snapshot.test.sh` (default and expanded decision-bucket membership, deferral explanations, blocker-overflow disclosure, working-hold dual surfaces, remote-summary schema invalidation, and the landed exclusion by surviving captain-hold annotations). -The exact commands and their summarized outputs are recorded in the shipping PR's evidence; run the four suites above plus `tests/fm-send-resolve-key.test.sh`, `tests/fm-bearings-board.test.sh`, and `bin/fm-lint.sh` to refresh this record. +The exact commands and their summarized outputs are recorded in the shipping PR's evidence; run the four suites above plus `tests/fm-send-resolve-key.test.sh`, `tests/fm-bearings-board.test.sh`, `tests/fm-procevent.test.sh`, and `bin/fm-lint.sh` to refresh this record, and `FM_BEARINGS_LAVISH_LIVE=1 tests/fm-bearings-board-lavish-live-e2e.test.sh` after a lavish-axi upgrade. diff --git a/docs/cd-guard.md b/docs/cd-guard.md index 9e02f9d9e26..b72ffb49679 100644 --- a/docs/cd-guard.md +++ b/docs/cd-guard.md @@ -81,7 +81,7 @@ It does not permit `cd /home/project`, because an absolute-path `cd` remains a p - GitHub Copilot CLI sends stdin JSON at `.tool_input.command` and adds `--copilot`, which renders its native permission decision. - Grok sends stdin JSON at `.toolInput.command`. - OpenCode sends the exact command string through `--command <exact string>`. -- Pi and pi-signed send the exact command string through `--command <exact string>`. +- Pi, pi-signed, and omp send the exact command string through `--command <exact string>`. - Cursor sends stdin JSON at `.tool_input.command` and adds `--cursor`, which renders the deny as Cursor's own returned decision object. Processing order is cheapest-first: a strict-superset prefilter, then the primary-checkout scope, then the Node policy owner. @@ -102,7 +102,7 @@ Identical in shape to `docs/arm-pretool-check.md`: - `--claude` suppresses stdout completely because Claude ignores a PreToolUse deny when stdout is nonempty. - Codex blocks on exit 2 and displays stderr. - OpenCode throws only when the checker exits 2. -- Pi and pi-signed return `{block: true}` only when the checker exits 2. +- Pi, pi-signed, and omp return `{block: true}` only when the checker exits 2. - GitHub Copilot CLI reads `{"permissionDecision":"deny","permissionDecisionReason":"..."}` from stdout and therefore returns exit 0 in `--copilot` mode. ## Shared classifier ownership @@ -122,6 +122,7 @@ The cd-guard never duplicates shell lexing; it adds only the cd-specific decisio | Grok | `.grok/hooks/fm-primary-cd-check.json` PreToolUse hook anchored on `${GROK_WORKSPACE_ROOT:-}` | Consumes the stdout `decision=deny` object. | | OpenCode | `.opencode/plugins/fm-primary-cd-check.js` `tool.execute.before` | Throws, which surfaces as the failed tool result. | | Pi | `.pi/extensions/fm-primary-turnend-guard.ts` `tool_call` handler | Returns `{block: true}`; piggybacks on the already-loaded primary extension so no extra `-e` flag is needed. | +| omp | `.omp/extensions/fm-primary-turnend-guard.ts` `tool_call` handler | Returns `{block: true, reason}` and omp surfaces the reason to the model; runs before the watcher-arm seatbelt in the same auto-discovered extension, so no `-e` flag is needed. | | Cursor | `.cursor/hooks.json` `preToolUse` hook matching `tool_name` `Shell`, forwarding stdin with `--cursor` | Prints Cursor's own `{"permission":"deny","user_message":...}` object on stdout and exits 0, because Cursor reads the returned object rather than the exit status. Without `--cursor` the Cursor-delivered payload is the Claude-settings duplicate Cursor also loads, and allows; `docs/arm-pretool-check.md` owns that shared predicate. | Each harness runs the cd-guard alongside the watcher-arm seatbelt; the two are independent checks, and either deny blocks the command. diff --git a/docs/configuration.md b/docs/configuration.md index 517e30fb01f..d45fffb4f2a 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -59,7 +59,8 @@ The effort list is a handful of levels and stays on Pi's plain selector dialog. Both picks change the supervision branch alone and never the captain's own conversation model or effort. It persists the model pick in gitignored `config/supervision-branch-model` and the effort pick in gitignored `config/supervision-branch-effort`, both under the effective Firstmate home, resolved from `FM_HOME`, then `FM_ROOT_OVERRIDE`, then the tracked code root derived from the extension path, or under `FM_CONFIG_OVERRIDE` when that test and specialized-setup override is present. Firstmate keeps no model catalog of its own; the list is the intersection of what Pi reports when the picker opens and what a fresh isolated branch runtime can run. -A provider that exists only because an extension registered it inside the captain's session is not offered, while stored OAuth and API-key credentials retain their native credential type because Firstmate never copies, converts, installs, or overwrites credentials for the branch runtime. +A provider that exists only because an extension registered it inside the captain's session, such as pi-devin-auth's `devin`, is offered and can be pinned or followed like any other; [pi-supervision-branch.md](pi-supervision-branch.md#cost-model-and-the-byte-stable-prefix) owns how that registration reaches the isolated branch runtime. +Stored OAuth and API-key credentials retain their native credential type because Firstmate never copies, converts, installs, or overwrites credentials for the branch runtime. The file holds one `<provider>/<model-id>` line followed by one newline, split at the first `/` so a provider-qualified model id such as `openrouter/anthropic/claude-sonnet-4-5` survives intact. An absent, unreadable, or unparseable file means no pin, and the branch then follows main's own current model, applied explicitly and live whenever main changes models mid-session. A valid pin wins over main and remains unaffected by main's model changes. @@ -93,8 +94,11 @@ When the default backend is selected and compatible `tasks-axi` is on `PATH`, fi When the automatic transition gate applies, dispatch and completion are not separate operator actions: each moves its work item inside the same run that creates or removes the task's record, so the ordinary successful path cannot leave the backlog and live task set out of sync ([`bin/fm-backlog-transition-lib.sh`](../bin/fm-backlog-transition-lib.sh)). Under that gate, dispatch accepts only an unheld, unblocked Queued or In flight item in this home; a missing, Done, held, or dependency-blocked item is refused before any endpoint or local copy is created. Completion refuses to report success until the item is closed, and session start reconciles this home's own books after an interrupted run. +When a spawn is interrupted after launch delivery began, its exit path re-reads the paired task record and the backlog row under the same per-task lock as the commit, repairs a row the commit believed it had moved, and reports only what was verified or honestly attempted, never intent phrased as outcome ([`bin/fm-spawn.sh`](../bin/fm-spawn.sh); [`tests/fm-backlog-atomicity.test.sh`](../tests/fm-backlog-atomicity.test.sh)). Automatic transition mutations address the configured `<data>/backlog.md` explicitly from the data directory's parent, keeping relocated backlog configuration, archives, and relative scout-report links together. -The gate does not apply to persistent secondmates, manual-backend homes, or homes without a backlog file, preserving their existing persistent-agent, manual, or ad-hoc lifecycle behavior. +That explicit markdown file belongs to the markdown backend only: a home whose resolved tasks-axi backend is non-markdown never receives a markdown file override and requires no markdown backlog file, so its reads, probes, and mutations address the backend its own configuration selects. +Migrated-hold resolution on a beads home reads its graph path, binary, and prefix from the root `.tasks.toml` `[beads]` section only, and refuses (rc=2) when the beads backend is selected elsewhere (a `TASKS_AXI_BACKEND` override or user-level config) with no root-level `[beads]` section. +The gate does not apply to persistent secondmates, manual-backend homes, or markdown homes without a backlog file, preserving their existing persistent-agent, manual, or ad-hoc lifecycle behavior. On an automatic-backend home with a backlog, missing or incompatible `tasks-axi`, an unresolvable configured data directory, or one containing a control byte fails lifecycle work before mutation. Secondmate handoffs bypass that routine-backend choice: `fm-backlog-handoff.sh` keeps only its own fleet-level validation, delegates the item move to `tasks-axi mv`, and requires a verified receiver wake after a new move becomes durable. It moves in-scope `## Queued` items only and refuses `## In flight` and historical `## Done` records, which stay with their home for pruning or archiving. @@ -110,7 +114,7 @@ The file format is unchanged in both modes; tasks-axi and manual edits produce t ## Runtime backend (config/backend / FM_BACKEND) For spawn-capable adapters, the runtime session-provider backend controls where task windows/endpoints are created, captured, sent to, watched, and killed. -`tmux` is the verified reference backend (see [`docs/tmux-backend.md`](tmux-backend.md)); `herdr`, `zellij`, `orca`, and `cmux` are experimental spawn backends (see [`docs/herdr-backend.md`](herdr-backend.md), [`docs/zellij-backend.md`](zellij-backend.md), [`docs/orca-backend.md`](orca-backend.md), and [`docs/cmux-backend.md`](cmux-backend.md)). +`tmux` is the verified reference backend (see [`docs/tmux-backend.md`](tmux-backend.md)); `herdr` has its own required CI lane (see [`docs/herdr-backend.md`](herdr-backend.md)); `zellij`, `orca`, and `cmux` remain experimental spawn backends with no dedicated real-backend CI lane (see [`docs/zellij-backend.md`](zellij-backend.md), [`docs/orca-backend.md`](orca-backend.md), and [`docs/cmux-backend.md`](cmux-backend.md)). Treehouse remains the worktree provider for tmux, herdr, zellij, and cmux, since herdr, zellij, and cmux are session providers only; Orca provides both the task worktree and terminal endpoint. New spawns choose the backend in this order: an explicit `--backend` flag that current authority for that exact task alone has authorized (a present captain instruction or the task's own accepted brief; never later-task precedent by analogy), then `FM_BACKEND`, then the first non-empty line of local gitignored `config/backend`, then runtime auto-detection from `$TMUX`, `HERDR_ENV=1`, or cmux runtime signals, then default `tmux`. If more than one runtime marker is present, detection resolves innermost-first: `$TMUX` is checked before `HERDR_ENV=1`, which is checked before cmux's primary `CMUX_WORKSPACE_ID` marker and its documented fallback signals - tmux or herdr started from inside a cmux terminal is the innermost, currently-executing layer, while cmux itself (a terminal application, not a nestable multiplexer) is always checked last. @@ -200,13 +204,10 @@ The flag is a home-local supervision-noise preference and is not inherited by se ## Gate defaults (.no-mistakes.yaml) -The tracked `.no-mistakes.yaml` sets `test.evidence.store_in_repo: true`, pins `commands.lint` to `bin/fm-lint.sh` so local lint matches CI, and pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated` so the gate's test baseline runs through the repository's own runner instead of a hand-chained walk of `bash tests/*.test.sh`. +The tracked `.no-mistakes.yaml` sets `test.evidence.store_in_repo: true` and pins `commands.lint` to `bin/fm-lint.sh`, the same owner CI invokes. Storing evidence in the repo publishes each run's test artifacts to the orphan `no-mistakes/evidence` branch and links them from the PR body, instead of keeping them on local disk under the no-mistakes home. That branch shares no history with code branches, so evidence never enters a pushed feature branch or the default branch; the worktree's `.no-mistakes/` stays local and CI rejects tracked entries under that path. -`commands.test` stays changed-file-scoped and must never become a complete `tests/*.test.sh` walk: `--changed` selects only the families the branch's changed files map to, runs concurrency-admitted scripts with bounded concurrency, keeps every unproven stateful or measured Windows-heavy script serial, and applies its own generous per-script bound. -The runner's `--help` output owns the exact selection, scheduling, and timeout rules. -It excludes `real-herdr-gated` on the same grounds the portable CI lanes do, because those scripts drive a live Herdr lab and the dedicated required Herdr lane owns that coverage. -Because firstmate always supplies `--intent`, that command is a baseline and the Test step still runs its intent-targeted evidence agent on top of it. +The [`firstmate-coding-guidelines` skill](../.agents/skills/firstmate-coding-guidelines/SKILL.md#no-mistakes-test-configuration) owns why `commands.test` stays absent and targeted validation belongs to the evidence path. `commands.test` executes code, so no-mistakes honors it only from the default-branch copy of `.no-mistakes.yaml`; a pushed branch cannot change what the gate runs. See [CONTRIBUTING.md](../CONTRIBUTING.md) for the firstmate-specific local test policy and entry points. Portable shard evidence and coverage rules are in [fm-test-portable-shards.md](fm-test-portable-shards.md); [herdr-backend.md](herdr-backend.md#destructive-lab-safety) owns the real-Herdr lane's isolation boundary, and [runtime-backends.md](verification/runtime-backends.md#herdr) owns active evidence. @@ -296,7 +297,7 @@ The full cmux home label also includes a short hash of the resolved `FM_ROOT` pa ## Harness support -claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, and cursor are empirically verified for crewmate and secondmate launches; gemini is verified for crewmate and scout launches only, and [README requirements](../README.md#requirements) own the set supported for the primary session. +claude, codex, copilot, opencode, pi, pi-signed, grok, kimi, cursor, and omp are empirically verified for crewmate and secondmate launches; gemini is verified for crewmate and scout launches only, and [README requirements](../README.md#requirements) own the set supported for the primary session. A Copilot primary or secondmate uses tracked `.github/hooks/firstmate.json` for session start, asynchronous watcher completion notifications, a nonblocking turn-end backstop, and primary pre-tool seatbelts; [`docs/supervision-protocols/copilot.md`](supervision-protocols/copilot.md) owns its supervision protocol. Generated Copilot worker hooks also publish semantic busy state and a prompt-submission acknowledgement, so `fm-send` does not depend on a backend-specific Copilot composer shape. A cursor secondmate or primary runs the tracked project-scope `.cursor/hooks.json` in its own home and must be launched with `--trust`, or no project hook loads; [`docs/supervision-protocols/cursor.md`](supervision-protocols/cursor.md) owns its supervision protocol. @@ -305,6 +306,7 @@ On Zellij, cmux, and Orca a typed-plane Cursor send (a harness-native invocation muse is verified for crewmate and scout launches ONLY, and `fm-spawn.sh` refuses it for a secondmate, because muse ships no usable hook surface for a primary session's turn-end supervision; [`docs/verification/muse.md`](verification/muse.md) owns that evidence. muse also needs a worker-reachable credential before spawning, and the portable fleet path is the `<config>/muse/auth.json` credential stored by `muse login`, because a caller-only `META_API_KEY` does not cross a long-lived backend daemon. gemini is likewise refused for secondmates because it has no primary supervision protocol; [its adapter reference](../.agents/skills/harness-adapters/references/harness/gemini.md) owns the credential precondition, canonical-launch wiring, and raw-launch limitations. +rovo is likewise verified for crewmate and scout launches ONLY, refused for a secondmate for the same reason - no turn-end hook and no primary supervision protocol; [`docs/verification/rovo.md`](verification/rovo.md) owns that evidence, including the OAuth token's silent background refresh from a stored refresh token and both tmux and herdr pane liveness (herdr placement is verified live, with a Herdr-side agent-detection gap left open for recovery classification). New harnesses get verified through a supervised trial task before joining the set. The verified adapter evidence - each harness's busy-state source, interrupt and exit behavior, skill-invocation syntax, and per-harness quirks - lives in the skill tree rooted at [`.agents/skills/harness-adapters/SKILL.md`](../.agents/skills/harness-adapters/SKILL.md). The executable interrupt and exit mechanics live in [`bin/fm-control-lib.sh`](../bin/fm-control-lib.sh), and [`docs/agent-control.md`](agent-control.md) owns their lifecycle-control architecture. @@ -313,7 +315,7 @@ Pi-family launches adapt the regular-TUI safeguard to the installed CLI's capabi Enabled primary-session turn-end guard integrations are tracked as repo-level hook files and documented in [`docs/turnend-guard.md`](turnend-guard.md). Kimi remains outside the primary turn-end guard integrations; [`docs/turnend-guard.md`](turnend-guard.md#compatibility-limits) owns its separate captain-approved crew wake hook. Primary-session watcher wake protocols are rendered at session start by [`bin/fm-supervision-instructions.sh`](../bin/fm-supervision-instructions.sh) from [`docs/supervision-protocols/`](supervision-protocols/). -Claude's Stop `asyncRewake` hook owns tokenless re-arm cycles, Copilot uses a tracked asynchronous watcher with shell-completion notification re-entry, Cursor's stop hook parks on the watcher, Grok uses background-notify cycles, Codex uses bounded foreground checkpoints, Pi and pi-signed use the same two tracked primary extensions, and OpenCode uses its TUI plugin. +Claude's Stop `asyncRewake` hook owns tokenless re-arm cycles, Copilot uses a tracked asynchronous watcher with shell-completion notification re-entry, Cursor's stop hook parks on the watcher, Grok uses background-notify cycles, Codex uses bounded foreground checkpoints, Pi and pi-signed use the same two tracked primary extensions, omp uses its own two tracked `.omp/extensions/` files with a blocking `session_stop` turn-end hook, and OpenCode uses its TUI plugin. `config/crew-harness` is a local, gitignored file containing one adapter name for crewmate and scout launches. When pi-signed is selected, Firstmate preserves `FM_PI_HARNESS=pi-signed` and refuses the launch if the selected executable is unavailable rather than falling back to pi; [`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns executable resolution and launch mechanics. Plain Pi launches set `FM_PI_HARNESS=pi`, so a signed primary's environment cannot relabel a plain Pi worker. @@ -337,6 +339,55 @@ Kimi continues to use the captain's normal Kimi home, including the existing con The Kimi installer requires an existing regular non-symlink `~/.kimi-code/config.toml`, `python3` with `tomllib`, and `jq`; it validates but never serializes the captain's TOML and refuses before writing when the config is missing, malformed, or surprising or when either tool requirement is unavailable. Its `remove` action excises only the marker-delimited Firstmate region and removes Firstmate's hook files. For Pi and pi-signed secondmate launches, `fm-spawn.sh` starts the selected executable with `-e` pointed at the secondmate home's own tracked `.pi/extensions/fm-primary-pi-watch.ts` and `.pi/extensions/fm-primary-turnend-guard.ts`, both already present from the secondmate home's git worktree. +For omp secondmate launches, `fm-spawn.sh` passes no `-e` at all: omp auto-discovers the home's tracked `.omp/extensions/` with no trust gate, and naming a discovered file with `-e` as well loads it twice; every omp launch instead carries the tracked `.omp/fm-worker-overlay.yml` posture overlay through `--config`, which [`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns. + +## Worker launch environment (config/launch-env-allowlist) + +The optional local, gitignored `config/launch-env-allowlist` limits the ambient environment passed to newly launched workers, scouts, and secondmates, including relaunches. +With no file, launch behavior is unchanged: selected harness markers are cleared, while the provider, long-lived terminal daemon, and shell initialization determine which other variables reach the worker. +Do not assume every worker inherits the invoking Firstmate process's current environment. +The file is inherited into secondmate homes through the [primary-authoritative configuration contract](../.agents/skills/secondmate-provisioning/SKILL.md). +Changes apply to subsequent launches; existing processes keep their environment. + +Create the file with one environment variable **name** per line, never credential values, assignments, wildcards, or shell commands. +Blank lines and lines beginning with `#` are allowed. +Invalid names, an unreadable or nonregular file, or a path inspection error (including an inaccessible configuration directory) stop the launch. +An empty file enables filtering with only Firstmate's operational floor. +For example, a provider using `OPENAI_API_KEY` and Git using an SSH agent could use: + +```text +# Provider credential already available in the destination pane +OPENAI_API_KEY +# Git over SSH using an existing agent +SSH_AUTH_SOCK +``` + +Firstmate retains basic home, executable search, terminal, locale, temporary-directory, and backend routing variables, plus its explicit launch assignments, its ship and scout task marker, and enabled task trace. +[`fm-spawn.sh --help`](../bin/fm-spawn.sh) owns the exact retained names and parsing mechanics. +Other ambient names must be listed explicitly, including custom credential-store locations, proxy settings, and certificate overrides when required by the selected tools. +The command shell and worker may still create their own variables. +Allowed values come from the destination pane at execution time; they are neither copied from the invoking Firstmate process nor written into the launch command. +Listing a name does not provision it in a daemon's environment or transfer credentials to another machine. + +Choose the minimum additions for the authentication method actually in use: + +| Provider or Git transport | Additional names needed | +| --- | --- | +| Provider login stored under the normal home directory | None for the environment contract; the same user still has access to that provider's stored login. | +| Provider configured through environment variables | The exact credential and endpoint names required by that provider, for example `OPENAI_API_KEY` or `ANTHROPIC_API_KEY`; a multi-provider tool needs each provider it will actually use. | +| Custom provider store | Its configured location variables, such as `CODEX_HOME`, `GROK_HOME`, or `XDG_CONFIG_HOME`; Firstmate's existing explicit Claude and Muse store assignments still apply. | +| Muse environment authentication | `META_API_KEY`, already present in the target tmux session environment; Firstmate's preflight requires the stored-login path on other backends. | +| Git over SSH with an agent | `SSH_AUTH_SOCK`; add `GIT_SSH_COMMAND` only if the chosen transport requires that override. | +| Git over SSH with a key file | No credential variable when normal SSH configuration selects the key; file permissions and any passphrase handling still apply. | +| Git over HTTPS with a credential helper | Whatever the configured helper requires; a GitHub CLI helper using an environment token needs its selected `GH_TOKEN` or `GITHUB_TOKEN`. | + +Verify the selected provider login and Git transport after opting in; Firstmate does not infer credentials from model names or install a secret manager. +Raw launch commands run under noninteractive POSIX `sh` with this option and must use compatible syntax. +The filter runs at the worker command boundary, after the terminal daemon and pane shell have started; it does not scrub either of those processes. +This is not a sandbox: it cannot revoke same-user access to credential files, prevent tools or later shells from loading credentials again, or isolate processes from the same user's other processes. +Regression coverage executes emitted launch commands with synthetic nonsecret values in [`tests/fm-spawn-dispatch-profile.test.sh`](../tests/fm-spawn-dispatch-profile.test.sh). + +Every claude launch's inline `--settings` JSON also carries `"attribution":{"commit":"","pr":"","sessionUrl":false}`, so a spawned worker never writes a Co-Authored-By trailer, Claude-Session link, or generated-with line into a commit or PR body regardless of which settings scopes end up loaded. ## Crew dispatch profiles (config/crew-dispatch.json) @@ -390,7 +441,7 @@ The universal toolchain is node, git, gh with GitHub auth via `gh auth login`, n This section is the single owner of that universal toolchain list; backend guides' prerequisites point here and add only their backend-specific tools. In that list, no-mistakes runs the validation pipeline, gh-axi, chrome-devtools-axi, and lavish-axi cover GitHub, browser, and rich-review operations, and tasks-axi plus quota-axi back backlog mutations and quota-aware array dispatch. The per-backend delta is required only for the backend resolved from `FM_BACKEND`, then `config/backend`, then runtime auto-detection, then default `tmux`, so a home is never told to install a tool an inactive backend or feature would need. -That delta is owned in code by `fm_backend_required_tools` in `bin/fm-backend.sh`: the resolved backend's own session-provider CLI (`tmux`, `herdr`, `zellij`, `orca`, or `cmux`), `jq` for the JSON-emitting experimental adapters (`herdr`, `zellij`, `cmux`) whose spawn and liveness paths parse the backend's JSON output, and the `treehouse` worktree provider for every session-provider-only backend (`tmux`, `herdr`, `zellij`, `cmux`). +That delta is owned in code by `fm_backend_required_tools` in `bin/fm-backend.sh`: the resolved backend's own session-provider CLI (`tmux`, `herdr`, `zellij`, `orca`, or `cmux`), `jq` for the JSON-emitting adapters (`herdr`, `zellij`, `cmux`) whose spawn and liveness paths parse the backend's JSON output, and the `treehouse` worktree provider for every session-provider-only backend (`tmux`, `herdr`, `zellij`, `cmux`). Backend tool availability uses the adapter's own executable resolver, so bootstrap and spawn agree on supported non-`PATH` locations such as cmux's bundled CLI. An unknown resolved backend emits `BACKEND_INVALID` and blocks dispatch instead of silently dropping its dependency delta or falling back to tmux. Orca provides both the task worktree and terminal endpoint (see "Runtime backend" above), so `backend=orca` requires only `orca` on top of the universal toolchain and skips both `treehouse` and every other backend's session CLI. @@ -418,7 +469,7 @@ When a running home advances and its loaded instruction surface (`AGENTS.md`, `b If that send fails, bootstrap keeps an idempotent retry marker and emits `NUDGE_SECONDMATES:` with the failure reason. The same bootstrap run emits `SECONDMATE_LIVENESS:` only when a registered secondmate is skipped or its relaunch fails; already-live and successfully relaunched secondmates are handled silently. For a mid-session inherited local-material edit where tracked-file sync is not needed, run `bin/fm-config-push.sh`. -It uses the same live secondmate discovery and propagation helper as bootstrap, prints each live home's `crew-dispatch.json`, `crew-harness`, `backlog-backend`, `backend`, `herdr-presentation-spaces`, `startup-memory-budget`, `trace-context`, and `data/captain-shared.md` result as `pushed`, `unchanged`, `skipped`, or `error`, and exits non-zero for real propagation errors or config-reread send failures. +It uses the same live secondmate discovery and propagation helper as bootstrap; its [help](../bin/fm-config-push.sh) owns reporting and exit semantics, and [`fm_config_inherit_items`](../bin/fm-config-inherit-lib.sh) declares the inherited items. When an allowlisted config item changes for an already-running local home, it sends the literal-content reread pointer described in [`secondmate-provisioning`](../.agents/skills/secondmate-provisioning/SKILL.md); unchanged allowlisted config sends no pointer unless a previous delivery is pending. A changed remote home instead receives one durably recorded marked re-read instruction after the allowlisted bytes have transferred because primary-local generation paths are not meaningful on another host. The locked bootstrap inheritance pass uses the same placement-specific behavior; see `secondmate-provisioning` for the single contract owner. @@ -472,7 +523,7 @@ See [`docs/examples/watched-tools.json`](examples/watched-tools.json) for a star Arm the check once per home with `bin/fm-tool-update-check.sh arm`. That writes `state/tool-updates.check.sh` and binds its bytes with `bin/fm-check-register.sh`, so the existing watcher polls it on its normal cadence and turns its one line into a `check:` wake; no separate schedule is involved. -The armed check runs whenever that home has a watcher running, and arming alone does not make watcher supervision required, so a home with no in-flight work and no other reason to watch does not start a watcher just for this check. +Registering the check is itself a reason to watch, so the home keeps a watcher for it after the last task is torn down, and `disarm` is what ends that need. `bin/fm-tool-update-check.sh disarm` removes the shim, its trust binding, and the report record. The check prints nothing when everything is current, and `state/.tool-updates` records the findings the last report was made from so the same pending update is reported once instead of on every poll. A changed or returning condition is reported again. @@ -683,8 +734,9 @@ Never run the registered blocking source command directly in a conversational tu A long-polling external process is registered as a *source* through its adapter, whose header and `--help` own the commands and flags. `bin/fm-procevent.sh` owns the generic contract; built-in adapters retain their tracked `bin/fm-procevent-<adapter>.sh` commands, while an explicitly bound external adapter routes through the trusted host contract above. `bin/fm-procevent-lavish.sh` is the first built-in adapter and wraps only the currently published `lavish-axi poll` interface. -That adapter, and only that adapter, retries the one exact transient response a cut-short listener returns while its marks remain available (`error: Lavish Editor poll response was interrupted` with `code: SERVER_ERROR`), up to 12 times at 5 second intervals, so an internal retry never reaches the runner as a captured result. -Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that same interruption still standing once the bound is spent are all captured and announced normally; `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second test override for the interval only, and the runner itself stays adapter-agnostic. +That adapter, and only that adapter, retries the one exact transient response a cut-short listener returns while its marks remain available (`error: Lavish Editor poll response was interrupted` with `code: SERVER_ERROR`), up to 12 times with poll starts at least 5 seconds apart, so an internal retry never reaches the runner as a captured result. +This start-to-start governor is a no-op after a normally blocking poll but caps an immediately returning poll under the shipped defaults independently of the owner lease and registration launch pacing. +Real feedback, ended and missing sessions, any other `SERVER_ERROR`, and that same interruption still standing once the bound is spent are all captured and announced normally; `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 1 to 60 second test override for the interval only, and the runner itself stays adapter-agnostic. An already-armed Lavish source keeps its registered listener command until it is retired and armed again, so re-arm a live board once to adopt this retry policy. The `when` adapter (`bin/fm-procevent-when.sh`) turns this channel into a condition->action primitive: it registers a deterministic condition and a deterministic action once, its blocking child polls the condition without waking firstmate, and a stable true fires the action at most once before one terminal outcome is durably captured and published as a wake that remains eligible for re-announcement until handled. @@ -731,20 +783,27 @@ Keyed captain answers from built-in adapters use one more seam of the same kind, Some built-in sources carry the captain's answer to a captain-held task, and what such an answer means is owned once by `bin/fm-captain-hold.sh`'s keyed-answer intake rather than by any channel. A built-in source bound with `bin/fm-captain-hold.sh bind` therefore has each captured result passed to `bin/fm-procevent-<adapter>.sh answers <result-file>`, and whatever that prints is piped straight into that intake. A binding can select one decision origin or the script's cross-origin mode; the command header owns the exact forms and key interpretation. -The built-in adapter reports only what the captain chose; the intake owns every rule about what happens next, so the runner names no adapter, parses no result, and carries no decision rule, and a future built-in source needs nothing here beyond an `answers` command and a binding. -Feeding is independent of handling: it never acknowledges a result and never suppresses a wake, because recording the answer is transcription while acting on it is firstmate's judgement. -An unbound built-in source, a built-in adapter with no `answers` command, and a failure on either side all leave the capture untouched and still announced. -External binding responses never enter this authority-bearing intake. +The built-in adapter reports only what the captain chose; the intake owns every rule about what happens next, so the runner names no adapter, parses no result, and carries no decision rule, and a future built-in answer source needs nothing here beyond an `answers` command and a binding. +The reserved Reconcile selection uses the parallel optional `reconciles` adapter command and binding-verified `reconcile-requests` intake rather than entering keyed answers; [`captain-hold-lifecycle.md`](captain-hold-lifecycle.md#reconcile-re-check-reality-never-a-blind-close) owns those semantics. +Feeding is independent of handling: it never acknowledges a result and never suppresses a wake, because recording the answer or request is transcription while acting on it is firstmate's judgement. +An unbound built-in source, a built-in adapter without the corresponding command, and a failure on either side all leave the capture untouched and still announced. +External binding responses never enter either authority-bearing intake. Ownership is machine-wide per canonical source, because separate homes can share one underlying source store. Claims live under `$XDG_STATE_HOME/firstmate/procevent-claims` (override with `FM_PROCEVENT_CLAIM_ROOT`). Each claim binds its caller-reported home and runner PID to a process identity, unique claim generation, exact registration-file generation, and resolved state-root identity. Registration, acquisition, replacement, retirement, and generation-bound release are serialized at one machine-wide boundary per source. A live identity-matched owner is never displaced, and release removes only the exact generation the caller acquired. -Retirement and orphan reconciliation signal a runner process group only while its recorded process identity still matches, or when the recorded leader is gone and only its own owned group survives. -A runner leads its own process group, so a claim counts as reclaimable only when that whole generation is gone: a crashed leader whose group still has members is not stale, and reconcile stops that surviving group and releases its generation before starting any replacement. +Retirement and orphan reconciliation select a runner process group for signalling only while its recorded process identity still matches and the live runner still leads that group. +A claim counts as reclaimable only when its owner is stale and an independent process-group check finds no members; a crashed leader or reused pid whose process group still has members cannot relax ownership cleanup, so reconcile preserves the claim without signalling the ambiguous group or starting a replacement. +Reclaiming a generation that IS gone is not gated on tidying its capture-reservation records. +Those records are keyed by claim token and every replacement claims a fresh one, so a leftover that can no longer be located - a state-root identity a claim recorded before its home was re-created, for example - is stale bytes rather than an ownership hazard. +Ordinary release and reclamation still attempt reservation cleanup and require it unless both owner staleness and whole-group absence prove the generation gone. +The narrow live-owner terminal-self-retirement path also attempts cleanup but tolerates its own still-in-flight reservation, which the runner removes on the normal end-of-capture path; exact home, PID, and claim-token ownership remains mandatory before the claim is released. If identity cannot be established for a live PID, or a surviving owned group cannot be proved stopped, the operation preserves the registration and claim for safe retry rather than adding a second owner. -A live PID whose identity no longer matches is a reused PID, so it is treated as stale and its process group is never signalled. +A live PID whose identity no longer matches is a reused PID, so cleanup refuses it before signalling. +Identity and process-group verification cannot be made atomic with signalling in portable shell: the reaper signals only a target it has verified as the recorded generation, but PID and group reuse remain possible in the narrow interval between verification and the signal. +Launch pacing is the primary host-wedge protection; watchdog cleanup is a backstop. Supported secondmate retirement preflights each target home's bounded `sweep-home` command before destructive teardown, snapshots its registrations outside the target, then runs the sweep at that home's final deletion or return boundary. If deletion or return fails, teardown restores those registrations and reconciles them before returning the refusal. @@ -753,6 +812,26 @@ The sweep retires local registrations and machine-wide claims whose recorded sta Teardown refuses with the home, lease, routing evidence, registrations, claims, and runners retained when identity is uncertain, ownership is unreadable or unreleased, or relevant state exists without a sweep-capable child script. Raw manual deletion of a Firstmate home is unsupported because it can orphan a blocking child. To recover, restore that home's tracked `bin/fm-procevent.sh`, run `FM_HOME=<home> <home>/bin/fm-procevent.sh sweep-home`, then rerun the supported teardown. +The owning-home lease below bounds how long such an orphan can run, but it is a backstop, not a substitute for the supported path. + +A runner is bound to the HOME that owns it, not to the one session that armed it. +That granularity is deliberate: a persistent source is meant to outlive the turn and the session that armed it, so binding a runner to its arming session would stop exactly the sources this mechanism exists to keep running. +Any activity in the same home refreshes the lease, so a replacement session, another watcher, or an ordinary inspection command keeps a runner of that home alive; a runner whose SOURCE is no longer wanted in a live home is stopped by reconcile when that source is retired, independently of the lease. +The lease is therefore the backstop for a home that is GONE - the torn-down test sandbox this change exists to bound - and not a per-session ownership check. +KNOWN LIMIT: while any activity continues in a home whose original owning session has ended, that activity refreshes the lease and a runner of that home keeps running until its source is retired or the home goes away. +Detaching a runner into its own process group is what lets a persistent source outlive the turn that armed it, and on its own it is also what lets a runner outlive its whole home: reparented to init, it keeps its blocking child - and every process that child spawns - running with nothing left to reap it. +So a home's process-event state carries a lease that registration, attached start, reconciliation, acknowledgement, and listing refresh, and the watcher's reconcile cycle is what keeps it fresh in a live home. +An attached public `start` continues refreshing the lease while its caller remains attached. +Each runner fails closed unless a small guard starts successfully beside it in a separate process group. +That guard accepts the lease only while the state root retains the device/inode identity recorded by the runner's claim, and stops the runner's whole process group after two consecutive checks cannot prove that identity and lease freshness. +The group signal reaches the blocking child and everything under it exactly as retirement does. +A runner exports the inherited `FM_PROCEVENT_IN_RUNNER` marker and every lease refresh is skipped under it, so a runner and its ordinary children do not certify their own owner, and the next reconcile in a live home simply starts a replacement runner. +That no-self-refresh rule is CONFUSED-AGENT-GRADE, the same deliberate captain-decided grade `bin/fm-lease-lib.sh` documents: it stops the accidental case this boundary exists for, an orphaned or test-scaffolding source tree that would otherwise keep its own owner alive. +A source that DELIBERATELY strips the marker from its environment can still refresh the lease, so adversarial-grade unforgeability is explicitly out of scope here and tracked as separate follow-up design work. +Scope is the owning state root and one runner generation, never a script or process name, so a live source in another home is untouched: that home refreshes its own lease. +`FM_PROCEVENT_OWNER_LEASE_SECONDS` (default 600, range 1..86400) is how long a runner keeps going with no sign of activity in its owning home, and `FM_PROCEVENT_OWNER_CHECK_SECONDS` (default 15, range 1..3600) is how often its guard re-reads the lease. +`FM_PROCEVENT_LAUNCH_FLOOR_SECONDS` (default 1, range 1..3600) is the minimum time between consecutive launches of one registration generation's stored command, bounding the launch rate of an immediately returning source during that lease window. +The generation's first launch is immediate, later launches share its monotonic pacing timestamp, a timestamp from before a reboot is treated as expired, and replacing the registration starts a fresh pacing generation. `FM_PROCEVENT_MAX_OUTPUT_BYTES` (default 1048576) bounds a single captured result while the source runs; oversized output is drained but truncated with a stderr notice rather than staged or published whole or dropped. @@ -802,6 +881,7 @@ FM_CONFIG_OVERRIDE= # alternate config dir, mainly for tests FM_PROC_ROOT_OVERRIDE= # alternate /proc root for Linux process-identity reads in fm-wake-lib.sh and fm-teardown.sh, mainly for tests FM_BACKEND= # optional runtime backend override for new spawns; tmux/herdr/zellij/orca/cmux support ship/scout spawns, codex-app is not accepted FM_TRACE_CONTEXT= # optional trace-context override; see "Trace context propagation" +FM_TASK_ID= # internal task-worker marker fm-spawn.sh exports into ship and scout panes, never set by hand; bin/fm-test-run.sh refuses to execute in the repository primary checkout while it is set HERDR_SESSION=default # herdr-only: named session for normal backend ops; not enough for destructive cleanup (docs/herdr-backend.md) FM_BACKEND_HERDR_SUBMIT_POLLS=6 # herdr-only: agent-state samples spread across each Enter attempt's budget when confirming a submit (docs/herdr-backend.md "Current transport behavior") FM_BACKEND_HERDR_SUBMIT_MIN_SLEEP=0.6 # herdr-only: minimum per-Enter confirmation budget before polling agent-state after an idle baseline @@ -840,6 +920,9 @@ FM_TOOL_UPDATE_BUDGET_SECS=20 # 1..120 seconds allowed for a whole watched-too FM_TOOL_UPDATE_NOW= # test override for the watched-tool sweep clock; the sweep budget still uses real time FM_PROCEVENT_MAX_OUTPUT_BYTES=1048576 # bound on one captured process-to-event result FM_PROCEVENT_CLAIM_ROOT= # machine-wide source claim root; default $XDG_STATE_HOME/firstmate/procevent-claims +FM_PROCEVENT_OWNER_LEASE_SECONDS=600 # how long a source runner keeps going with no activity in its owning home; 1..86400 +FM_PROCEVENT_OWNER_CHECK_SECONDS=15 # how often a runner's guard re-reads that lease; 1..3600 +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 # minimum interval between launches of one registration generation's source command; 1..3600 FM_WHEN_OUTPUT_TAIL_BYTES=8192 # bound on the command-output tail inside one condition->action outcome document FM_CODEX_WATCH_CHECKPOINT=180 # seconds per foreground watcher checkpoint in Codex primary supervision FM_CREW_STATE_NM_TIMEOUT=10 # seconds allowed per no-mistakes query inside fm-crew-state.sh @@ -857,7 +940,7 @@ FMX_X_THREAD_MAX=25 # maximum messages in one auto-split reply thread FMX_FOLLOWUP_MAX_AGE_SECS=604800 # local window for posting Relay completion follow-ups (7 days) FMX_FOLLOWUP_MAX_COUNT=3 # local cap on Relay completion follow-ups per linked mention FM_PF_RETRY_BACKOFF_SECS=900 # seconds before the next attempt after a retryable promised-public-reply delivery error -FM_LOCK_STALE_AFTER=2 # seconds before dead-pid lock records can be reclaimed; mid-acquire locks keep at least 2s grace +FM_LOCK_STALE_AFTER=2 # grace seconds for missing or nonnumeric lock-owner PIDs (minimum 2s); dead numeric PIDs have no age grace FM_GUARD_GRACE=300 # seconds before guard warnings, arm health checks, and the primary turn-end guard treat a watcher beacon as stale FM_CLAUDE_AUTOARM_ATTEMPTS=2 # bounded Stop-owned arm attempts per Claude auto-arm cycle; accepted values are 1, 2, or 3 FM_CLAUDE_AUTOARM_SYNC_WAIT_MS=800 # milliseconds the --claude turn-end guard waits for watcher health, an open Stop auto-arm generation claim, or a fresh epoch before deciding recovery ownership or failure progression @@ -873,15 +956,15 @@ FM_WATCH_REARM_RETRY_MAX_MS=4000 # Pi/OpenCode adapter cap for exponential con FM_WATCH_REARM_RETRY_LIMIT=5 # Pi/OpenCode adapter launch-failure retries before surfacing restoration failure FM_WATCH_CYCLE_LOG_MAX_BYTES=262144 # size cap for the arm-owned watcher lifecycle ledger FM_WATCH_CYCLE_LOG_KEEP_LINES=1000 # newest complete lifecycle rows considered when the ledger is capped -FM_WATCHER_STALE_GRACE=300 # defaults to FM_GUARD_GRACE; seconds a live watcher lock may have a stale beacon before re-arm errors +FM_WATCHER_STALE_GRACE=300 # defaults to FM_GUARD_GRACE if set, else the poll-derived grace (docs/turnend-guard.md "Guard grace and the poll cadence"); seconds a live watcher lock may have a stale beacon before re-arm errors FM_SIGNAL_GRACE=30 # seconds to coalesce nearby status and turn-end signals into one wake FM_TURNEND_CHURN_ABSORB_SECS=900 # longest one endpoint's bare turn-ends may be deferred on pane-churn evidence alone; only consulted when config/turnend-churn-absorb is present FM_CAPTAIN_RE='done:|needs-decision:|blocked:|failed:|PR ready|checks green|ready in branch|merged' # captain-relevant status regex; nonterminal progress verbs remain excluded even when their prose matches FM_CLASSIFY_PAUSED_VERB=paused # leading status verb for a declared external wait; excluded from FM_CAPTAIN_RE and distinct from blocked FM_STALE_ESCALATE_SECS=240 # idle seconds before a provably-working stale pane escalates; stale panes whose crew is not provably working surface immediately unless admitted directly to the declared-wait cadence, while a live idle declared wait still surfaces once before that cadence bounds repeats FM_BUSY_TURN_MAX_SECS=3600 # maximum age of a busy pane's latest state/<id>.turn-ended marker, or its state/<id>.meta spawn record before any turn completes, before the same wedge escalation used for a provably-working non-busy stale takes over; inspection-only, never an automatic interrupt or restart; a declared external wait or verified captain-held transfer takes the FM_PAUSE_RESURFACE_SECS recheck below instead -FM_PAUSE_RESURFACE_SECS=3600 # seconds between bounded rechecks of a declared external wait or verified captain-held transfer, including a live idle pane after its first inconclusive stale wake and a live busy pane past FM_BUSY_TURN_MAX_SECS; the away-mode daemon uses the same setting, ageing its window against the crew's own latest status line rather than pane busy state -FM_SECONDMATE_WAKE_STALL_SECS=60 # minimum age of the oldest valid foreign wake-queue row before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification; zero or invalid values use 60 +FM_PAUSE_RESURFACE_SECS=3600 # seconds between bounded rechecks of a declared external wait or verified captain-held transfer, and between repeated new-hash stale alarms for an ordinary crew task with an open backlog captain call; this includes a live idle pane after its first inconclusive stale wake and a live busy pane past FM_BUSY_TURN_MAX_SECS, while the away-mode daemon uses the same setting and ages its window against the crew's own latest status line rather than pane busy state +FM_SECONDMATE_WAKE_STALL_SECS=180 # minimum interval with no change of the oldest actionable foreign wake-queue row (it advances as the mate drains, and a queue reprovisioned under the same task id starts a fresh interval at whatever sequence it restarts) before an endpoint-recorded local secondmate produces one durable parent wake-loop-stall notification for that no-progress episode; a mate that is provably inside an active turn (an exact busy verdict, bounded by the same FM_BUSY_TURN_MAX_SECS above) never escalates whatever this interval says, declared external-wait pause rows are excluded, and zero or invalid values use 180 FM_WEDGE_DEMAND_INSPECT_COUNT=3 # consecutive provably-working stale escalations on the same unchanged pane before demand-deep-inspection is added FM_WORKTREE_WRITE_PRUNE='.git node_modules .venv venv __pycache__ .mypy_cache .pytest_cache .ruff_cache .tox target dist build .next .cache vendor' # directory names the wedge detector's task-worktree write probe skips; the default keeps .git out so a supervisor's own read-only git command can never look like crew progress; set it to the empty string to prune nothing, which widens the probe to the whole depth-bounded tree rather than disabling it FM_WORKTREE_WRITE_MAXDEPTH=6 # depth that same probe walks below the recorded worktree; it runs only at the moment a wedge escalation would otherwise fire, never on every poll; no probe knob applies to a secondmate, whose recorded worktree is a provisioned home the probe skips entirely diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index b4853aafc27..2b05276954a 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -208,6 +208,10 @@ "path": ".agents/skills/harness-adapters/references/harness/muse.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/harness-adapters/references/harness/omp.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/harness-adapters/references/harness/opencode.md", "audience": "agent-runtime" @@ -216,6 +220,10 @@ "path": ".agents/skills/harness-adapters/references/harness/pi.md", "audience": "agent-runtime" }, + { + "path": ".agents/skills/harness-adapters/references/harness/rovo.md", + "audience": "agent-runtime" + }, { "path": ".agents/skills/process-event-sources/SKILL.md", "audience": "agent-runtime" @@ -404,6 +412,10 @@ "path": "docs/supervision-protocols/grok.md", "audience": "agent-runtime" }, + { + "path": "docs/supervision-protocols/omp.md", + "audience": "agent-runtime" + }, { "path": "docs/supervision-protocols/opencode.md", "audience": "agent-runtime" @@ -432,6 +444,10 @@ "path": "docs/verification/dispatch-auth.md", "audience": "maintainer-verification" }, + { + "path": "docs/verification/lint-option-a.md", + "audience": "maintainer-verification" + }, { "path": "docs/verification/muse.md", "audience": "maintainer-verification" @@ -445,13 +461,17 @@ "audience": "maintainer-verification" }, { - "path": "docs/verification/secondmate-parent-channel.md", + "path": "docs/verification/rovo.md", "audience": "maintainer-verification" }, { "path": "docs/verification/runtime-backends.md", "audience": "maintainer-verification" }, + { + "path": "docs/verification/secondmate-parent-channel.md", + "audience": "maintainer-verification" + }, { "path": "docs/verification/stow-memory.md", "audience": "maintainer-verification" diff --git a/docs/extension-bindings.md b/docs/extension-bindings.md index 1884b2081cf..a8947afa942 100644 --- a/docs/extension-bindings.md +++ b/docs/extension-bindings.md @@ -216,7 +216,7 @@ The process-event durability boundary remains owned by [`configuration.md`](conf ## Runtime independence The host runs in the Firstmate home that owns the source, never in a task worker or its session container. -Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, and Muse therefore expose no package-loading surface for this capability. +Claude, Codex, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, and Rovo therefore expose no package-loading surface for this capability. The result reaches every supported primary through the existing bounded `check` wake path, including the unknown-protocol fallback used where no specialized primary continuation exists. The tmux, Herdr, Zellij, Orca, and cmux session providers are not consulted because a process-event source has no task endpoint. Remote and local secondmate homes bind and install independently, and the primary never executes a missing remote-home package locally. `remote-bind` carries one canonical `firstmate.extension-package-transfer.v1` JSON envelope over the existing bounded `fm-on` stdin/stdout job. Its hashed manifest pins the extension id, version, complete package-tree digest, entry count, total bytes, and byte-sorted entries. Entries are limited to normalized relative directories at mode 0755 and single regular files at mode 0644 or 0755, each with an exact size and SHA-256 payload digest. The receiver accepts at most 128 entries, 256 KiB per file, 512 KiB of package bytes, and 900,000 serialized bytes; it rejects malformed or truncated JSON, duplicate keys or paths, collisions, absolute or traversing names, links and special files, noncanonical modes, hash or size mismatches, and duplicate transfer identities. diff --git a/docs/fm-test-portable-shards.md b/docs/fm-test-portable-shards.md index dca996bf1cb..c966f1b9c27 100644 --- a/docs/fm-test-portable-shards.md +++ b/docs/fm-test-portable-shards.md @@ -50,7 +50,7 @@ The two parallel lanes use longest-processing-time assignment from those measure ## Portable serial remainder `portable-serial` includes every `tests/*.test.sh` that is neither proven-isolated nor `real-herdr-gated`. -It keeps watcher, lock, AFK, real tmux, daemon, secondmate lifecycle, bootstrap, live-harness opt-in, GUI-backend, and other unproven work serial. +It keeps watcher, lock, AFK, real tmux, daemon, secondmate lifecycle, bootstrap, the `live-harness-optin` family, GUI-backend, and other unproven work serial. Membership is derived rather than enumerated, so a newly added test lands here by default. ## Portable serial CI shards @@ -64,10 +64,10 @@ Each shard is still strictly serial in itself, and separate runners mean no two `.github/workflows/ci.yml` derives the same `n` from `strategy.job-total` rather than a literal, so changing the shard count in either file without the other fails the lane loudly instead of leaving part of the required suite unrun. Assignment is longest-processing-time bin packing over per-script duration hints embedded in `bin/fm-test-run.sh`. -The 139 current hints are the slowest measurements retained from the `fm-test-timing-portable-serial-*` artifacts of three green CI runs on 2026-09-01, [33558082172](https://github.com/kunchenguid/firstmate/actions/runs/33558082172), [33523597838](https://github.com/kunchenguid/firstmate/actions/runs/33523597838), and [33463326167](https://github.com/kunchenguid/firstmate/actions/runs/33463326167). -Those per-script maxima total 3825047 ms of conservative balance weight. -Taking the slowest of several runs rather than a single run keeps the balance honest on a slow runner: individual scripts varied by up to 20% between those three runs. -A script with no hint gets the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default; the current 140-script lane has one such script, bringing its assignment weight to 3852047 ms. +The 142 current hints include the slowest measurements retained from the `fm-test-timing-portable-serial-*` artifacts of three green CI runs on 2026-09-01, [33558082172](https://github.com/kunchenguid/firstmate/actions/runs/33558082172), [33523597838](https://github.com/kunchenguid/firstmate/actions/runs/33523597838), and [33463326167](https://github.com/kunchenguid/firstmate/actions/runs/33463326167), plus the 5121 ms native-Windows focused runner measurement for `tests/fm-pi-windows-shell-invocation.test.sh` from 2026-09-06T21:02Z. +Those per-script maxima total 3836189 ms of conservative balance weight. +Taking the slowest of several CI runs rather than a single run keeps the balance honest on a slow runner: individual scripts varied by up to 20% between those three runs. +A script with no hint gets the conservative `PORTABLE_SERIAL_DEFAULT_WEIGHT_MS` default; the current 152-script lane has ten such scripts, bringing its assignment weight to 4106189 ms. Hints only affect balance: the coverage guard keeps the partition complete and disjoint whatever they say, so a stale hint costs a slower shard rather than lost coverage. Balance is still worth keeping current, because enough unmeasured scripts let one shard carry more than twice another shard's real work and reach the job cap while another runner sits idle. That is not hypothetical: by 2026-09-01 the lane had grown from 116 to 139 scripts and from ~42 to ~63 minutes, 17 scripts were still unmeasured, and several hints were low by 2-5x, so shard 3 of 4 ran 17-20 minutes against its 20-minute cap while shard 1 ran 11.5 minutes and run [33574154856](https://github.com/kunchenguid/firstmate/actions/runs/33574154856) timed out seconds after a passing test. @@ -76,19 +76,19 @@ Refresh the hints whenever the serial lane gains scripts, rather than waiting fo | Lane | Script count | Estimated duration | |---|---:|---:| -| `portable-serial-1of5` | 27 | 770410 ms (~12.84 min) | -| `portable-serial-2of5` | 29 | 770416 ms (~12.84 min) | -| `portable-serial-3of5` | 30 | 770417 ms (~12.84 min) | -| `portable-serial-4of5` | 26 | 770405 ms (~12.84 min) | -| `portable-serial-5of5` | 28 | 770399 ms (~12.84 min) | -| imbalance | | 18 ms | - -The current table is generated from the runner's retained maxima plus its default for the one unhinted script. +| `portable-serial-1of5` | 29 | 821231 ms (~13.69 min) | +| `portable-serial-2of5` | 30 | 821243 ms (~13.69 min) | +| `portable-serial-3of5` | 31 | 821236 ms (~13.69 min) | +| `portable-serial-4of5` | 31 | 821247 ms (~13.69 min) | +| `portable-serial-5of5` | 31 | 821232 ms (~13.69 min) | +| imbalance | | 16 ms | + +The current table is generated from the runner's retained maxima plus its default for the ten unhinted scripts. The last complete replay against the three source runs put the then-current partition's worst shard at 12.54 min, 63% of the 20-minute job cap. The single longest script, `tests/fm-watch-triage.test.sh` at 262626 ms, is the floor for any shard count. -Refresh the hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: +Refresh the CI-derived hints by downloading the per-shard timing artifacts from several green CI runs, replacing the `portable_serial_weight_hints` table in `bin/fm-test-run.sh` with the slowest measured `duration_ms` per `path`, and updating the table above: ```sh for run in <run-id> <run-id> <run-id>; do @@ -101,6 +101,7 @@ bin/fm-test-run.sh --check-coverage ``` A timed-out shard uploads no artifact, so pick runs where every serial shard is green or the lane's slowest scripts go unmeasured in exactly the shard that needs them most. +Measure native-Windows-only scripts through the focused Git Bash runner and retain that `duration_ms` separately, because the portable CI shards skip them. ## Coverage guard @@ -125,7 +126,7 @@ Portable shards, each portable serial shard, and the Herdr lane upload runner-ge | Lane | Bound | Rationale | |---|---|---| | portable parallel 1/2 | job `timeout-minutes: 10` | The measured shard sums are about three minutes and the timeout is a hang tripwire. | -| portable serial 1-5 | job `timeout-minutes: 20` | Each balanced shard carries about 12.84 minutes of conservative assignment weight, leaving roughly 1.6x hang-tripwire margin for job setup and runner-speed spread. | +| portable serial 1-5 | job `timeout-minutes: 20` | Each balanced shard carries about 13.69 minutes of conservative assignment weight, leaving roughly 1.5x hang-tripwire margin for job setup and runner-speed spread. | | Herdr | family-run step `timeout-minutes: 20`; job `timeout-minutes: 75` backstop | Healthy runs finished around 7 minutes before this lane gained `fm-backend-herdr-focus-flash-e2e`, which measures about 2 minutes against a real lab locally, so the step bound is still the hang tripwire (cleanup and timing artifacts still upload) while the job cap stays a last-resort backstop. Refresh this figure from the lane's uploaded timing artifact. | Timeouts are hang tripwires rather than expected healthy durations. diff --git a/docs/herdr-backend.md b/docs/herdr-backend.md index 1b55881dffc..e40de497ba6 100644 --- a/docs/herdr-backend.md +++ b/docs/herdr-backend.md @@ -1,6 +1,6 @@ # Herdr runtime backend -Herdr is an experimental agent-native terminal backend with native per-pane agent state and push events. +Herdr is an agent-native terminal backend with native per-pane agent state and push events. Firstmate requires Herdr protocol 14 or newer; broad backend verification covers versions 0.7.1, 0.7.3, 0.7.4, 0.7.5, and 0.8.0, while protocol-16 features remain gated by availability. Default-on presentation spaces have a higher floor of Herdr 0.8.0 for the reason given under [Presentation spaces](#presentation-spaces). Herdr provides the terminal session while Treehouse continues to provide task worktrees. @@ -8,7 +8,7 @@ Herdr provides the terminal session while Treehouse continues to provide task wo ## Setup -Pick Herdr when you want native busy, idle, and blocked state and accept the experimental limits below. +Pick Herdr when you want native busy, idle, and blocked state and accept the active limits below. Prerequisites: @@ -340,7 +340,6 @@ Tests use thin compatibility wrappers in `tests/herdr-test-safety.sh` and never ## Active limits -- Herdr remains experimental. - Presentation ordering needs protocol 16 and Python and is best-effort only. - Mutable labels can collide; they are never placement or destructive authority. - A Firstmate outside Herdr cannot resolve a launcher workspace, so a colliding home label refuses new spawns until the collision is cleared. diff --git a/docs/pi-supervision-branch.md b/docs/pi-supervision-branch.md index 0b0c1580786..76cb84a8b85 100644 --- a/docs/pi-supervision-branch.md +++ b/docs/pi-supervision-branch.md @@ -7,7 +7,7 @@ This document stays the owner and the contract. Fleet supervision on the Pi primary harness runs on a second conversation - the supervision branch - inside the same `pi` process as the captain's chat. Supervision is default-on: once a Pi primary session owns this home's fleet lock, the branch handles eligible task-local rows from ordinary actionable wakes plus heartbeat scans that the cheap bash-level scan flags as possibly captain-relevant, then merges each outcome back into the captain conversation's transcript. -Ordinary main-only rows remain on main even when eligible task-local rows share their queue. +Ordinary main-only rows remain on main even when eligible task-local rows share their queue, except that a decision-owned signal or stale trigger keeps its entire coalesced trigger batch on main. An unresolvable row makes the scan unsafe and returns the whole wake to main, and every watcher-failure alarm also stays on main. Captain-relevant branch outcomes persist as exact, sequence-keyed visible transcript entries and then open one sequence-keyed processing turn on main, which stays open until main acknowledges that sequence. The design source is the captain-approved forked-supervision architecture board, a captain-private fleet record (a self-contained HTML explainer with the measured cache and judgment evidence); this document records the shape it landed as, and the delivering PR cites the board artifact itself. @@ -23,6 +23,13 @@ The supervision branch itself is Pi-only by construction: - Wake dispatch: `.pi/extensions/fm-primary-pi-watch.ts` stays the dispatcher; `.pi/extensions/lib/fm-branch-dispatch.ts` owns the offer handshake and row eligibility, while [`watcher-continuity.md`](watcher-continuity.md#per-actor-acknowledgement) owns the per-actor consume contract. A successful row grant transfers ownership of exactly the currently branch-eligible rows to the branch; a check-kind triggering close (merge-confirmation polls, Relay mentions, credential/auth failures, and every other legitimately main-only class) is never offered even when other rows are eligible, no acceptor (extension absent, away mode, branch broken) keeps today's wake-to-main path for that close, and watcher-failure alarms always go to main because only main can repair the watcher cycle. + A decision-owned event surfaced by `bin/fm-watch.sh`'s signal path gets the identical treatment even though it keeps the ordinary `signal` kind. + `signal_files_actionable` marks the queued payload `needs-decision:` for a newly surfaced `needs-decision`, a `captain-held` declaration surfaced through the no-verb fallback, or a pending-reply second-mate escalation; `scopeForUnreadWake` excludes every marked row from what the branch may claim. + For a stale row, `scopeForUnreadWake` folds the mapped task's status log and excludes the row when any `needs-decision` remains open or the current meaningful declaration is `captain-held`; an unreadable or symlinked status log fails the scope closed rather than influencing routing. + The dispatcher resolves trigger keys and every currently unread excluded decision row to task identity before cross-referencing them: any signal or stale trigger containing a decision-owned task goes wholly to main, including a batch that also contains routine rows, and an unread decision for one task keeps every later signal or stale trigger for that same task on main until the decision row is read, regardless of whether the rows use its status-file key or window alias. + Other tasks remain independently eligible. + The wake message itself retains its existing shape, so other harness-arm scripts remain unchanged. + Heartbeat handling remains independent. A fleet-wide heartbeat keeps its own all-or-nothing rule (see "Heartbeat routing" below): it takes every branch-ownable unread row or none of them. A co-present main-owned check row no longer defers that review to main, because it is not fleet context the branch is missing and main is woken for it on its own triggering close. - The branch itself: `.pi/extensions/fm-branch-supervision.ts` creates the branch session, serializes wakes, mirrors dialog, and merges outcomes. @@ -51,9 +58,9 @@ The supervision branch itself is Pi-only by construction: - Consistency: `bin/fm-lease-lib.sh` owns the per-task lease contract, the main-only role partition, and the deliberate CONFUSED-AGENT-GRADE threat model these guards target (captain-decided; adversarial-grade separation is out of scope and tracked as follow-up design work); `bin/fm-lease.sh` is the command surface. The guards are wired into `fm-send.sh`, `fm-control.sh`, and `fm-teardown.sh` (overlap, lease-checked, with claim serialization retained through the mutation) and `fm-pr-merge.sh`, `fm-merge-local.sh`, and `fm-spawn.sh` (main-owned, branch refused; a relaunch through `fm-control` stays branch-legal recovery). - Autonomy: supervision is default-on for every task once a Pi primary session owns the fleet lock (docs/configuration.md "Pi supervision branch"); no captain grant file is required. - A fleet-wide heartbeat is separately eligible only when every non-check row in the unread queue is a heartbeat row or a resolvable task-local row (see "Heartbeat routing" below); every other fleet-wide or unresolvable wake, and every watcher-failure alarm, stays on main. + A fleet-wide heartbeat is separately eligible only when every row other than a check or decision-owned signal/stale row is a heartbeat row or a resolvable task-local row (see "Heartbeat routing" below); every other fleet-wide or unresolvable wake, and every watcher-failure alarm, stays on main. The branch recomputes eligibility immediately before prompting the branch to drain and publishes the exact eligible row set to `state/.branch-eligible-rows` through `writeEligibleRowsSnapshot`. - A newly-arrived main-owned row observed at that recheck no longer defers the whole queue to main: it is excluded from the eligible set, so whatever else is currently eligible still reaches the branch, and the main-owned row stays queued for main's own later drain. + After an independently eligible wake has already been offered, a newly-arrived main-owned row observed at that pre-drain recheck does not revoke the offer: it is excluded from the eligible set, so whatever else is currently eligible still reaches the branch, and the main-owned row stays queued for main's own drain. [`watcher-continuity.md`](watcher-continuity.md#per-actor-acknowledgement) owns the consume-side guarantee that neither actor can present or acknowledge the other's claim. Heartbeat keeps its own all-or-nothing recheck over the rows it can claim: it takes every branch-ownable unread row or none of them, and an unresolvable task-local row still defers the whole review to main. A producer can still append a row in the instant between that final check and drain startup; this accepted residual follows the confused-agent-grade boundary above rather than claiming adversarial queue isolation. @@ -121,10 +128,10 @@ Main can read the durable outcome store on demand through its `fm_branch_outcome ## Heartbeat routing The cheap bash-level heartbeat scan absorbs a genuinely no-op pass before it reaches Pi, unchanged from before. -Only a scan already flagged as possibly captain-relevant emits the bare `heartbeat` wake; `.pi/extensions/fm-primary-pi-watch.ts` flags that offer `heartbeat: true`, and the branch accepts it without a project only when every non-check row observed in the unread-queue eligibility check is either heartbeat-kind or a resolvable task-local signal or stale event. +Only a scan already flagged as possibly captain-relevant emits the bare `heartbeat` wake; `.pi/extensions/fm-primary-pi-watch.ts` flags that offer `heartbeat: true`, and the branch accepts it without a project only when every branch-ownable row observed in the unread-queue eligibility check is either heartbeat-kind or a resolvable task-local signal or stale event. -A heartbeat is never vetoed or ridden into main by a co-present check row. -A check row is permanently main-owned in every mode: it is excluded from what the branch may claim and left queued for main, which is woken for it on that check's own watcher cycle, so nothing starves by being left behind. +A heartbeat is never vetoed or ridden into main by a co-present check row or decision-owned signal/stale row. +Those rows are permanently main-owned in every mode: they are excluded from what the branch may claim and left queued for main, which is woken for each on its own watcher cycle, so nothing starves by being left behind. Deferring the fleet review to main merely because some unrelated merge poll or Relay mention happened to be sitting unread put a routine review in the captain's chat for a reason that had nothing to do with the fleet, and that coupling is gone. What all-or-nothing still guarantees is unchanged: the branch takes every branch-ownable unread row or none of them, and an unresolvable task-local row, an unknown row kind, or an unreadable queue still defers the whole review to main. The branch runs its normal operating procedure for the wake (`bin/fm-branch-prompt.sh` "Handling a wake") and performs the deeper fleet review that main previously performed. @@ -137,6 +144,8 @@ Every other fleet-wide or unresolvable wake - including watcher-failure alarms, The captain accepted the normal provider prompt-caching strategy: a byte-identical branch prefix generated once per firstmate version, the same tool set in the same order on every request, and one shared `prompt_cache_key` per home for all branch sessions (set in a `before_provider_request` hook, and only for providers whose requests already carry that field); main keeps its own per-session key. Budget roughly 60% cache hits on a new branch conversation's first call and 95% on later calls within that conversation; the shared per-home key is what carries the byte-identical prefix across the conversation each main session start opens, and reuse is best-effort, never guaranteed. The branch can also run on a cheaper model and a shallower reasoning effort than main, both pinned with the Pi `/supervision-model` command; [configuration.md](configuration.md#pi-supervision-branch-model-and-effort-configsupervision-branch-model-configsupervision-branch-effort) owns those pins' operator-facing schema and unpinned behavior. +A provider an extension registered only into main's runtime, such as pi-devin-auth's `devin`, reaches the isolated branch runtime by copying its provider config from main's captured `ModelRegistry` into the branch `ModelRuntime` at model-resolution time and in the `/supervision-model` picker, so the provider's own `streamSimple` transport and OAuth wiring are reused by reference rather than reimplemented. +That carve-out is scoped to provider registration alone: the branch keeps its `noExtensions`, `noSkills`, and `noContextFiles` isolation, the copy is never persisted, a provider whose registration fails to compose is simply unavailable, and `tests/fm-pi-branch-extension.test.sh` pins the pin-and-fallthrough behavior. No caching machinery beyond this exists, deliberately: any later dynamic content in the branch prefix silently removes most of the cache benefit, which is why `bin/fm-branch-prompt.sh`'s header is the contract's single owner and `tests/fm-branch-supervision.test.sh` pins the output to byte identity. ## Away mode @@ -146,12 +155,13 @@ What is new is only the attended path: outside away mode, the branch absorbs the ## Verification -Portable regressions: `tests/fm-pi-branch-extension.test.sh` covers dispatch, signal and stale report scoping with unscoped heartbeat reports, the new branch conversation at every main session start with continuation inside one session, the mirror re-anchor that pairs with it, requested-versus-unsolicited delivery, exact visible entry content, no unkeyed model turn, the sequence-keyed processing request and its acknowledgement, re-presentation after an empty reply and after an unrelated prior answer, the triggered-then-next-turn pacing, session-start re-presentation, routine outcomes staying turn-free, the processed-marker migration, idle and busy main state, incident-shaped compaction and unrelated-assistant context, cold-start post-lock recovery, crash-before-cursor reload recovery, repeated-reload idempotency, mirroring, post-construction provider-error and no-report fallback, the consecutive-error latch, cooldown probe, exponential backoff, report-plus-settlement recovery, report-before-error re-latch, cache key, and model and effort selection. +Portable regressions: `tests/fm-pi-branch-extension.test.sh` covers dispatch, signal and stale report scoping with unscoped heartbeat reports, the new branch conversation at every main session start with continuation inside one session, the mirror re-anchor that pairs with it, requested-versus-unsolicited delivery, exact visible entry content, no unkeyed model turn, the sequence-keyed processing request and its acknowledgement, re-presentation after an empty reply and after an unrelated prior answer, the triggered-then-next-turn pacing, session-start re-presentation, routine outcomes staying turn-free, the processed-marker migration, idle and busy main state, incident-shaped compaction and unrelated-assistant context, cold-start post-lock recovery, crash-before-cursor reload recovery, repeated-reload idempotency, mirroring, post-construction provider-error and no-report fallback, the consecutive-error latch, cooldown probe, exponential backoff, report-plus-settlement recovery, report-before-error re-latch, cache key, model and effort selection, and (in `test_branch_dispatch_classifies_main_only_rows_and_writes_the_eligible_snapshot`) decision-owned signal and stale rows' exclusion from `eligibleSeqs`, their presence in `needsDecisionKeys`, task alias resolution, reserved-key configuration, status-log race and symlink refusal, non-vetoing behavior for unrelated eligible rows, and decision-only queues reading as ordinary main-only absence. `tests/fm-branch-supervision.test.sh` covers prompt stability, store append-only behavior, the captain cursor barrier, the processed marker's sequence bounds, leases, guards, and non-branch-home invariance. `tests/fm-wake-drain-outcome-backstop.test.sh` covers keyless resurfacing, causal suppression, same-second ordering, one-shot presentation, first-drain index self-healing under the outcome lock, store-fault fail-closed behavior, bounded history cost and output, and the oversized-line limit. `tests/fm-teardown.test.sh` covers removal of the retired task's outcome index and the append-side rule that a post-teardown report does not recreate it. -The branch-offer, heartbeat-offer, heartbeat-not-ridden-by-a-check, and main-only-check-class tests remain in `tests/fm-pi-watch-extension.test.sh`, the recovery test remains in `tests/fm-session-start.test.sh`, and the per-actor consume regression remains in `tests/fm-wake-queue.test.sh`. +The branch-offer, heartbeat-offer, heartbeat-not-ridden-by-main-only-rows, main-only-check-class, captain-held-stale-stays-on-main, and mixed-signal-routing tests remain in `tests/fm-pi-watch-extension.test.sh` (the last two routing classes exercise `offerWakeToBranch`'s trigger-key cross-reference end to end), the recovery test remains in `tests/fm-session-start.test.sh`, and the per-actor consume regression remains in `tests/fm-wake-queue.test.sh`. It also covers the off-thread delivery contract behaviorally: that a delivery leaves the event loop running rather than blocking it, that interleaved reports stay ordered and exactly once, that a session replaced mid-delivery neither loses nor duplicates an outcome, and that a failing store script surfaces without losing or doubling one. +`tests/fm-watch-triage.test.sh` covers `bin/fm-watch.sh`'s side of the contract end to end: needs-decision, no-verb captain-held, and pending-reply second-mate escalation signal rows are marked `needs-decision:`, a needs-decision whose key transition was rejected by the reserved-key vocabulary (`fm-classify-lib.sh`'s `reconciliation-required:` wrapper) is still marked, and ordinary blocked or captain-relevant signals stay unmarked. Live guards: `FM_PI_BRANCH_LIVE_E2E=1 tests/fm-pi-branch-live-e2e.test.sh` exercises the real installed Pi SDK's immediate active-transcript appendEntry rendering, persistence, custom-entry model exclusion, branch-session surfaces, and watcher-owned fallback after rejected branch settlement. `FM_PI_BRANCH_RESPONSIVENESS_E2E=1 tests/fm-pi-branch-responsiveness-live-e2e.test.sh` answers the question only a real TUI can: it types into an isolated Pi pane while outcomes are delivered and fails if keystroke echo leaves the class of the same machine's extension-free floor. Record dated current results in [docs/verification/runtime-backends.md](verification/runtime-backends.md). diff --git a/docs/scripts.md b/docs/scripts.md index 90dff6df5df..1e0b231ba48 100644 --- a/docs/scripts.md +++ b/docs/scripts.md @@ -33,12 +33,12 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-captain-hold.sh` | Hold tasks for the captain, record the captain's answers, gate investigation completion, and report record divergence between the status log and the backlog | | `fm-decision-hold.sh` | One-release compatibility shim mapping the retired decision commands onto fm-captain-hold.sh | | `fm-brief.sh` | Scaffold ship (explicit `--mode`), scout, secondmate-charter, and Herdr-lab briefs, with Captain's intent and Firstmate spec subsections on ship/scout | -| `fm-dod-lib.sh` | One owner of the ship definition of done and of the no-mistakes `--intent` contract | +| [`fm-dod-lib.sh`](../bin/fm-dod-lib.sh) | Own ship/scout worker role scope, ship definitions of done, and the no-mistakes `--intent` contract | | `fm-herdr-lab.sh` | Provision and guardedly operate an isolated, never-default Herdr lab session | | `fm-install-herdr.sh` | Install CI's exact-version Herdr pin with official asset URL, SHA-256, and protocol checks | | `fm-install-treehouse.sh`| Install CI's exact-version Treehouse pin for real-Herdr E2E that needs spawn worktrees | | `fm-herdr-ci-cleanup.sh` | Snapshot and tear down only job-owned `fm-lab-*` sessions in the Herdr CI lane | -| `fm-test-run.sh` | Behavior-test runner: selection, portable lanes, bounded concurrency, budgets, coverage guard, timing/JSON | +| `fm-test-run.sh` | Behavior-test runner: selection, portable lanes, bounded concurrency, budgets, coverage guard, timing/JSON; refuses to execute in the repository primary checkout when `FM_TASK_ID` marks a task worker | | `fm-test-isolation-proof.sh` | Concurrent isolation harness and portable candidate set owner | | `fm-ensure-agents-md.sh` | Ensure a project's real `AGENTS.md`, its `CLAUDE.md` `@AGENTS.md` pointer, and self-governance guidance (explicit project mark documented in the helper's header and help) | | `fm-guard.sh` | Warn on primary-checkout tangles, main-session pending wakes, and unhealthy supervision | @@ -61,7 +61,7 @@ The shared no-mistakes gate refusal for fleet lifecycle entrypoints is summarize | `fm-backend-hometag-lib.sh` | Shared per-installation home-tag derivation for zellij tab and cmux workspace titles | | `fm-composer-lib.sh` | Single fleet-wide owner of composer shapes, capability-aware screen classification, and verdicts | | `backends/tmux.sh` | Verified tmux session-provider adapter | -| `backends/herdr.sh` | Experimental herdr session-provider adapter | +| `backends/herdr.sh` | Herdr session-provider adapter with its own required CI lane | | `backends/zellij.sh` | Experimental zellij session-provider adapter | | `backends/orca.sh` | Experimental Orca backend adapter owning both worktree and terminal | | `backends/cmux.sh` | Experimental cmux session-provider adapter | diff --git a/docs/sessionstart-nudge.md b/docs/sessionstart-nudge.md index 7680d37b790..7b93d579618 100644 --- a/docs/sessionstart-nudge.md +++ b/docs/sessionstart-nudge.md @@ -7,7 +7,7 @@ Firstmate ships two session-open tiers, and the tier is a property of the harnes | Tier | What the adapter does | Used by | | --- | --- | --- | -| Run | Executes `bin/fm-session-start.sh` through the native session-open adapter and gates its ordered digest into model context before the first turn. | Claude, `codex exec`, GitHub Copilot CLI, Pi / pi-signed, Cursor | +| Run | Executes `bin/fm-session-start.sh` through the native session-open adapter and gates its ordered digest into model context before the first turn. | Claude, `codex exec`, GitHub Copilot CLI, Pi / pi-signed, omp, Cursor | | Nudge | Asks the agent to run the digest through the native adapter or the tracked session-start instruction. | Grok, OpenCode, and run-tier sources routed to the nudge | Codex's interactive TUI has no tracked session-open, compaction, or re-emit channel and is not covered by either tier. @@ -76,8 +76,9 @@ A lock another session holds and a truncated digest therefore surface as digest | OpenCode | Nudge | `.opencode/plugins/fm-primary-sessionstart-nudge.js` listens for `session.created`, runs once per session id, and calls `client.session.promptAsync` only when the wrapper prints a nudge. | Interactive TUI delivery is supported, while headless `opencode run` is intentionally fail-open and cannot use the run tier because the process can exit before the queued turn. | | Grok | Nudge | `.grok/hooks/fm-primary-sessionstart-nudge.json` registers a project `SessionStart` hook and invokes the wrapper through inline-defaulted `${GROK_WORKSPACE_ROOT:-}`. | The project hook runs when the checkout is trusted, but Grok currently discards hook stdout from model context, so this path is intentionally fail-open and cannot use the run tier. | | GitHub Copilot CLI | Run | `.github/hooks/firstmate.json` registers `sessionStart` and invokes `bin/fm-ghcp-hook.sh session-start` through Bash or the PowerShell-to-Git-Bash transport. | The adapter returns the complete digest as `additionalContext`; it is inert unless `COPILOT_CLI=1`, so repository hooks do not take over non-CLI Copilot surfaces. | -| Cursor | Run | `.cursor/hooks.json` registers `sessionStart`, anchored through `$CURSOR_PROJECT_DIR` with a 180s timeout, invoking `bin/fm-sessionstart-cursor.sh`. | Cursor's payload has no `source` field, so the registration supplies `--source` itself and returns the digest as `additional_context`, while project hooks load only when the workspace is launched with `--trust`. | -| Cursor compaction | Uncovered | None. | Cursor's `preCompact` response can return only `user_message` and is absent from Cursor's `additional_context` step set, so re-emit delivery remains deferred and a Cursor primary does not receive a fresh digest after compaction. | +| Cursor | Run | `.cursor/hooks.json` registers `sessionStart`, anchored through `$CURSOR_PROJECT_DIR` with a 180s timeout, invoking `bin/fm-sessionstart-cursor.sh`. | Cursor's payload has no `source` field, so the registration supplies `--source` itself, and the adapter returns the digest as `additional_context`. Project hooks load only when the workspace is launched with `--trust`. | +| omp | Run | `.omp/extensions/fm-primary-turnend-guard.ts`, auto-discovered from the home with no trust gate, starts the wrapper at `session_start` and has `before_agent_start` await it and return one persistent context message before the first provider call, exactly as Pi's does; `session_compact` is the compaction equivalent. | omp's `session_start` carries no reason field (verified 18.1.11), so the source is derived following the Cursor precedent: the first start of the process is `startup`, or `resume` when the launch line carried `--continue`/`-c` or `--resume`/`-r`; a later in-process start (`/new`, `/resume`, `/fork`) is `clear`, which re-emits only when this lock owner completed a full startup. `before_agent_start` message delivery was verified to reach model context on 18.1.11. | +| Cursor compaction | Uncovered | None. | Cursor's `preCompact` response can return only `user_message` and is absent from Cursor's `additional_context` step set, so it cannot inject a re-emit digest. Delivering one needs its own design and is deliberately deferred to a follow-up; a Cursor primary does not re-emit its digest after a compaction. | Cursor's `sessionStart` fires at every session open with no source distinction, including a resumed session, so a resume re-runs the full digest; that is redundant and idempotent rather than a lost helm. Cursor's compaction surface is uncovered in the same sense as Codex's interactive TUI above: Firstmate registers nothing for `preCompact`, so a compacted Cursor session keeps whatever context survived rather than receiving a fresh digest. diff --git a/docs/subagent-guard.md b/docs/subagent-guard.md index c3a0f139a63..fd0aa544d2b 100644 --- a/docs/subagent-guard.md +++ b/docs/subagent-guard.md @@ -176,7 +176,7 @@ When that script is absent the message still defers to intake classification and ## Harness wiring -Every supported primary harness was reviewed. +Every supported primary harness was reviewed except omp, whose row below rests on its bundled material rather than a live enumeration. Applicability turns on one question: does the harness expose built-in delegation tools that a primary session could use instead of `bin/fm-spawn.sh`? | Harness | Delegation surface | Status | @@ -185,6 +185,7 @@ Applicability turns on one question: does the harness expose built-in delegation | Codex | none | Not applicable, verified empirically below. Codex 0.144.1 exposes no subagent, sub-task, or delegated-agent tool, so there is nothing to remove or intercept. `.codex/hooks.json` is unchanged. | | GitHub Copilot CLI | `Agent` and `Task` | Scoped guard wired in `.github/hooks/firstmate.json`; both names reach the same native Copilot deny renderer through matcher `Agent|Task`. | | Grok | present, exact tokens unconfirmed | Not wired pending live verification. See below. | +| omp | present, per bundled material | Not wired and unverified. omp ships a built-in task delegation tool: its bundled docs list `tools/task.md` and the captain-level `task.maxConcurrency` setting governs it. No Firstmate delegation seatbelt is wired for it yet, and its status stays unverified until a live tool enumeration is recorded the way the Codex row was. | | OpenCode | present, exact tokens unconfirmed | Not wired pending live verification. See below. | | Pi | none reported | Not wired pending live verification. See below. | diff --git a/docs/supervision-protocols/grok.md b/docs/supervision-protocols/grok.md index f27ae302e13..305e1802a16 100644 --- a/docs/supervision-protocols/grok.md +++ b/docs/supervision-protocols/grok.md @@ -25,7 +25,7 @@ When you see a background-task-completed system reminder for the arm: 1. Run `bin/fm-wake-drain.sh` first. 2. Optionally fetch arm output with `get_command_or_subagent_output(<task_id>)` for the reason line. 3. Handle `signal`, `stale`, `check`, or `heartbeat` using the harness-neutral contract in `AGENTS.md`. -4. Ordinary wake: re-arm the next cycle with the same background `bin/fm-watch-arm.sh` call if work remains in flight or Relay still needs polling. +4. Ordinary wake: re-arm the next cycle with the same background `bin/fm-watch-arm.sh` call if the home still needs supervision, as `bin/fm-supervision-lib.sh` defines it. 5. Do not invent a wake from an attach-status line alone. Drain the queue and act only on real wake records, the drain's `OPEN DECISIONS` and `UNREAD STATUS` entries, or a real watcher reason line. Re-arm attaches to an existing healthy cycle when one is already present and follows its verified successor chain. diff --git a/docs/supervision-protocols/omp.md b/docs/supervision-protocols/omp.md new file mode 100644 index 00000000000..eddacc3ff6d --- /dev/null +++ b/docs/supervision-protocols/omp.md @@ -0,0 +1,30 @@ +Mode: omp (Oh My Pi) extension background wake. + +When this session owns supervision and away mode is not active: +1. Drain first with `bin/fm-wake-drain.sh`. + After handling all emitted wakes and reconciling open decisions and unread status lines, run the exact `--ack-through` command printed as `WAKE_ACK_REQUIRED`; until then the work remains durable for idempotent re-handling after interruption. +2. Confirm the omp primary auto-loaded both project extensions from `.omp/extensions/`; omp has no project-trust gate, so a plain `omp` started with this home as its working directory loads them with no dialog. + If `bin/fm-session-start.sh` reported the omp extensions as not loaded, restart omp inside this home; pass `-e __FM_OMP_TURNEND_EXT__ -e __FM_OMP_EXT__` only when omp must start from another directory, because omp loads a file named both ways twice. +3. Initial process cycle only: make the one required `fm_watch_arm_omp` call; if startup already owned the fleet lock, this is an ownership-based no-op. + Use `/fm-watch-arm-omp` only as a human-entered fallback. + Never run `bin/fm-watch-arm.sh` through omp's bash tool because that foreground arm can wedge the agent and bypasses extension-owned cleanup. +4. If the extension says no live session holds the lock, run `bin/fm-session-start.sh` to reclaim the session lock, then call `fm_watch_arm_omp` again. +5. The extension starts `bin/fm-watch-arm.sh --restart`, keeps the child attached to the live omp process, and owns every later successor launch. +6. Ordinary same-process session replacement (`/new`, `/resume`, `/fork`) retires only the prior generation; when the replacement owns the fleet lock, its `session_start` arms the new generation without a model turn or another `fm_watch_arm_omp` call. + The generation-owner contract and in-flight actionable-close handoff live in `.omp/extensions/fm-primary-omp-watch.ts`; because omp reports no shutdown reason, every shutdown with a pending actionable close persists the handoff, and the next owning `session_start` in any process replays it. +7. After an actionable child close, the extension rechecks session-lock ownership and verifies one successor before it delivers the follow-up wake; its bounded fallback is defined in `docs/watcher-continuity.md`. +8. Ordinary work, turn completion, and ordinary signal, stale, check, heartbeat, or other wake handling: do not call `fm_watch_arm_omp` again because continuity is extension-owned rather than model-memory-owned. +9. An unexpected child close enters bounded exponential retry, and an exhausted retry or lost session lock is surfaced as a watcher failure instead of disappearing. +10. Missing, failed, or unhealthy cycle only: if a later notification explicitly reports one of those repair conditions, drain queued wakes, inspect the failure text, call `fm_watch_arm_omp`, and restart omp inside this home if the extensions are not loaded. + A redundant call while the extension owns an arm child or scheduled retry is an ownership-based `watcher: unchanged` no-op, not an independent health claim. +11. Never use shell `&` for watcher supervision. + The arm mechanism above is extension-owned, not a model tool call, but a manual recovery probe that backgrounds, pipes, or bundles the arm is denied automatically by the pre-tool seatbelt (`bin/fm-arm-pretool-check.sh`, wired into the turn-end guard extension at `__FM_OMP_TURNEND_EXT__`). + +The turn-end guard on omp is structural, not advisory: `__FM_OMP_TURNEND_EXT__` answers omp's blocking `session_stop` hook, and when `bin/fm-turnend-guard.sh` returns 2 it forces one continuation carrying the guard text, bounded to one per turn by the `stop_hook_active` flag omp sets on the continuation's own stop. +An interrupted turn never raises `session_stop`, so a supervisor-initiated interrupt is not guarded; `bin/fm-control.sh` owns that postcondition. + +The Pi supervision branch (`docs/pi-supervision-branch.md`) is out of scope for the omp primary: every actionable wake is delivered to this conversation, exactly as on Claude, and the lease, outcome-store, and `fm_branch_processed` contracts do not apply here. + +The turn-end guard extension lives at `__FM_OMP_TURNEND_EXT__`. +The watcher extension lives at `__FM_OMP_EXT__`. +Both are tracked, project-local `.omp/extensions/*.ts` files that omp auto-discovers from this home with no trust dialog; `bin/fm-session-start.sh` reports when the running omp session has not loaded both required extensions. diff --git a/docs/supervision-protocols/pi.md b/docs/supervision-protocols/pi.md index b01f5c514c9..1f9a1ea6d44 100644 --- a/docs/supervision-protocols/pi.md +++ b/docs/supervision-protocols/pi.md @@ -20,6 +20,7 @@ When this session owns supervision and away mode is not active: The arm mechanism above is extension-owned, not a model tool call, but a manual recovery probe that backgrounds, pipes, or bundles the arm is denied automatically by the PreToolUse seatbelt (`bin/fm-arm-pretool-check.sh`, wired into the turn-end guard extension at `__FM_PI_TURNEND_EXT__`). The supervision branch is default-on (docs/pi-supervision-branch.md): whenever this session owns the fleet lock and away mode is not active, the watcher extension hands eligible task-local rows from ordinary actionable wakes, plus selected fleet-wide heartbeat reviews, to the in-process supervision branch while main-only rows remain queued for this conversation. +Decision-owned signal and stale routing, including whole-batch precedence and the independent heartbeat exception, is owned by [docs/pi-supervision-branch.md](../pi-supervision-branch.md#components-and-their-owners). A no-change heartbeat outcome explicitly reported with `task=fleet` and `silent=true` is delivered silently with no rendered note, while every other routine outcome returns as an appended, rendered note that leads with ⛵ then the dim outcome text. A captain-facing outcome instead appears as one exact, sequence-keyed visible transcript entry, and then arrives in this conversation as one hidden supervision processing request listing each `[seq N] task: summary` it covers. That request is the one turn in which MAIN processes the outcome: give the captain a visible response where one is due, answer or escalate a decision, act on a blocker or failure, or record that no further action is needed, then call the `fm_branch_processed` tool with the highest sequence the request listed, exactly once. diff --git a/docs/tmux-backend.md b/docs/tmux-backend.md index 287fbcad4fa..fcc306f8450 100644 --- a/docs/tmux-backend.md +++ b/docs/tmux-backend.md @@ -48,7 +48,7 @@ Verify setup by spawning a small task and confirming its `fm-<id>` window appear A target-existence check proves only that the pane exists. The deeper tmux agent-liveness probe first verifies exact window membership, then reads process names to distinguish a running harness from a bare idle shell. -It classifies recognized Claude, Codex, Copilot, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, and Muse process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. +It classifies recognized Claude, Codex, Copilot, OpenCode, Pi, pi-signed, Grok, Kimi, Cursor, Muse, Rovo, and omp process identities as `alive`, common shells as `dead`, an authoritatively absent window as `missing`, unreadable state as `unreadable`, and every other process as `ambiguous`. Only `dead` and `missing` authorize recovery because a false dead result could launch a duplicate agent. For positive attribution, the probe combines two independent name sources rather than making either one load-bearing. @@ -60,6 +60,7 @@ Scoping the second source to the foreground process group rather than to the pan The same scoping covers multi-process launchers without a special case, so the Pi Launcher path is attributed through its `pi-signed` wrapper and `pi` engine even though its title is the exact foreground command `pi-launcher`. Direct executable identities `pi`, `pi-signed`, and `Pi` remain accepted exactly, and similar or prefixed process names are not accepted through those exact Pi-family entries. Muse is likewise anchored to the exact `muse` launcher identity or the installed `muse-bin-<version>` prefix, so unrelated names such as `musescore` and `amuse` remain ambiguous. +omp is anchored to the exact `omp` identity for the same reason, so `ompd` and `comp` remain ambiguous. Cursor is identified from its exact `cursor-agent` identity or versioned install tree in the foreground process path or structured argv[0]; a bare `node` or unrelated `agent` remains ambiguous. The CI-enforced portable regression and opt-in real-harness drift guard follow the split owned by `.agents/skills/firstmate-coding-guidelines/SKILL.md`. @@ -108,6 +109,7 @@ tests/fm-composer-ghost.test.sh tests/fm-kimi-harness.test.sh tests/fm-cursor-harness.test.sh tests/fm-muse-harness.test.sh +tests/fm-omp-harness.test.sh tests/fm-tmux-submit-busy.test.sh tests/fm-bootstrap.test.sh ``` diff --git a/docs/trace-context.md b/docs/trace-context.md index fcaefd52571..d493e05d9a1 100644 --- a/docs/trace-context.md +++ b/docs/trace-context.md @@ -23,7 +23,7 @@ When enabled, for each spawn Firstmate resolves one W3C `traceparent` carrier fo This feature parents no SDK span by itself. Because the injected carrier and the recorded carrier are the same string, an observer that reads the metadata reconstructs exactly the identity the child received. -The injection sits at the unconditional pre-launch export site, so it covers ship and scout spawns across `claude`, `codex`, `copilot`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, `gemini`, and `muse`, plus Secondmate spawns across that same set except the deliberately crewmate-only `gemini` and `muse` adapters. +The injection sits at the unconditional pre-launch export site, so it covers ship and scout spawns across `claude`, `codex`, `copilot`, `opencode`, `pi`, `pi-signed`, `grok`, `kimi`, `cursor`, `gemini`, `muse`, `rovo`, and `omp`, plus Secondmate spawns across that same set except the deliberately crewmate-only `gemini`, `muse`, and `rovo` adapters. This is the same coverage `GOTMPDIR` already has and requires no trace-specific `launch_template()` behavior. Ship and scout spawns reach that site on every spawn backend (`tmux`, `herdr`, `zellij`, `orca`, `cmux`); a Secondmate reaches it on every backend that accepts a Secondmate spawn (`tmux`, `herdr`, `zellij`), because `bin/fm-spawn.sh` rejects a Secondmate on `orca` and `cmux`. diff --git a/docs/turnend-guard.md b/docs/turnend-guard.md index 7ac8dab050b..a2f1564ac88 100644 --- a/docs/turnend-guard.md +++ b/docs/turnend-guard.md @@ -13,7 +13,7 @@ Do not infer this guard's scope, loop safety, or compatibility tradeoffs for tho `bin/fm-guard.sh` is a pull-based warning that runs only when another supervision command invokes it. The turn-end guard closes the remaining gap at the primary's own turn boundary. -When work, a process-event source, or Relay polling needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. +When work, a process-event source, a registered custom check, or Relay polling needs supervision at that boundary and no identity-matched watcher has a fresh beacon, the harness integration must either block the turn end or force one bounded follow-up that uses the recovery instruction from the emitted session-start protocol. The mid-turn pull warning uses the model-aware supervision verdict described below, while the turn-end guard keeps the PID-strict watcher predicate. Away mode is the one place the turn-end guard accepts a different supervisor: while `state/.afk` exists the away-mode daemon owns supervision, so a live identity-matched daemon with a fresh beacon satisfies that boundary in place of a watcher process holding the lock. The guard remains a backstop; [`watcher-continuity.md`](watcher-continuity.md) owns normal continuity. @@ -31,16 +31,17 @@ For an in-scope primary, the guard counts in-flight work from `state/*.meta`. Registered `state/procevent/*.source` records also require supervision even though they have no task metadata. The default cross-harness mode exits silently with no supervision need. Every mode treats `state/x-watch.check.sh` as supervision need, so Relay polling remains guarded without an in-flight task. +A custom check registered with `bin/fm-check-register.sh` counts the same way, so an operator's home-level poll keeps running after the last task is torn down. Otherwise it calls `fm_watcher_healthy <state-dir> <watch-path> [grace-seconds] [home]` from `bin/fm-wake-lib.sh`, the same PID-strict identity-matched lock and fresh-beacon check used by `bin/fm-watch-arm.sh`: a stale beacon blocks even when a watcher pid is live, and a fresh leftover beacon blocks when the lock is missing, dead, or identity-mismatched. The turn-end guard needs that strict check because it fires at the turn boundary, where the auto-arm is bringing a fresh watcher up for the upcoming idle period, and it cooperates with that arm rather than trusting a beacon left by the cycle that just ended. `bin/fm-guard.sh`, the pull warning, instead uses the model-aware `fm_watcher_supervision_verdict` from the same library, because it fires mid-turn when the auto-arm model runs no watcher at all. Under the Claude Stop auto-arm model a beacon fresh within grace is healthy even with no live watcher process, and only a beacon stale beyond grace (or absent) alarms. -Under the Pi extension model a live identity-matched watcher is the ordinary healthy state, but a genuinely unheld lock with a beacon fresh within grace is also healthy while a live Pi session provably owns continuity, because `.pi/extensions/fm-primary-pi-watch.ts` tears the watcher down on every actionable wake and spawns the replacement itself. +Under the extension model (Pi, pi-signed, and omp) a live identity-matched watcher is the ordinary healthy state, but a genuinely unheld lock with a beacon fresh within grace is also healthy while a live Pi or omp session provably owns continuity, because `.pi/extensions/fm-primary-pi-watch.ts` and `.omp/extensions/fm-primary-omp-watch.ts` tear the watcher down on every actionable wake and spawn the replacement themselves. A lock is genuinely unheld only when the lock directory or its symlinked owner directory is absent, or when the existing lock records no pid at all. Any lock with a recorded pid remains down when its pid, home, watcher path, or process identity fails the strict watcher health check. -That ownership proof is `fm_pi_extension_owns_supervision` in `bin/fm-wake-lib.sh`: both Pi primary extensions must be recorded in their state markers at their current on-disk builds by the process named in `state/.lock`, and that process must still be alive. +That ownership proof is `fm_extension_owns_supervision` in `bin/fm-wake-lib.sh`, which accepts either the Pi pair (`fm_pi_extension_owns_supervision`) or the omp pair (`fm_omp_extension_owns_supervision`): both primary extensions of one family must be recorded in their state markers at their current on-disk builds by the process named in `state/.lock`, and that process must still be alive; omp never inherits the Pi tolerance because its proof is keyed on its own two files and markers. Requiring the turn-end guard extension as well as the watch extension is deliberate, because a home without that structural backstop has no benign hand-off to tolerate. -Without that proof an unheld lock alarms exactly as it did before, so an unloaded, version-drifted, or exited Pi session is loud immediately, and a cycle the extension never restores is loud once the beacon passes grace. +Without that proof an unheld lock alarms exactly as it did before, so an unloaded, version-drifted, or exited Pi or omp session is loud immediately, and a cycle the extension never restores is loud once the beacon passes grace. Under every persistent-watcher harness a live identity-matched watcher with a fresh beacon is still required, so the pull guard keeps the same strict semantics there. Its banner names the true failing condition, either a missing live watcher process or a genuinely stale beacon with its real age, and keys the once-per-episode dedup on that condition rather than the beacon mtime. @@ -48,19 +49,31 @@ While `state/.afk` exists the away-mode daemon (`bin/fm-supervise-daemon.sh`) ow The turn-end guard therefore accepts `fm_afk_daemon_owns_supervision` from `bin/fm-wake-lib.sh` as proof of supervision on that path: away mode must be active, and this home's `state/.supervise-daemon.lock` must name a live pid whose current process identity still matches the identity the daemon recorded for itself. That is the same identity discipline the watcher lock uses, so a recycled pid, a lock left behind by a killed daemon, and a daemon that never recorded its identity all fail it. A daemon that cannot record its own identity at startup logs a warning and keeps running, because a supervisor must not refuse to run over an unreadable `ps`; that warning is what names the cause when the guard then keeps blocking away-mode turn boundaries for the rest of that daemon's life. -The proof covers ownership only, never freshness: the fresh-beacon half of the predicate is unchanged, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. +The proof covers ownership only, never freshness: the guard still requires a fresh beacon, so a daemon that stops restarting its watcher still blocks once the beacon passes grace, and a home with no daemon and no watcher blocks exactly as it did before. +That beacon check uses the poll-derived grace described below rather than the flat `FM_GUARD_GRACE` default, because the daemon starts a fresh one-shot watcher only after it finishes handling the previous wake, and that handling can legitimately outrun a fixed 300-second window under load (a slow registered check, a busy supervisor pane) with the daemon perfectly healthy throughout. With away mode off the daemon lock proves nothing and the strict watcher predicate is unchanged. `FM_STATE_OVERRIDE` wins over `FM_HOME/state`, and `FM_HOME` wins over repository-root `state/`. `FM_GUARD_GRACE` controls beacon freshness and defaults to 300 seconds. If `jq` is missing or hook stdin is empty, the guard exits 0 because it cannot safely read loop-guard fields. +### Guard grace and the poll cadence + +`bin/fm-watch.sh` touches `state/.last-watcher-beat` once per cycle, immediately before its terminal wait (`event_wait_or_sleep`) as well as at the top of the next cycle, so a healthy watcher's beacon can legitimately age up to `FM_POLL` seconds between touches. +A fixed 300-second grace default stops correctly bounding staleness once a home's `FM_POLL` reaches or exceeds it: a perfectly healthy watcher mid-wait would then read stale at the edge of every full poll cycle by definition, which is exactly what a long-poll home (`FM_POLL=300`) hit against the Claude Stop-hook auto-arm (`bin/fm-claude-stop-autoarm.sh`). +That hook and `bin/fm-watch.sh`'s own pre-acquisition staleness check (the "lock held by live pid but heartbeat is stale" refusal) both derive their default grace from the configured poll instead of a bare constant: `max(300, FM_POLL + 60)`, so the default never drops below the historical 300-second floor for the common short-poll case but grows with the poll cadence once that cadence would otherwise outrun it. +`fm_poll_derived_grace` in `bin/fm-wake-lib.sh` is the single owner of that formula. +The auto-arm hook additionally exports its resolved `FM_GUARD_GRACE` when it forks `bin/fm-watch-arm.sh`, so the arm wrapper and the watcher it may start judge staleness with the exact same value the hook just judged it with, whether that value came from an operator override or the poll-derived default. +`bin/fm-turnend-guard.sh`'s away-mode branch (`fm_afk_daemon_owns_supervision`, above) also derives its beacon grace from `fm_poll_derived_grace` rather than falling back to the bare 300-second default, for the same reason: the daemon's watcher-restart cadence there is not a fixed poll loop, so a flat grace misreads a daemon that is genuinely still cycling as down. +Every other direct `FM_GUARD_GRACE` reader (`bin/fm-guard.sh`, the strict-watcher checks in `bin/fm-turnend-guard.sh` and its harness-specific wrappers, `bin/fm-wake-lib.sh`) still falls back to the bare 300-second default unless `FM_GUARD_GRACE` is set explicitly in the environment. + ## Harness integrations - Claude registers two `Stop` hooks in `.claude/settings.json`, both anchored through `CLAUDE_PROJECT_DIR`: `bin/fm-turnend-guard.sh --claude`, and `bin/fm-claude-stop-autoarm.sh` with `asyncRewake: true` and `timeout: 28800`. - Codex registers a `Stop` hook in `.codex/hooks.json`, anchors the executable to the hook process working directory, verifies a Firstmate-shaped hook-bearing root, and passes the original payload to the shared guard. - OpenCode listens for `session.idle` in `.opencode/plugins/fm-primary-turnend-guard.js`, lets the watcher coordinator act first, and calls `client.session.promptAsync` once when the guard returns 2. - Pi listens for `agent_settled` in `.pi/extensions/fm-primary-turnend-guard.ts`, runs once per logical agent run, and calls `pi.sendUserMessage(..., { deliverAs: "followUp" })` once when the guard returns 2. +- omp answers its blocking `session_stop` hook in `.omp/extensions/fm-primary-turnend-guard.ts`, passing the payload's own `stop_hook_active` to the shared guard and returning `{ continue: true, additionalContext }` when the guard returns 2, so the continuation is compelled rather than requested; the continuation's stop carries `stop_hook_active: true`, which bounds it to one per turn, and omp's own cap of eight consecutive continuations is the second backstop. `session_stop` never fires for an interrupted turn or a task session, so those boundaries are deliberately unguarded. - GitHub Copilot CLI registers shell-completion `notification` and `agentStop` hooks in `.github/hooks/firstmate.json`. The primary runs one tracked asynchronous watcher shell task, the notification hook injects actionable completion context, and `bin/fm-ghcp-hook.sh primary-stop` delegates only the nonblocking missing-watcher backstop to `bin/fm-copilot-stop.sh`. - Cursor registers a `stop` hook in `.cursor/hooks.json` and delegates the whole turn boundary to `bin/fm-turnend-guard-cursor.sh`, the park described below. @@ -107,6 +120,7 @@ A Claude failure notice describes the automatic mechanism as broken and does not OpenCode, Pi, and pi-signed expose passive callbacks for this purpose. Their adapters fail open at the hook boundary to protect the user session but schedule one bounded follow-up when the predicate blocks. +omp is the exception among the Pi-derived harnesses: its `session_stop` hook blocks like Codex's `Stop` hook, so no passive latch is needed and the `stop_hook_active` loop guard applies unchanged. The generated prompts use the canonical `turn-end-guard` kind after the U+2063 `FIRSTMATE_OP: ` prefix, so Ahoy does not treat them as captain messages. Each passive adapter owns a loop latch. Pi keeps the latch across internal tool turns and clears it only when the generated follow-up settles or delivery fails. @@ -177,7 +191,7 @@ That warning uses `bin/fm-supervision-instructions.sh --repair-line`, so it alwa ## Regression coverage -`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. +`tests/fm-turnend-guard.test.sh` covers the predicate, main and secondmate primary scope, child-worktree exclusion, `FM_HOME` and `FM_STATE_OVERRIDE` precedence, the live-lock and fresh-beacon guard predicate, the cooperative `--claude` open-generation claim wait, monotonic failed-epoch progression, bounded attended fail-open, post-alarm continuation suppression, positive recovery reset, generation and legacy claim cases that must block or clear instead of allowing a blind stop, away-mode daemon ownership between watcher cycles and over a watcher lock left behind by an exited watcher, plus its dead, pid-reused, absent, stale-beacon, and away-mode-off negatives, the away-mode beacon's poll-derived grace widening for a live daemon still mid-cycle and its bound against a dead daemon, a beacon older than that wider grace, and FM_POLL's inapplicability with away mode off, Pi logical-run latching, missing-`jq` behavior, all five primary registrations, Grok native and legacy selection, typed field precedence, malformed input, and exactly-one-path safety. `tests/fm-guard-stale-banner.test.sh` covers the pull-guard predicate, including the persistent-model fresh-leftover-beacon negative control, the auto-arm model's healthy fresh-beacon-without-a-watcher case and stale-beacon alarm, and the extension model's live-watcher path, ownership-qualified fresh hand-off, held-lock failures, independently broken ownership signals, stale-beacon alarm, queued-wake warning, and Pi and pi-signed harness routing. It also covers true-reason banner wording and reason-keyed episode dedup surviving a beacon mtime change. `tests/fm-cursor-primary.test.sh` covers the Cursor park end to end over real processes with no harness installed: each tracked Claude-shaped entrypoint standing down on a Cursor payload, both follow-up sources, the bounded repair nag and its reset, the nested loop bounds, supersession, away-mode and lock-ownership inertness, Pi-host stand-down without Cursor identity and continued parking when `PI_CODING_AGENT` leaks alongside `CURSOR_AGENT` or `CURSOR_INVOKED_AS`, child-worktree exclusion, and that the adapter never exits 2. @@ -187,4 +201,5 @@ It also covers true-reason banner wording and reason-keyed episode dedup survivi `tests/fm-kimi-harness.test.sh` covers the separate Kimi crew hook's format preservation, idempotence, refusal cases, token guard, spawn registration, and teardown cleanup. `tests/fm-supervision-instructions.test.sh` covers recovery-line ownership and pi-signed's identity-preserving reuse of Pi's protocol. `FM_PI_LIVE_E2E=1 tests/fm-pi-primary-live-e2e.test.sh` is the opt-in isolated Pi path. +`tests/fm-omp-harness.test.sh` covers the omp extension pair over a fake omp API (forced continuation on exit 2, the `stop_hook_active` bound, the seatbelt block, the ownership proof), and `FM_OMP_LIVE_E2E=1 tests/fm-omp-primary-live-e2e.test.sh` is the opt-in isolated omp path. [`verification/supervision.md`](verification/supervision.md#turn-end-guard) records the active cross-harness empirical evidence, including the 2026-07-24 Claude `asyncRewake` revalidation. diff --git a/docs/verification/lint-option-a.md b/docs/verification/lint-option-a.md new file mode 100644 index 00000000000..ded3b0ac3eb --- /dev/null +++ b/docs/verification/lint-option-a.md @@ -0,0 +1,71 @@ +# Local ShellCheck option A measurement + +The 2026-09-05 lint-cost audit measured the seven roots from the missed-reply incident at commit `f09de8a3d3a550b13b4d535346fbc7b9ac0d6c19`: + +```text +bin/fm-brief.sh +bin/fm-parent-channel-lib.sh +bin/fm-pending-reply-lib.sh +bin/fm-secondmate-report.sh +tests/fm-brief.test.sh +tests/fm-classify-corr-token.test.sh +tests/fm-pending-reply.test.sh +``` + +ShellCheck was the repository-pinned 0.11.0 Darwin arm64 build. +The baseline was one source-aware invocation containing all seven roots. +Option A used one process per root, omitted `--external-sources`, retained extended dataflow, and applied the local cross-file exclusion list. +Both variants were measured in the same quiet-host window: + +| Variant | User + system CPU | Reduction | Worst-process RSS | Reduction | +| --- | ---: | ---: | ---: | ---: | +| source-aware baseline | 140.1 s | n/a | 8.30 GB | n/a | +| option A, per-root processes | 9.8 s | 93.0% | 0.56 GB | 93.3% | + +## Reproduction + +Check out the recorded commit, install the pinned binary with `bin/fm-install-shellcheck.sh`, put it first on `PATH`, and run the following on macOS. +No `--extended-analysis=false` flag is present, so dataflow remains on. +Diagnostics are discarded because only process cost is under measurement. + +```bash +set -eu +[ "$(bin/fm-lint.sh --required-version)" = "$(shellcheck --version | awk '/^version:/ {print $2; exit}')" ] +roots=( + bin/fm-brief.sh + bin/fm-parent-channel-lib.sh + bin/fm-pending-reply-lib.sh + bin/fm-secondmate-report.sh + tests/fm-brief.test.sh + tests/fm-classify-corr-token.test.sh + tests/fm-pending-reply.test.sh +) +rm -rf .lint-option-a-measurement +mkdir .lint-option-a-measurement +/usr/bin/time -lp -o .lint-option-a-measurement/baseline.time \ + shellcheck --norc --external-sources -- "${roots[@]}" >/dev/null || true +index=0 +for root in "${roots[@]}"; do + index=$((index + 1)) + /usr/bin/time -lp -o ".lint-option-a-measurement/option-a.$index.time" \ + shellcheck --norc --exclude=SC1091,SC2034,SC2153,SC2329 -- "$root" \ + >/dev/null || true +done +awk ' + /^user / {cpu += $2} + /^sys / {cpu += $2} + /maximum resident set size/ {if ($1 > rss) rss=$1} + /bytes allocated/ {allocated += $1} + END {printf "cpu_seconds=%.2f worst_rss_bytes=%.0f bytes_allocated=%.0f\n", cpu, rss, allocated} +' .lint-option-a-measurement/baseline.time +awk ' + /^user / {cpu += $2} + /^sys / {cpu += $2} + /maximum resident set size/ {if ($1 > rss) rss=$1} + /bytes allocated/ {allocated += $1} + END {printf "cpu_seconds=%.2f worst_rss_bytes=%.0f bytes_allocated=%.0f\n", cpu, rss, allocated} +' .lint-option-a-measurement/option-a.*.time +``` + +CPU and RSS vary with host load, so percentage claims must compare runs from one measurement window. +When results must be compared across windows, use the reported `bytes_allocated` totals as the stable work proxy rather than quoting a CPU or RSS ratio. diff --git a/docs/verification/muse.md b/docs/verification/muse.md index 2a2637b3c65..38d12653a6a 100644 --- a/docs/verification/muse.md +++ b/docs/verification/muse.md @@ -204,7 +204,7 @@ That is the same terminal shape the `echo`-provider interrupt produced, now conf ## Refreshing this record -Run both opt-in live guards after any muse upgrade, because the version-suffixed process name, session protocol, and styled composer are vendor-controlled surfaces: +Run both live guards after any muse upgrade, because the version-suffixed process name, session protocol, and styled composer are vendor-controlled surfaces: ``` FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh diff --git a/docs/verification/process-event-sources.md b/docs/verification/process-event-sources.md index f33ba4fea5e..c37a0d89cf8 100644 --- a/docs/verification/process-event-sources.md +++ b/docs/verification/process-event-sources.md @@ -97,10 +97,11 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | adapter-owned terminal verdict | two fixture adapters - one that ends on any result, one with no terminal knowledge - decide the outcome alone: the first has its registration and claim retired automatically after one capture and is never restarted, the second stays armed | | adapter-owned application of a captured result | a remote-secondmate reply captured through the real relay in an isolated home reaches that secondmate's local status mirror, settles its correlated pending-reply expectation, re-arms the next cursor-anchored source, and is acknowledged, with no handler step or duplicate `check` wake; its new mirrored bytes remain visible to the watcher's signal gate, while a cursor-loss whole-log recapture that adds no bytes is acknowledged quietly; for an already-escalated request, the same path closes the exact decision so the open-decision fold clears and remains clear; a capture whose adapter application fails because local storage for a referenced remote document is obstructed is left unacknowledged and receives the fallback `check` wake, and the handler's own `handle` still applies it in full after storage recovers | | generic built-in keyed-answer feed | `tests/fm-captain-hold-lifecycle.test.sh` drives a bound built-in source through the real runner with a fixture adapter that only prints keyed lines, proving any bound built-in channel reaches the one keyed-answer intake: named captain-held tasks close at capture time, a card-declared release mode frees held work, keys naming no captain-held task skip, freeform prose forges nothing, matching answer-and-mode replays are idempotent while mode mismatches refuse, an unbound source closes nothing, and capture remains independent of the handler wake. | +| structured reconcile feed | The same suite drives the optional `reconciles` adapter seam through the real runner and proves only a bound captured source can create a request; the ordinary keyed-answer and chat paths refuse the reserved value without closing or creating a request, versioned selection stays separate from its note, rollout-compatible ordinary legacy answers still pass, and legacy reconcile-shaped values feed neither intake. | | adapter-owned silence verdict | an armed Lavish source driven against a stand-in poll that returns an empty ended session captures its result, records it durably handled, appends no wake, and stays silent through a later `reconcile` that would otherwise republish it, while still retiring its ended source; the same real path with a `Send & End` response carrying the captain's choice still publishes its `check` wake and is left unacknowledged for the handler | | silence fails closed | the adapter's published `silent` command suppresses only an `ended` session with no queued content block, and announces a real answer, freeform prose, any recognized content block regardless of its declared count, a malformed top-level content header, a `waiting` or `missing` session, a server error, an unreadable result, and indented payload text imitating an empty content block; the `remote-reply` and `when` adapters, which implement no `silent` command, announce every result | | terminal retirement preserves the result | the retired source's captured output, its announced event, its handled acknowledgement, and later explicit `retire` all still behave normally | -| registration-generation retirement | an old terminal runner preserves a concurrently replaced registration and releases ownership so the replacement runs independently; injected registration-removal failure retains a terminal claim, performs no second poll, and completes idempotently once removal recovers | +| registration-generation retirement | an old terminal runner preserves a concurrently replaced registration and releases ownership so the replacement runs independently; injected registration-removal failure retains a terminal claim, performs no second poll, and completes idempotently once removal recovers; a live owner retiring its own terminal source mid-capture tolerates only its transient reservation-removal failure and still removes the registration under exact ownership | | one `Send & End`, one result | an armed Lavish source driven against a stand-in for the published poll, which delivers the final `session_ended` feedback once and empty ended sessions afterward, polls exactly once, captures exactly one result, publishes one distinct event, and retires itself | | bounded re-announcement until handled | a durably captured result with no handled acknowledgement is re-announced by `reconcile` with the same source and sequence on every call - not only the first restart after a crash - and a presented-but-unacknowledged wake resurfaces identically after a simulated replacement session | | handled acknowledgement | `fm-procevent.sh handled <source-id> <sequence>` atomically and idempotently records handling at mode `0600`, fails without leaving a marker when private-mode enforcement fails, reports the first call distinctly from every repeat, stops further re-announcement once recorded, and never authorizes a paired effect twice across repeat calls | @@ -112,9 +113,13 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | one owner per canonical source | a second home's `start` for the same source id reports `already owned` and publishes nothing | | canonical physical identity | a final-component symlink and its target produce the same Lavish source id | | isolated public start boundary | direct `start` establishes a new runner-led process group before claiming the source, so retirement cannot signal an unrelated process inherited from the caller's group | -| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, and cross-home replacement removes the old generation's staging file from its recorded state directory | -| crashed leader with a live owned group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile then stops that surviving group before any replacement starts, never leaves two source processes running for one canonical source, and a generation with no leader and no surviving group is still reclaimed | -| PID-reuse safety | retirement refuses to signal a live PID whose identity differs from the claim, and a reused PID never reaches the group-stop path because its leader is alive | +| guarded runner startup | the source command does not launch when the detached owner guard rejects an invalid lease configuration, proving the runner waits for positive guard readiness and fails closed when initialization fails | +| attached owner continuity | a foreground `start` with a one-second lease remains alive beyond that lease while its caller stays attached, then captures normally when the blocking source completes | +| owner-home lifetime and scope | a detached runner and its spawning descendant are observed reparented before an expired owner lease stops their whole process group and process churn; replacing the state directory at the same path cannot keep the old runner alive with a new lease because its recorded device/inode no longer matches, while an identical runner in an unchanged home whose reconcile cycle keeps its lease fresh remains alive | +| launch pacing during owner-loss grace | an immediately returning source that attempts detached self-relaunches is held to the configured minimum interval between command launches and remains bounded until its expired owner lease stops the generation; replacement starts a fresh pacing generation, prunes prior pacing state, and prevents a superseded sleeping runner from recreating it | +| stale reclaim without displacement | concurrent contenders replacing one stale claim start exactly one runner, cross-home replacement removes the old generation's staging file from its recorded state directory, and a generation whose stale owner and independently empty process group prove it gone remains reclaimable when its recorded state-root identity can no longer be revalidated | +| crashed leader with a live group | `SIGKILL` on only the runner leader leaves its blocking child group alive; reconcile treats that leaderless group as ambiguous, preserves its claim without starting a replacement, and still reclaims a generation with no leader and no surviving group | +| PID-reuse safety | retirement refuses a live PID whose identity differs from the claim before signalling, and a surviving process group prevents stale-generation cleanup on both ordinary and failed reservation-removal paths | | coherent ownership reads | a claim replacement held inside the source boundary blocks `list` until one complete generation is visible | | retire-start exclusion | a queued start revalidates registration after the serialized retirement boundary and executes no child | | uncertain identity | a live owner whose identity probe transiently fails is not signaled or released, and its registration remains for retry | @@ -134,7 +139,7 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | inertness | a home with no registered source generates no state, starts no process, and does not need supervision | | absent extension registry parity | `tests/fm-extension-binding.test.sh` drives `list` and `verify` in a fresh home while the current directory contains project files and Pi packages and an environment variable names fake package data; both commands report no bindings, create no home path, and discover nothing outside `config/extensions.d` | | complete package and binding identity | the same suite drives the public bind and verify commands through manifest duplicate/unknown/version failures, project and task-copy confinement, canonical path and symlink rejection, hard-link rejection, owner/mode checks, a non-executable entrypoint, binding mode drift, complete-tree mutation, exact executable mutation, and a missing executable; the foreign-owner fixture executes when the platform permits constructing another uid and otherwise reports that privilege limitation, while ordinary non-privileged CI does not exercise it or claim it ran | -| external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves claim release and dead-owner reconciliation remove pending or consumed capture reservations only from the recorded revalidated state root; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | +| external evidence write confinement | the same suite substitutes `state/procevent/` and `state/procevent-inbox/` with post-registration symlinks and proves an external start fails before bytes reach either outside target; it proves public lifecycle entry, environment, paths, and descriptors cannot forge capture authority; it proves live-generation claim release removes pending or consumed capture reservations only from the recorded revalidated state root, while a generation independently proved gone may leave an unreachable token-keyed reservation rather than wedging ownership; and it proves the absent-registry built-in capture path retains its legacy state-path behavior | | strict handshake and negotiation | manifests offering versions 2 and 1 select host protocol 1 and `process-event-adapter/1`, unknown-only versions refuse, and wrong request ids, unknown or duplicate fields, malformed JSON, and nonzero handshake exits publish no binding | | strict invocation envelope | malformed UTF-8, a byte-order mark, unescaped controls, malformed or multiple JSON documents, duplicate or unknown fields, oversized stdout, oversized stderr, wrong request ids, crashes, nonzero exits, a successful parent that leaves a foreground descendant in its host-created invocation group, and authority-shaped result fields are rejected; leaked group members are reaped and package diagnostic text is not copied into the bounded host-produced error evidence | | extension timeout and process-group cleanup | a bound adapter that ignores `TERM`, spawns a foreground descendant that ignores `TERM`, and exceeds its invocation timeout returns deterministic timeout evidence only after its exact invocation group is gone; deliberate process-group escape is outside this trusted-same-user protocol guarantee | @@ -144,13 +149,14 @@ Exercised by `tests/fm-procevent.test.sh` against a fake blocking source whose c | owner-matched replacement safety | two registrations for the same external source receive distinct owner tokens; unconditional external retirement and the first token cannot retire the replacement, the replacement token can, bounded home sweep derives and uses that exact token, and legacy built-in registrations retain unconditional behavior plus exact `--if-matches` retirement | | independent homes | two homes bind the same package id/version to different content-addressed absolute paths and independently capture results and extension state, with no cross-home fallback or result path | -Run the focused external-binding evidence with: +Run the focused external-binding evidence and the live Bearings session guard with: ```sh node --version bin/fm-test-run.sh tests/fm-extension-binding.test.sh FM_EXTENSION_BINDING_SEGMENT=lifecycle-invocation-cleanup bin/fm-test-run.sh tests/fm-extension-binding.test.sh bin/fm-test-run.sh tests/fm-procevent.test.sh +FM_BEARINGS_LAVISH_LIVE=1 bin/fm-test-run.sh tests/fm-bearings-board-lavish-live-e2e.test.sh bin/fm-doc-audience-check.sh ``` @@ -171,19 +177,29 @@ The 2026-08-27 review inspected `bin/fm-harness.sh`, `bin/fm-supervision-instruc ## Runner lifetime and cleanup A runner started by `reconcile` is its own process group leader and is reparented to init, so it outlives the shell that started it by design. -That means nothing about the starting context can reap it: removing a home's state directory does not stop an already-running child, and signalling only the runner leaves the blocking child alive. +Removing a home's state directory does not stop an already-running child, and signalling only the runner leaves the blocking child alive. -Two paths therefore stop a runner, and both verify the runner-owned process group, escalate to `KILL` while that group still exists, and refuse to release ownership until the whole group is gone: +Three paths stop a runner generation through its verified process group: +- The runner starts only after its separate owner guard confirms initialization; the guard stops the runner group after two consecutive checks cannot prove the owning home's recorded physical identity and lease freshness. - `retire` resolves the runner PID and identity from this home's machine-wide claim, so retirement still works when the home's state is already gone. - `reconcile` stops a runner this home owns whose source registration has been removed, and reports it as `stopped=N`. -The same group rule decides when a claim may be reclaimed, not only when a runner may be signalled. -A leader that died while its owned group kept running is not a stale generation, so `reconcile` stops that surviving group and releases its generation before starting any replacement, and preserves the claim for a later retry when it cannot prove the group stopped or another home owns it. -Signalling that group is safe precisely because only an absent leader reaches this state: a reused PID leaves the leader alive, which the identity comparison classifies as stale or uncertain, and no group signal follows. +The owner guard and explicit cleanup paths reach the blocking source and its descendants through the runner's group. +The registration launch floor independently bounds repeated runner launches while an owner-loss lease is still valid. +The Lavish adapter's start-to-start poll governor separately bounds its internal retry loop under shipped defaults without delaying a normally blocking poll. +An attached public `start` maintains the lease for its caller's lifetime. +At the accepted confused-agent/accidental grade, the inherited `FM_PROCEVENT_IN_RUNNER` marker prevents detached runners and their ordinary children from refreshing it; adversarial unforgeability against a source that deliberately strips that marker is out of scope. -This was found by four orphaned runners, elapsed 6-13 minutes, left by a suite whose fixture source never completed. -`tests/fm-procevent.test.sh` now covers both paths, and three consecutive suite runs leave zero runners, zero fixture children, and zero stray claims. +The same group rule decides when a claim may be reclaimed, not only when a runner may be signalled. +A leader that died while its process group kept running is not a gone generation. +Because the leaderless group cannot be proved to belong to the recorded generation, `reconcile` preserves its claim without signalling it or starting a replacement. +Once a stale owner and an independent group check prove the whole generation gone, an unreachable token-keyed capture reservation cannot veto reclamation. +Known limit: when either a live reused PID or an absent leader makes group ownership ambiguous, the reaper does not act because it cannot prove the group is the orphan generation; storm-rate containment plus ordinary lease and reconcile cleanup are the confused-agent-grade backstop. +Known limit: identity and process-group verification cannot be made atomic with signalling in portable shell. +The reaper signals only a target it has verified as the orphan generation, but PID and group reuse remain possible in the narrow interval between verification and the signal; launch pacing is the primary host-wedge protection and watchdog cleanup is a backstop. + +`tests/fm-procevent.test.sh` covers owner-loss reaping, descendant churn cessation, cross-home scope, launch pacing, guard startup failure, attached-start continuity, explicit retirement, and stale-group reconciliation. ## Portability finding diff --git a/docs/verification/rovo.md b/docs/verification/rovo.md new file mode 100644 index 00000000000..8a588e3ad8f --- /dev/null +++ b/docs/verification/rovo.md @@ -0,0 +1,237 @@ +# Verification: the rovo (Atlassian Rovo CLI) crewmate/scout adapter + +Active empirical evidence for firstmate's rovo adapter. +The skill tree rooted at [`.agents/skills/harness-adapters/SKILL.md`](../../.agents/skills/harness-adapters/references/harness/rovo.md) owns the operating facts; this record owns how they were established and what is still unproven. + +## Subject + +| Field | Value | +|---|---| +| Version | `Rovo CLI: 202609.1.2` | +| Verified | 2026-09-02 (herdr backend liveness added 2026-09-03) | +| Binary | `~/.local/bin/rovo`, a bash wrapper that execs a PyArmor-obfuscated PyInstaller bundle under `~/.local/share/rovo/active/` | +| Platform | macOS arm64 (Darwin 25.6.0) | + +An earlier scout task (`fm-rovo-smoke-s1`) established the baseline empirical facts through a hand-written PTY VT emulator, no adapter code, and no dispatchable wiring. +This task landed the executable owners against those facts and re-verified the load-bearing ones live, including two facts the scout could not test (auth refresh under an expired access token, and a mid-tool-call Escape). +Every command below ran unsandboxed, because rovo's OAuth credentials live in the macOS keychain, which a sandboxed shell cannot read. + +## Detection + +``` +$ rovo --version +Rovo CLI: 202609.1.2 +``` + +`bin/fm-harness.sh` tests `ATLASSIAN_AGENT_TYPE=rovo` and `ROVODEV_CLI=1` before the `CLAUDECODE` line and matches ancestry `comm=rovo` otherwise; `tests/fm-rovo-harness.test.sh` pins both the marker-precedence order (a rovo marker outranks an inherited `CLAUDECODE`) and the markerless-ancestry fallback with faked `ps` output. + +## Launch: bare launch-then-send, the kimi shape + +`fm-spawn.sh` builds `env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS <rovo-bin> run --yolo <model/effort flags>` - BARE, with no positional brief - wrapped by the shared `env -u CURSOR_AGENT -u CURSOR_INVOKED_AS` prefix every non-cursor harness gets. +The brief is then typed in after the TUI comes up, the same launch-then-send shape kimi uses, through the same shared readers (`fm_backend_capture`, `fm_backend_composer_state`, `fm_backend_send_text_submit`): + +1. `rovo_wait_for_ready` polls for the `Welcome to Rovo!` banner (primary) or a composer-empty verdict (weaker fallback, see the composer-ghost-text gap below). +2. The pointer `Read the brief at <absolute-path> and follow it exactly.` is submitted via `fm_backend_send_text_submit`. +3. `rovo_wait_for_delivery` confirms composer-empty AND either the echoed `Read the brief at` text or a nonzero `Context:` percentage (`context:[^%]*[1-9][^%]*%`, tolerant of the footer's bar glyph but anchored before the `%` so the `.../922K` denominator cannot false-positive). + +Each gate fails the spawn loudly (a `failed:` line in the task status file) if it never resolves, so a never-ready or silently-dropped delivery is a visible spawn failure rather than a half-wired pane. + +### Why not a positional brief + +A positional brief is dead-on-arrival. `rovo run --yolo "<brief>"` loads a spinner, never enters a working state, prints no reply, and drops back to a bare idle shell prompt within about 10-15 seconds. This was reproduced independently four times over a raw PTY (varying `TERM`, window size, workspace, and 60-150s windows) and once more under real tmux 3.6a driven with the exact `fm-spawn.sh` send-keys shape (new window, `send-keys -l` the full launch line, then `Enter`). `--startup-receipt` cannot rescue that shape either - it is rejected before start alongside any message: + +``` +$ rovo run --startup-receipt receipt.json --yolo "Reply with PONG" +Invalid value: --startup-receipt requires prompt-free interactive mode in a terminal +``` + +so it can only gate a bare (no-message) launch, and this adapter always delivers a message, so it is not used. + +### The launch-then-send shape, confirmed live end to end + +Bare `rovo run --yolo`, driven over a raw PTY, was confirmed to: render the `Welcome to Rovo!` readiness banner; accept the typed pointer and act on it (a brief instructing a real `sleep 25` bash tool call drove the `Rovo is thinking` busy line); accept a mid-tool-call Escape that printed `Agent cancelled` (see the interrupt section); and exit cleanly on `/exit` with the `Run rovo --restore <id> to resume your conversation` hint. `tests/fm-rovo-signals-live-e2e.test.sh` is that end-to-end guard. + +`tests/fm-rovo-harness.test.sh` pins the portable half against a stateful fake `tmux` and a fake `rovo` binary (no real network or credentials): the launch command is bare (`run --yolo`, no positional brief, no `--startup-receipt`); the pointer typed after readiness is exactly `Read the brief at <absolute-path> and follow it exactly.`; delivery confirms via the context-percentage or echoed-pointer signal; a never-ready fake screen fails the spawn loudly; a dropped-submit fake screen fails the spawn loudly; and the model/effort flags, marker-clearing, missing-binary refusal before any pane exists, and crew/scout-only secondmate refusal all hold. + +## Busy state: the "Rovo is thinking" fallback + +Live, over a raw PTY, submitting a prompt that runs a real `sleep 25` bash tool call rendered the busy line and footer: + +``` +⬢ Rovo is thinking... +Enter to queue, Ctrl+Enter to steer +``` + +`fm_busy_rovo_tail_busy` (`bin/fm-busy-lib.sh`) matches that exact rendered text; `fm_busy_classify` was confirmed live-and-portably to read it as `busy rovo-regex`, and an idle footer with no busy line as `idle rovo-regex`. +This is a rendered-tail fallback exactly like Grok's, not a semantic source: rovo's `eventHooks` (`~/.rovo/config.yml`) fire at tool granularity (`on_tool_start`/`on_tool_end`) only, never at turn-end, so no writer is armed and none is seeded. +Grok was previously the only rendered-text arm the redesigned busy contract allowed; this task extends that same documented exception to rovo, scoped to `harness=rovo` exactly like Grok is scoped to `harness=grok`, and neither can classify the other (`tests/fm-rovo-harness.test.sh`'s isolation case). + +## Composer ghost text: measured, deliberately left unfixed + +A live idle-composer capture over a raw PTY located the inline placeholder chip inside the actual bordered content row, not merely in a suggestion list below it: + +``` +row 10 ╭──────────────────────────────────────╮ +row 11 │ Summarize my open tasks │ fg 38;2;162;163;165 (luminance ~163) +row 12 ╰──────────────────────────────────────╯ +``` + +Real typed text in the same row, captured separately, renders at `38;2;206;207;210` (luminance ~207). +Both values sit above `bin/fm-composer-lib.sh`'s default `FM_COMPOSER_GHOST_LUMA_MAX` of 128, so `fm_composer_strip_ghost` leaves the placeholder unstripped and a fresh rovo composer can misclassify as `pending` rather than `empty`. +Raising the shared default was considered and rejected: muse's own real, must-not-be-stripped prompt glyph measures luminance ~149.9 (`muse.md`), below rovo's ghost luminance of ~163, so no single global threshold can keep muse's glyph real while dropping rovo's ghost chip. +This is recorded as a known gap rather than patched, because the safe fix needs a harness-scoped signal the shared composer classifier does not carry today, and a threshold change risks regressing muse's already-credentialed behavior for a rovo-scoped fix. +The blast radius is bounded to composer-emptiness consumers such as steering delivery, which already retries through the doorbell ladder on a non-`empty` read. +It does not block readiness: readiness leads with the `Welcome to Rovo!` banner, so the ghost chip is never the deciding signal there. Delivery, however, requires composer-empty as one conjunct (alongside the echoed pointer or a nonzero `Context:` percentage), and on the herdr backend this conjunct may fail to settle within its poll window (the composer read non-empty even mid-turn in the live herdr run below), so `rovo_wait_for_delivery` can fail the gate and tear the pane down there. tmux delivery is separately verified working (see the tmux backend-liveness section below). This is a known limitation whose fix is tracked as a separate follow-up, not fixed in this change. + +## Interrupt: confirmed under real tmux + +The `fm-rovo-smoke-s1` scout report recorded a single Escape printing `Agent cancelled` during a running tool call, using its own hand-rolled PTY VT emulator. +A follow-up live check under real tmux 3.6a reproduced the scout's exact finding: a single Escape sent during a genuine mid-flight bash tool call printed `Agent cancelled` in the captured pane, in an isolated `tmux -L <private-socket>` session/window, not the shared fleet session. +The launch-then-send live guard (`tests/fm-rovo-signals-live-e2e.test.sh`) also reproduces it over a raw PTY. A single fixed-timer Escape had landed unreliably there - the exact instant the interrupt is delivered is timing-sensitive over a bare PTY, so one fixed Escape can fall between states - so the guard now sends Escape across the live `sleep 25` tool-call window until the cancel renders. That reproduced `Agent cancelled` on every run (it consistently landed within the first few attempts, ~8s into the tool call); a session that never rendered the cancel would exhaust every attempt and fail. +The session was never wedged: `/exit` still exited cleanly with the `Run rovo --restore <id> to resume your conversation` hint immediately after the Escape. +`bin/fm-control-lib.sh` records rovo's `fm_control_interrupt_ack_source` as `none`, the same conservative choice already made for claude, codex, grok, kimi, and cursor - a control-plane fact independent of whether the render happens to appear, because a rendered acknowledgement is not something the control plane depends on for any of those adapters. +Escape is the recorded interrupt key, and its rendered evidence is now corroborated both under real tmux and over a raw PTY rather than in tension with the code. + +## OAuth token lifetime and silent refresh + +The captain corrected this task's initial brief, which had treated the ~1h access-token lifetime as a hard mid-task blocker; this task's own live evidence confirms the corrected model. + +``` +$ rovo auth status +authenticated — Access token expired (2026-09-02 13:39:55 UTC), but a refresh +token is present. + +$ rovo run --yolo --output-file out.json "Reply with exactly the single word PONG and nothing else." +Run rovo --restore <session-id> to resume your conversation +$ cat out.json +PONG + +$ rovo auth status +authenticated — Access token valid, expires in 3574s (2026-09-02 15:15:40 UTC). +``` + +No browser prompt, no interactive step, and no visible interruption occurred between the first and second `rovo auth status` calls; the run in between silently refreshed the access token from the stored refresh token. +Treat the ~1h access-token lifetime as an ordinary operational fact rather than a non-negotiable-safety blocker: `rovo auth login` (interactive browser OAuth) is needed only after roughly four weeks of disuse or an invalidated refresh token, not mid-task. + +## Effort and model + +`agent.efficiencyLevel` accepts `low|medium|high|max` live via `--config-override`; a requested `xhigh` (unsupported) is recorded in task metadata but omitted from that JSON object, both verified against the fake-binary suite. +`--config-override` is single-value - a second occurrence silently discards the first rather than merging, confirmed live by reversing the order of two `--config-override` flags and observing the earlier one's effect disappear - so `fm-spawn.sh`'s `rovo_config_override_flag` folds `agent.efficiencyLevel` into the SAME JSON object as the mandatory `allowedExternalPaths` grant below rather than emitting two flags; the fake-binary suite pins that exactly one `--config-override` occurrence carries both. +Model discovery is per-account (`/models` or ACP `session/new`); the observed live list is recorded in `references/harness/rovo.md` and must never be hardcoded. + +## Worktree confinement and the allowedExternalPaths fix + +The standard crewmate flow needs a rovo worker to read its own brief and steering messages, and to write its status and report - all of which live in the firstmate home, outside the task's git worktree. +By default rovo confines every file-tool operation (`open_files`, `create_file`, `grep`, `expand_folder`, ...) to the workspace it was launched in, and its bash tool independently refuses the same external paths regardless of any grant. +Confirmed live with a plain `rovo run --yolo` launched inside an isolated scratch workspace, against an unrelated file in a separate outside directory: + +``` +$ cat "$LAB/outside/secret.txt" +OUTSIDE_SECRET_TOKEN_12345 +$ rovo run --yolo "Use your file-opening tool (not bash) to open and read the file $LAB/outside/secret.txt, then report its exact contents." --output-file out.txt +$ cat out.txt +Sorry, I can't access or read files outside the current workspace, including that temporary-system path. If you copy the file into the workspace or paste its contents here, I can help inspect it. + +$ rovo run --yolo "Run this exact bash command and nothing else: cat $LAB/outside/secret.txt" --output-file out.txt +$ cat out.txt +Captain, I can't run that command because it attempts to read a file outside the workspace, which I'm not permitted to access. Would you like to provide the file's contents here instead? +``` + +`toolPermissions.allowedExternalPaths` (`~/.rovo/config.yml`, default `[]`) is the only lift, and it must be granted at launch through `--config-override`: there is no live escalation once the process is already running. Confirmed live with the grant, same file, same file tool: + +``` +$ rovo run --yolo --config-override '{"toolPermissions":{"allowedExternalPaths":["'"$LAB"'/outside"]}}' \ + "Use your file-opening tool (not bash) to open and read the file $LAB/outside/secret.txt, then report its exact contents." --output-file out.txt +$ cat out.txt +`OUTSIDE_SECRET_TOKEN_12345` +``` + +The grant lifts the file tools only. rovo's bash tool stays confined to the worktree regardless, confirmed live with the identical grant still active: + +``` +$ rovo run --yolo --config-override '{"toolPermissions":{"allowedExternalPaths":["'"$LAB"'/outside"]}}' \ + "Run this exact bash command: echo hello >> $LAB/outside/status.txt" --output-file out.txt +$ cat out.txt +I can't run that command because it modifies a file outside the workspace. If you provide a workspace-relative path, I can run the equivalent command there - would you like to do that? +``` + +This matters because the standard crewmate contract's literal status line is a bash `echo ... >> status_file` command. Given that exact literal instruction, rovo recovered on its own by falling back to its native file tool for the same append, and succeeded, preserving the file's existing content: + +``` +$ echo "existing: prior" > "$LAB/outside/status2.txt" +$ rovo run --yolo --config-override '{"toolPermissions":{"allowedExternalPaths":["'"$LAB"'/outside"]}}' \ + 'Report status by appending one line: echo "working: test line" >> '"$LAB"'/outside/status2.txt' --output-file out.txt +$ cat out.txt +Appended `working: test line` to `status2.txt`. What would you like to do next? +$ cat "$LAB/outside/status2.txt" +existing: prior +working: test line +``` + +The same grant, at directory granularity, also covers listing a directory, reading a file inside it, and moving (not copying) it into a `handled/` subdirectory - the exact shape the steering-inbox acknowledgement contract (`bin/fm-task-inbox-lib.sh`) needs - confirmed live in one pass against a pre-existing `inbox/handled/` directory: + +``` +$ rovo run --yolo --config-override '{"toolPermissions":{"allowedExternalPaths":["'"$LAB"'/outside"]}}' \ + "List the directory $LAB/outside/inbox for *.msg files, read 001.msg, then move it into $LAB/outside/inbox/handled/001.msg (a rename/move, not a copy-and-delete you narrate but don't do)." --output-file out.txt +$ ls "$LAB/outside/inbox/handled" +001.msg +``` + +`fm-spawn.sh`'s `rovo_config_override_flag` builds one merged JSON object per rovo launch (see "Effort and model" above for why it must be one), always granting `toolPermissions.allowedExternalPaths` for exactly three real (symlink-resolved) paths scoped to this task: the brief directory (`data/<id>/`, covering `brief.md`/`launch-brief.md`/`report.md`), the steering inbox directory (`state/<id>.inbox/`, covering every steer and its `handled/` acknowledgement), and the status file itself (`state/<id>.status`). +`tests/fm-rovo-signals-live-e2e.test.sh` extends the launch-then-send live guard with exactly this shape end to end: a real rovo process launched with the production `--config-override` grant reads an external brief and appends to an external status file (preserving its prior content), and the same brief and status file, launched WITHOUT the grant, are left untouched while the transcript shows rovo's own refusal - proving the fix closes the gap rather than merely adding an untested flag. +`tests/fm-rovo-harness.test.sh` pins the portable half against the fake-binary suite: the grant's three paths appear in every rovo launch (including when the requested effort is unsupported and omitted), and exactly one `--config-override` occurrence ever appears. + +## Backend liveness: tmux verified live, herdr placement verified live with a herdr-side agent-detection gap + +tmux 3.6a is now installed and was exercised live in an isolated `tmux -L <private-socket>` session, so tmux pane liveness is fully verified rather than pending. +`bin/backends/tmux.sh`'s `fm_backend_tmux_classify_process_name` matches `*rovo*` alongside the other globbed harness names, so a rovo pane classifies `agent` (not `other`). +The two independent name sources behaved as designed: `#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class [`runtime-backends.md`](runtime-backends.md) already documents for codex/kimi's own patch-release name drift - while the foreground ps-based `comm` correctly reported `rovo`, and `fm_backend_tmux_agent_state` correctly returned `alive` through that primary source. The two-independent-name-sources design is exactly why the truncation quirk does not break the verdict. +`tmux capture-pane` correctly rendered the box composer and the `Rovo is thinking...` busy line while a real `sleep`-based bash tool call ran; `fm_busy_rovo_tail_busy` classified it busy, then idle once the tool call completed and the reply landed. The Escape/`Agent cancelled` evidence in the interrupt section above was captured in this same live tmux session. +`/exit` closed the tmux window cleanly, and `fm_backend_tmux_agent_state` reported `missing` immediately afterward - a clean, unambiguous exit verdict. + +A full `fm-spawn.sh --backend herdr` placement still cannot be driven to completion from this host: this task's own agent process runs inside the shared production Herdr session, so `fm-spawn.sh`'s cross-session launcher-identity guard correctly refuses to place a worker pane from that ambient parent identity into any other session, and forcing placement into the shared `default` session was rejected as unacceptable interference with the live, human-observed fleet. +That guard scopes fm-spawn.sh's own task/worktree orchestration, not the lower-level primitives it calls, so this task instead drove those same primitives directly against an isolated non-`default` session created by `bin/fm-herdr-lab.sh` (fleet-state tripwire confirmed the live `default` session was unchanged before and after): `herdr workspace create`/`pane list` for placement, the `rovo_capture`/`rovo_wait_for_ready`/`rovo_delivery_is_confirmed`/`rovo_wait_for_delivery` gate functions extracted verbatim from `fm-spawn.sh`, and `fm_backend_capture`/`fm_backend_send_text_submit`/`fm_backend_agent_state` from `bin/fm-backend.sh` and `bin/backends/herdr.sh` directly. + +``` +$ herdr workspace create --label rovo-verify --cwd ~/.fm-herdr-rovo-verify-scratch --no-focus --session fm-lab-... +{"result":{"root_pane":{"pane_id":"w1:p1",...},"workspace":{"workspace_id":"w1",...},...}} +$ herdr pane list --workspace w1 --session fm-lab-... +{"result":{"panes":[{"pane_id":"w1:p1","cwd":"/Users/.../.fm-herdr-rovo-verify-scratch",...}]}} +``` + +A rovo pane was placed in that isolated workspace, launched bare (the same `env -u ... rovo run --yolo` template documented above), and `rovo_wait_for_ready` returned success on the `Welcome to Rovo!` banner. +The typed pointer (`Read the brief at <path> and follow it exactly.`) was echoed into the pane, and rovo read a trivial no-op brief, ran a real `sleep 15` bash tool call, and replied `PONG` - the same launch-then-send shape already verified over tmux and a raw PTY, now also confirmed live over Herdr. +`fm_backend_herdr_capture` correctly rendered the `Rovo is thinking...` busy line during the tool call (`fm_busy_rovo_tail_busy` matches that captured text), and the pane read idle with `PONG` visible once the tool call finished; `rovo_wait_for_delivery`'s own composer-empty conjunct did not settle within its poll window, consistent with the already-documented composer-ghost-text gap below rather than a new defect. + +`fm_backend_agent_state` is the one signal this run disproves rather than confirms: it reported `dead` throughout - at the ready banner, mid-tool-call busy, and idle-with-`PONG` alike - even though rovo was demonstrably alive and responding the whole time. +The cause is on Herdr's side, not firstmate's: `fm_backend_herdr_pane_agent_state` calls `herdr agent get <pane>`, which returned `{"error":{"code":"agent_not_found","message":"agent target w1:p1 not found"}}` for the live rovo pane, because `herdr integration status` lists no `rovo` entry at all (only `pi`, `omp`, `claude`, `codex`, `copilot`, `devin`, `droid`, `kimi`, `opencode`, `kilo`, `hermes`, `qodercli`, `qwen`, `cursor`, `mastracode`, `antigravity-cli`, and `grok` are known integrations on the installed Herdr build). +Herdr has not shipped agent detection for rovo, so the classifier that recovery logic depends on (`fm_backend_agent_state`'s `alive`/`dead` distinction, and `fm_backend_herdr_tab_is_husk`'s reuse of it) cannot currently tell a live rovo pane apart from an empty one on the herdr backend; a live rovo worker placed on `backend=herdr` risks being misclassified as an agent-less husk by any recovery path that trusts this classifier. +This is recorded as a known Herdr-side integration gap rather than a firstmate bug, and is deliberately left unpatched here: no herdr-scoped workaround is safe to add without risking a false-positive `alive` verdict for some unrelated idle shell, so `backend=herdr` remains usable for launching a rovo crewmate/scout but unverified for automatic dead/husk recovery until Herdr ships rovo detection (or `bin/backends/herdr.sh` gains an independent process-based fallback the way `bin/backends/tmux.sh` already has). +`/exit` returned the pane to an idle shell prompt rather than closing it, unlike tmux which closes the whole window, so `fm_backend_agent_state` reading `dead` after exit is the textually correct verdict for an agent-less-but-present pane; it is only the ready/busy/idle misclassification while rovo was actually running that is the real finding above. + +## Skill-loading interop gap (documented, not fixed) + +``` +⚠ Invalid skill definition in .../.agents/skills/bootstrap-diagnostics/SKILL.md: 'metadata -> internal': Input should be a valid string +``` + +rovo's skill loader rejects every firstmate skill because `metadata.internal` is a boolean in firstmate's frontmatter and rovo's schema wants a string. +This blocks `/no-mistakes` and every other firstmate skill invocation inside a rovo worker until firstmate's `SKILL.md` frontmatter is made rovo-compatible, a separate deferred follow-up that touches every skill file and the installer contract (`.agents/skills/firstmate-coding-guidelines/SKILL.md`). +A `no-mistakes`-mode rovo ship crewmate is blocked by this gap; a rovo scout, which invokes no skill, is unaffected. + +## quota-axi provider mapping: not established + +`bin/fm-quota-choose.sh`'s `provider_for_harness` has no `rovo` entry. +rovo routes to several distinct underlying model families (OpenAI, Anthropic, Gemini) through Atlassian's own account, and this task found no live evidence of how, or whether, `quota-axi` models that relationship. +Rather than guess a provider family and risk a wrong quota verdict, `rovo` stays absent from that mapping, so a `rovo` candidate in a quota-balanced dispatch array fails closed with `unknown harness: rovo` instead of being silently misjudged; establishing the real mapping is follow-up work, not part of this adapter. + +## Refreshing this record + +Run the portable suite and the live guard after any rovo upgrade, because the process name, marker set, and rendered busy/interrupt text are all vendor-controlled surfaces: + +``` +bin/fm-test-run.sh tests/fm-rovo-harness.test.sh +FM_ROVO_SIGNALS_LIVE=1 bin/fm-test-run.sh tests/fm-rovo-signals-live-e2e.test.sh +``` + +The live guard requires a real, authenticated `rovo` binary but drives it through a raw PTY rather than tmux, so it runs on hosts without tmux installed; tmux and herdr pane placement and liveness were both verified separately in live isolated sessions (see the backend-liveness section above), where the herdr agent-state classifier's rovo blind spot is recorded as a Herdr-side integration gap to track, not a live-guard coverage gap this refresh command needs to close. diff --git a/docs/verification/runtime-backends.md b/docs/verification/runtime-backends.md index d880baaf4ab..1a77089ab3e 100644 --- a/docs/verification/runtime-backends.md +++ b/docs/verification/runtime-backends.md @@ -55,7 +55,7 @@ Observed identities, and the resulting verdict: | grok | 0.2.118 | `grok-0.2.118-ma` | `grok` | alive | | kimi | 0.31.1 | `kimi` | `kimi` | alive | -Claude Code is the harness whose title no longer attributes it at all; every other adapter is currently attributed by both sources. +In that 2026-08-03 seven-adapter run, Claude Code was the only harness whose title did not attribute it; every other adapter was attributed by both sources. Codex reported `codex-aarch64-a` at 0.145.0 and `codex` at 0.146.0, and Kimi Code reported `kimi-code` as its foreground `comm` at 0.29.1 and `kimi` at 0.31.1, so these identities move between ordinary patch releases in both directions. That is the evidence for treating any single process name as a surface under vendor control rather than a stable contract. @@ -63,6 +63,9 @@ The crewmate-only Muse Code 0.1.0-R708.1 adapter was verified separately on 2026 Its installed `muse-bin-0.1.0-R708.1` foreground identity classified `alive`, while `musescore`, `amuse`, `muse-binary`, and `muse-bind` remained ambiguous in the portable regression. [`muse.md`](muse.md#process-identity) owns the artifact identity and launcher evidence for that verification. +The crewmate/scout-only Rovo CLI 202609.1.2 adapter added `*rovo*` to the same glob family as `*grok*`/`*kimi*` in `fm_backend_tmux_classify_process_name`, and was relaunched live under tmux 3.6a in an isolated private socket. +`#{pane_current_command}` reported the truncated on-disk binary name `atlassian_cli_r` - macOS's 15-char `comm` truncation cuts `atlassian_cli_rovodev` off just before the `rovo` substring begins, the same truncation-volatility class codex/kimi's own patch-release name drift shows above - while the foreground ps-based `comm` correctly reported `rovo`, so `fm_backend_tmux_agent_state` returned `alive` through that primary source; the two-independent-name-sources design is exactly why the truncated title does not break the verdict. +[`rovo.md`](rovo.md#backend-liveness-tmux-verified-live-herdr-placement-verified-live-with-a-herdr-side-agent-detection-gap) owns the fuller record, including the busy/interrupt/exit facts captured in that same live tmux session and the herdr agent-detection gap found when herdr placement was verified live in an isolated lab session. GitHub Copilot CLI 1.0.81-7 was verified separately on 2026-08-21 under Git for Windows. Its native process is `copilot.exe`, while the Git Bash child can terminate at `PPID=1` before reaching that process. The adapter therefore accepts `COPILOT_LOADER_PID` only when `ps -W` proves that exact Windows PID is live and named `copilot.exe`; the portable tmux classifier accepts both `copilot` and `copilot.exe`. @@ -73,7 +76,6 @@ The native Windows Bearings command path was verified on 2026-08-27 against GitH The regression puts a twelve-second non-Git `bash.exe` first on `PATH`, preserves native Copilot process markers, invokes the canonical local-only snapshot through `bin/fm-windows-git-bash.ps1`, and requires the expected schema and `FM_HOME` within eight seconds. The exact command was `. .\bin\fm-windows-git-bash.ps1; $bash = Resolve-FirstmateGitBash; & $bash .\tests\fm-copilot-harness.test.sh` from PowerShell. Its relevant output was `ok - Copilot Windows bearings transport bypasses a hanging ambient bash`, followed by `all fm-copilot-harness tests passed`. - Bounded observed output: ```text @@ -91,14 +93,36 @@ alive On macOS the pane command reflected the rewritable title while the full install path could survive in `ps -o comm=`; in the Linux portable regression those roles reversed for the version-named native executable, with the identifying path retained in argv[0]. The classifier therefore accepts a harness basename first, then an exact harness path component in the full executable path, then the same component in argv[0], without depending on which field carries it on a given platform. -The portable regression is CI-enforced, while the real-harness drift guard is opt-in under the policy in `.agents/skills/firstmate-coding-guidelines/SKILL.md`. +The portable regression is CI-enforced. +The real-harness drift guard spends no model tokens, so under the policy in `.agents/skills/firstmate-coding-guidelines/SKILL.md` it runs by default wherever tmux is installed and reports a capability skip elsewhere; `FM_HARNESS_LIVENESS_DRIFT=1` additionally turns an absent tool into a failure. Run the live guard after any harness upgrade and before trusting or refreshing the table above: ```sh FM_HARNESS_LIVENESS_DRIFT=1 bin/fm-test-run.sh tests/fm-harness-liveness-drift-live-e2e.test.sh ``` -Bounded output from the run that produced the table: +### 2026-09-06 default-on drift refresh, and the Cursor editor CLI collision + +Running the guard with no variable set on macOS 26.5.2 arm64 checked 8 installed harnesses and classified every one `alive`: + +```text +# claude 2.1.263 (Claude Code): title='2.1.263' foreground=[/Users/kunchen/.local/bin/claude <defunct> <defunct> ] +# codex codex-cli 0.147.0: title='codex' foreground=[/opt/homebrew/bin/codex ] +# opencode 1.18.29: title='opencode' foreground=[/opt/homebrew/bin/opencode ] +# pi 0.84.4: title='pi-launcher' foreground=[/opt/homebrew/bin/pi-signed .../pi ] +# pi-signed 0.84.4: title='pi-launcher' foreground=[/opt/homebrew/bin/pi-signed .../pi ] +# grok grok 1.0.13 (5e9a58528b76) [stable]: title='grok-1.0.13-mac' foreground=[/Users/kunchen/.local/bin/grok ] +# cursor 2026.09.02-c22c1a3: title='node' foreground=[/Users/kunchen/.local/bin/cursor-agent ] +# muse Muse Code 1.0.3 (1.0.3-R2198.1): title='muse-bin-1.0.3-' foreground=[/Users/kunchen/.local/bin/muse-bin-1.0.3-R2198.1 ] +# checked 8 installed harness(es) +``` + +The first default-on run failed on Cursor with `LIVENESS DRIFT: cursor unknown is running but classifies 'missing'`, observed title `zsh`. +The classifier was not at fault: the guard resolved the harness through a generic `command -v cursor`, which on a machine that also has the Cursor editor finds `~/.local/bin/cursor` - the editor launcher, not the agent. +That binary exits immediately, leaving a bare shell in the pane. +The guard now asks `fm_cursor_resolve_binary` first for `cursor`, which is the same verified owner `bin/fm-spawn.sh` uses, so the probe launches `cursor-agent` and the editor CLI can no longer masquerade as the harness. + +Bounded output from the 2026-08-03 run that produced the first table above: ```text ok - harness liveness: claude 2.1.220 (Claude Code) classifies alive @@ -1610,3 +1634,51 @@ It names the installed version and the floor rather than degrading quietly, and The same guard against the pre-change extension in the same lab measured a 676.9 ms worst keystroke echo while delivering two outcomes and a 295.3 ms worst echo with nothing to deliver, against a 49.2 ms extension-free floor, and failed as designed. Measured through the same real `fm_branch_report` tool and real `bin/` scripts with a 1 ms interval timer, the largest single block of the JavaScript thread fell from 273 ms to 2.0 ms for a routine outcome, from 286 ms to 2.0 ms for a captain outcome, and from 134 ms to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle-loop floor. Those absolute figures are specific to this host and Pi version; the guards assert the relationship (delivery must stay in the class of the same machine's own floor) rather than a remembered millisecond number. + +## Oh My Pi (omp) + +omp runs crewmate, scout, secondmate, and primary work; [`supervision.md`](supervision.md#omp-oh-my-pi-native-delivery-2026-09-05) owns the primary evidence. +The evidence below was produced on 2026-09-05 against omp 18.1.11 (`~/.local/bin/omp`, a Bun-compiled single binary) on macOS 26 arm64 through the Herdr backend with the `openai-codex/gpt-6-astra` model, building on the 2026-09-02 adapter investigation against 18.1.2. + +### Process identity and markers + +`ps -o comm=` reports the bare name `omp` for the agent process, from both its `!` bash path and the model's bash tool, so identity is the anchored name; `ompd` and `comp` never match. +omp publishes no harness marker: `PI_CODING_AGENT` is absent from the binary, and the default profile sets neither `PI_CODING_AGENT_DIR` nor `OMP_PROFILE` in the process environment. +`FM_OMP_HARNESS=omp` is Firstmate's own launch marker and wins over an inherited `CLAUDECODE` only under a real omp ancestor; `tests/fm-omp-harness.test.sh` pins both directions with real processes. + +### Composer + +Under the captain's `unicode` symbol preset the idle screen through Herdr was a bare `❯` (U+276F) row followed directly by the status row: + +```text +❯ + π · ◔ GPT-6-Astra · 🌳 …-workspace · ⑂ detached · ◫ 15.4%/272K ⟲ · (sub) +``` + +Before the status-row rule the shared classifier folded that row into the bare composer's wrap region and read the idle pane `pending`, so `bin/fm-send.sh` skipped its doorbell on the first live omp worker. +After the rule, the same live Herdr capture read `empty`, a steer's doorbell landed, and the worker opened a turn on it. +`tests/fm-composer-lib.test.sh` pins the unicode idle row, the nerd-preset idle row, the busy spinner row, and typed text over the same fixture in both locales. + +### Busy state and lifecycle + +| Fact | Observed | +| --- | --- | +| Semantic source | `omp-ext`: `busy source=omp-ext event=agent-start` on the brief, `idle source=omp-ext event=agent-end` at its natural end, `busy` again on a steer, `idle` after a control-plane interrupt | +| Rendered busy row | `⎋ Working…` (U+2026) above the composer and a braille spinner plus elapsed cell (`⠧ 36s`) in the status row; the omp busy regex accepts only those two TUI signals, not the `Working...` that headless `-p` writes to stderr, since no supervised omp pane runs headless | +| Interrupt | `bin/fm-control.sh <id> interrupt` delivered a single Escape (`verified=agent-alive cancel=unconfirmed`), the composer read `empty`, and omp raised `agent_end` | +| Exit | `bin/fm-control.sh <id> exit` typed `/quit`; Herdr then reported the pane `dead` | +| Extension loading | a file named both by `-e` and by `<cwd>/.omp/extensions` loads twice; discovery is top-level and cwd-only | +| Extension tools | the openai-codex model invokes a registered tool by writing `xd://<tool>` through omp's virtual-file bridge | + +### End-to-end + +A throwaway scout was spawned through `bin/fm-spawn.sh --scout --harness omp --model openai-codex/gpt-6-astra --effort low` on Herdr and driven to completion: + +1. the launch delivered its brief positionally under the tracked posture overlay and the agent executed it with no approval prompt; +2. the busy record moved seed, busy, idle exactly as the extension contract states; +3. the report landed and the `done:` status line was appended; +4. `bin/fm-send.sh` rang the doorbell once the composer read `empty`, and the worker opened a turn on the inbox record; +5. `bin/fm-control.sh <id> interrupt` cancelled the running turn; +6. `bin/fm-control.sh <id> exit` stopped the agent and `bin/fm-teardown.sh` returned the worktree and closed the item. + +`FM_OMP_LIVE_E2E=1 tests/fm-omp-primary-live-e2e.test.sh` refreshes the primary evidence; the worker path above is refreshed by repeating the scout dispatch after any omp upgrade. diff --git a/docs/verification/supervision.md b/docs/verification/supervision.md index 98c487d43d9..388b09c82f2 100644 --- a/docs/verification/supervision.md +++ b/docs/verification/supervision.md @@ -51,6 +51,13 @@ That cold positional-prompt check established eventual custom-message delivery, The installed pi-signed 0.82.0 wrapper repeated the shared Pi primary extension and session-start path on 2026-07-27. [`runtime-backends.md`](runtime-backends.md#tmux) owns the shared-ancestry evidence and authoritative selection-marker boundary. +### omp (Oh My Pi) native delivery, 2026-09-05 + +The omp Run-tier adapter was verified on 2026-09-05 with omp 18.1.11 and the openai-codex `gpt-6-astra` model through `FM_OMP_LIVE_E2E=1 tests/fm-omp-primary-live-e2e.test.sh`, which drives a real omp in its JSON-RPC stdio mode inside an isolated lab clone. +Both tracked `.omp/extensions/*.ts` files loaded by auto-discovery alone (no `-e`, no trust dialog), `before_agent_start` returned the digest as a persistent context message, the model quoted the lab's `SESSION START -` heading back on its first turn, `state/.session-start-complete` was recorded, and `state/.lock` named the omp process, so ancestry detection identified the markerless binary. +omp's `session_start` payload carries no reason field, so the adapter derives the source: the first start of the process is `startup` (or `resume` from a `--continue`/`--resume` launch line) and a later in-process start is `clear`; `tests/fm-omp-harness.test.sh` pins that mapping over a fake omp API. +A file named both by `-e` and by auto-discovery loads twice (two factory calls, doubled `session_stop` continuations), which is why the secondmate launch names no `-e` and the per-task worker extension lives in `state/`. + ### Run-tier source vocabulary and context-reset injection The run tier depends on three facts only the vendor can supply: the session-open source it reports, whether hook stdout reaches model context on a context-RESET open rather than only a cold one, and whether a worker the hook detaches survives the hook returning. @@ -208,6 +215,7 @@ Each pass polled `state/<id>.busy-state` while a real turn ran. | Harness | Version verified | Semantic source | Observed result | | --- | --- | --- | --- | | Pi | 0.82.0 | Extension `agent_start` / `agent_settled` with `ctx.isIdle()` | The spawn seed `busy source=fm-spawn`, then `busy source=pi-ext event=agent-start`, then `idle source=pi-ext event=agent-settled`; the turn-end marker was still touched. | +| omp | 18.1.11 | Extension `agent_start` / `agent_end` without `willContinue` | Live Herdr scout on `openai-codex/gpt-6-astra` (2026-09-05): the spawn seed `busy source=fm-spawn`, then `busy source=omp-ext event=agent-start`, then `idle source=omp-ext event=agent-end` at the natural end of the brief; a steer through `fm-send` reopened `busy … agent-start`, and a control-plane interrupt closed it with `idle … agent-end` (omp fires `agent_end` on an interrupted turn). `ctx.isIdle()` is deliberately not consulted because it reads false at a natural TUI `agent_end`. | | OpenCode | 1.17.18 | Plugin `session.status` | In a real TUI pane: seed, then `busy source=opencode-plugin event=session-busy`, then `idle source=opencode-plugin event=session-status-idle`. | | Claude | 2.1.220 (Claude Code) | Hooks `UserPromptSubmit`, `Stop`, `StopFailure`, `SessionEnd` | `UserPromptSubmit` fired for the argv launch prompt and each steer, and `Stop` closed every completed turn. A mid-stream Escape interrupt fired no closing hook, which is why the firstmate-controlled clear exists. `StopFailure` and `SessionEnd` are wired from the four hook names present in the installed binary; only the abnormal paths they cover were not reproduced live. | | Copilot | 1.0.81-7 | Hooks `userPromptSubmitted`, `agentStop`, and `sessionEnd` | Genuine payloads established all three events: prompt submission writes `busy source=copilot-hook event=user-prompt-submitted` and a submission token, while stop and session end write idle and touch the turn-ended marker. | @@ -237,7 +245,7 @@ tests/fm-crew-state.test.sh ## Turn-end guard -The blocking and bounded-follow-up mechanisms were validated across seven harnesses on 2026-07-08 through 2026-09-01, with Claude's replacement Stop-owned path revalidated on 2026-07-24, Cursor's stop-hook park validated on 2026-08-13, and Copilot's asynchronous shell-completion re-entry validated on 2026-09-01. +The blocking and bounded-follow-up mechanisms were validated across eight harnesses on 2026-07-08 through 2026-09-05, with Claude's replacement Stop-owned path revalidated on 2026-07-24, Cursor's stop-hook park validated on 2026-08-13, Copilot's asynchronous shell-completion re-entry validated on 2026-09-01, and omp's blocking `session_stop` hook validated on 2026-09-05. | Harness | Version verified | Mechanism | Observed result | | --- | --- | --- | --- | @@ -246,6 +254,7 @@ The blocking and bounded-follow-up mechanisms were validated across seven harnes | Copilot | 1.0.83-0 | Tracked asynchronous shell task plus completion notification; nonblocking `agentStop` repair backstop | The initiating turn continued after starting the background PowerShell task, `shell_completed` fired the repository notification hook, hook-provided context produced a follow-up turn, and the portable contract kept missing-watcher repair continuations below the vendor's eighth-block override without parking in the hook. | | OpenCode | 1.17.6 | Passive `session.idle` callback | Throwing could not block, while `promptAsync` scheduled one TUI follow-up; headless remained fail-open. | | Pi | 0.80.5 | Passive `agent_settled` callback | Exactly one guard follow-up ran for an unhealthy cycle, with no recursion across tool turns. | +| omp | 18.1.11 | Blocking `session_stop` hook returning `{ continue: true, additionalContext }` | In the isolated rpc lab (2026-09-05), the successor watcher was frozen with `SIGSTOP` until its beacon passed the lab `FM_GUARD_GRACE` of 20s while its arm child stayed attached (a killed watcher closes its arm child and the extension re-arms before the guard can fire); the next turn end raised the guard, the guard spy recorded `rc=2` followed by a stop carrying `stop_hook_active: true`, omp compelled a continuation carrying the `turn-end-guard` operational text, the `fm_watch_arm_omp` invocation count then rose to at least two, and a live watcher held the home lock after the thaw; the flagged stop was allowed, so exactly one continuation ran. `session_stop` never fired for an interrupted turn. | | Grok | 0.2.112 native and 0.2.73 pre-native | Running-payload adaptive `Stop` | Native false-to-true continuation stayed in one process with two model turns and zero resume launches; the field-absent pre-native process launched exactly one guarded resume. | | Cursor | 2026.08.11-e8db854 | Awaited `stop` hook park returning one `followup_message` | Exit 2 ended the turn normally, proving it cannot block; a returned follow-up ran a genuine second turn; a sleeping hook held the boundary open and the wake landed after it; `loop_limit` stopped the hook being invoked at its ceiling. | @@ -484,6 +493,7 @@ grok 0.2.103 (89c3d36fb6f1) [stable] | Codex | `FM_CODEX_LIVE_E2E=1 tests/fm-codex-continuity-live-e2e.test.sh` | The one-second foreground checkpoint returned without switching to the arm wrapper. | | OpenCode | `FM_OPENCODE_LIVE_E2E=1 tests/fm-opencode-primary-live-e2e.test.sh` | A verified successor existed before prompt handling, with no model re-arm or turn-end fallback. | | Pi | `FM_PI_LIVE_E2E=1 tests/fm-pi-primary-live-e2e.test.sh` | One initial tool call led to extension-owned successors and clean child retirement on exit. | +| omp | `FM_OMP_LIVE_E2E=1 tests/fm-omp-primary-live-e2e.test.sh` | One initial `fm_watch_arm_omp` invocation (the openai-codex model reaches extension tools through omp's `xd://` virtual-file bridge, a `write` to `xd://fm_watch_arm_omp`, counted as the same invocation) started a live watcher; an actionable close spawned a ledger-linked successor and woke main exactly once; the lab is reaped by path, and omp 18.1.11 did not exit within 30s of its rpc stdin closing, recorded as a note. omp 18.1.11, 2026-09-05. | | Grok | `FM_GROK_LIVE_E2E=1 tests/fm-grok-continuity-live-e2e.test.sh` | Native task completion surfaced the actionable close and the cycle ledger recorded `reason=actionable-signal`. | Pi 0.81.1 repeated the continuity and clean-exit lifecycle on 2026-07-23 after the Calm presentation changes. diff --git a/docs/verification/trace-context.md b/docs/verification/trace-context.md index c6af19f8d41..f7b1edaa743 100644 --- a/docs/verification/trace-context.md +++ b/docs/verification/trace-context.md @@ -9,7 +9,7 @@ Comparison base: `main` at `976d97f`. The colocated unit suite `tests/fm-trace-context-lib.test.sh` (26 assertions) exercises validation (valid accepted; malformed, wrong-length, uppercase, all-zero, `ff` version, and shell-metacharacter values rejected), root minting with every mint a distinct sampled root and no parent-adoption input, the recovery reuse path with the recorded carrier winning over the ambient environment, default-off omission, the enable precedence of `FM_TRACE_CONTEXT` over `config/trace-context` with unset or empty deferring to the file, normalized home-session state, atomic replacement of a read-only prior record, stale-session rejection after failed publication, missing or invalid state defaulting off, the Secondmate home-session boundary with later file state plus the per-task trace boundary (two resolves under one persistent ambient `TRACEPARENT` root two distinct traces and adopt neither), forced entropy failure omitting safely, and the minted-root fixed-shape check. -The spawn-path integration suite `tests/fm-trace-context-spawn.test.sh` (12 assertions), hermetic against an ambient `FM_TRACE_CONTEXT`, drives `bin/fm-spawn.sh` end to end with a fake tmux pane and a real isolated git worktree: enabled, one resolved carrier is recorded as `traceparent=` in the meta only after the identical `TRACEPARENT` export is sent before the launch literal; disabled, neither is written nor sent (only `GOTMPDIR` is); a failed carrier delivery leaves no `traceparent=` claim while the source task still launches; an unsafe delivery whose partial input cannot be cleared stops before appending the launch command; a failed metadata append removes the carrier from the launched task without aborting it; duplicate Secondmate preflight leaves inherited trace configuration unchanged; a relaunch reuses the recorded carrier verbatim; and spawns ignore later config and environment edits in favor of the frozen home-session decision. +The spawn-path integration suite `tests/fm-trace-context-spawn.test.sh` (12 assertions), hermetic against an ambient `FM_TRACE_CONTEXT`, drives `bin/fm-spawn.sh` end to end with a fake tmux pane and a real isolated git worktree: enabled, one resolved carrier is recorded as `traceparent=` in the meta only after the identical `TRACEPARENT` export is sent before the launch literal; disabled, neither is written nor sent (`GOTMPDIR` still is); a failed carrier delivery leaves no `traceparent=` claim while the source task still launches; an unsafe delivery whose partial input cannot be cleared stops before appending the launch command; a failed metadata append removes the carrier from the launched task without aborting it; duplicate Secondmate preflight leaves inherited trace configuration unchanged; a relaunch reuses the recorded carrier verbatim; and spawns ignore later config and environment edits in favor of the frozen home-session decision. The per-task boundary regression models the reviewed Secondmate scenario exactly: two unrelated tasks spawned sequentially from one home while the same fixed `TRACEPARENT` sits in the spawning environment (a persistent Secondmate's launch-time carrier) record and inject valid carriers whose trace ids differ from each other and from the ambient carrier, and a relaunch of the first task reuses its original carrier verbatim for both the meta record and the injected export. Two further assertions drive a genuine two-level primary -> Secondmate -> worker chain, running `bin/fm-spawn.sh` twice with the exact environment the primary injects into the Secondmate, and prove the primary's effective override governs the nested worker both ways: env-on with no config file keeps the nested worker enabled while it roots its own per-task trace distinct from the Secondmate's carrier, and env-off with the file present keeps the nested worker disabled even though the `config/trace-context` file was copied into the Secondmate home. A final assertion drives the file-decided path (`FM_TRACE_CONTEXT` unset) and proves the Secondmate's recorded/injected carrier and its delivered `FM_TRACE_CONTEXT=on|off` snapshot are always derived from one frozen decision, so a carrier is never paired with the opposite enable state. diff --git a/docs/watcher-continuity.md b/docs/watcher-continuity.md index f4034fbeaeb..2db5c5482e6 100644 --- a/docs/watcher-continuity.md +++ b/docs/watcher-continuity.md @@ -5,11 +5,12 @@ Must-work continuity now lives above that process boundary instead of depending ## Ownership -Pi's `.pi/extensions/fm-primary-pi-watch.ts` and OpenCode's `.opencode/plugins/fm-primary-watch-arm.js` own continuous re-arm after an actionable child close. +Pi's `.pi/extensions/fm-primary-pi-watch.ts`, omp's `.omp/extensions/fm-primary-omp-watch.ts`, and OpenCode's `.opencode/plugins/fm-primary-watch-arm.js` own continuous re-arm after an actionable child close. Each adapter starts the next arm before delivering the wake prompt, checks current session-lock ownership at launch, preserves one child or scheduled retry at a time, and applies bounded exponential retry after an unexpected or failed close. A failed follow-up never cancels continuity restoration. Pi same-process session replacement follows the generation-owner contract in `.pi/extensions/fm-primary-pi-watch.ts`: an owning `session_start` arms the replacement generation without waiting for a model turn, and a state-scoped replacement handoff carries every actionable close whose delivery overlapped `session_shutdown`, including a main follow-up Pi accepted but had not yet consumed, branch handling, and a retiring child that reports after the bounded shutdown wait. A main follow-up counts as delivered once Pi accepts it, never once the model reads it, because a follow-up queued while main is streaming joins the running run without a `before_agent_start`; the extension header owns how consumption is observed and why it only decides what a replacement replays. +omp's replacement follows the same generation-owner contract in `.omp/extensions/fm-primary-omp-watch.ts`, whose header owns the one difference: omp reports no shutdown reason, so every shutdown with a pending actionable close persists the handoff for the next owning `session_start` to replay. Cursor's `.cursor/hooks.json` `stop` hook (`bin/fm-turnend-guard-cursor.sh`) owns routine tokenless re-arm for a Cursor primary by parking that awaited hook on `bin/fm-watch-arm.sh` and returning an actionable close as one follow-up; [`turnend-guard.md`](turnend-guard.md#harness-integrations) owns its Pi-host stand-down, loop bounds, and supersession baton. Copilot runs `bin/fm-watch-arm.sh` through its shell tool's native asynchronous mode, with `bin/fm-watch-arm.ps1` as the native Windows bridge. The tracked `notification` hook converts `shell_completed` and `shell_detached_completed` into typed Firstmate input only when the completed watcher left queued work or required supervision unhealthy. @@ -26,7 +27,7 @@ While supervision is still needed and away mode remains inactive, an actionable ## Actionable wake ordering -After an actionable Pi or OpenCode child close, the adapter starts and verifies one singleton successor before it delivers the original wake. +After an actionable Pi, omp, or OpenCode child close, the adapter starts and verifies one singleton successor before it delivers the original wake. It confirms the handling handoff against that successor before scheduling the follow-up, retries once against the current generation and successor, and treats a failed confirmation as a restoration failure: it classifies the error, retires a successor that is no longer alive, and surfaces exactly one typed message. A failed confirmation is never swallowed. It waits at most one readiness timeout per attempt, then sends TERM and waits a bounded retirement confirmation before the next lock-verified exponential retry. @@ -35,7 +36,7 @@ When that retained arm later closes, its actual close is classified as a new sup After the configured retry bound is exhausted, it delivers the original wake with a typed continuity-restoration failure even if every successor arm hung without reporting readiness. This is deliberate Option B ordering: the fleet is protected before the model handles the wake whenever restoration succeeds, but the model is never left blind when it does not. -Claude's Stop hook starts the successor arm at the next Stop after the handling turn, rather than before notification as Pi and OpenCode do. +Claude's Stop hook starts the successor arm at the next Stop after the handling turn, rather than before notification as Pi, omp, and OpenCode do. The durable wake queue preserves actionable events during the residual active-turn window, and the bounded turn-end guard enforces recovery at Stop when no watcher is live and no open generation claim is still deciding, so a finished, hung, or identity-mismatched claim cannot suppress it ([`turnend-guard.md`](turnend-guard.md#harness-integrations) owns that boundary). The recovery-episode contract below owns once-per-generation announcement. A handling successor does not re-announce; it enters its poll loop immediately and keeps scanning signals, stale panes, and checks. @@ -77,6 +78,13 @@ Main records its presented set in `state/.main-eligible-rows`. A branch grant is published through `bin/fm-wake-grant.sh` under that same lock in `state/.branch-eligible-rows`, bound to the live branch process and extension generation recorded in `state/.branch-eligible-owner`, and publication is refused if main already claimed any requested row. A main drain validates that owner evidence under the queue lock and reclaims the grant when its process is gone or its identity no longer matches. A main drain claims every currently unclaimed row and excludes an active branch grant from both presentation and acknowledgement. +Because that exclusion makes those rows invisible to main, `bin/fm-guard.sh`'s queued-wake warning counts only the rows the calling actor can itself present or retire, so an actor is never sent to a drain that provably has nothing for it. +`bin/fm-wake-lib.sh` owns that per-actor count (`fm_wake_actor_pending_count`) alongside the grant row-list and owner-record reads that the drain and `bin/fm-wake-grant.sh` share. +A row a live grant reserves is therefore never counted as drainable for main; rather than going silent about a visibly non-empty queue, the guard prints a distinct advisory naming the live supervision branch as the holder and saying not to drain those rows from here. +The branch actor's queued-wake output stays suppressed in every case. +A main drain with nothing of its own left, and a live grant still holding the queue, says so in one bounded line instead of exiting silently. +A row that lost the five appended fields or its numeric sequence can never be claimed, presented, or named by an `--ack-through` cutoff, so a main drain retires it under the queue lock and reports how many it removed together with those rows verbatim, bounded to the first 20 and a count of the rest, because the queue was their only durable record; a branch drain never does, because a grant can only name sequences that were structurally valid when it was published. +A retirement that cannot be read or written is reported and never fails the drain: the rows that remain usable are still presented with their acknowledgement command, the unusable ones stay queued for a later drain to retire, and failing the whole drain would strand the usable rows too. Its `--ack-through <SEQ>` deletes only claimed main rows at or below the cutoff, while a branch acknowledgement deletes only claimed branch rows at or below its cutoff. Every settled branch prompt releases any residual grant, so an omitted or failed acknowledgement leaves the durable row available to a later main drain; a successful acknowledgement has already removed it. An acknowledgement whose cutoff removes none of the actor's rows while a presented row above the cutoff still waits is reported as having acknowledged nothing, together with the exact `--ack-through` and `--recovery-generation` command for that presented row; the presented set is read before any re-claim, so a row that arrived after presentation is never named for unseen acknowledgement. @@ -87,6 +95,7 @@ A check-kind row is main-owned in every mode, including a heartbeat review, so i A missing or empty branch snapshot is refused loudly rather than read as "nothing eligible", because reaching the drain without the non-empty handoff promised by the extension is a wiring bug. Because branch claims contain no check-kind rows, a branch acknowledgement skips check-specific receipt scans. `tests/fm-wake-queue.test.sh`'s mixed-queue actor, stale-acknowledgement remedy, and presentation-deadline tests drive the real scripts: branch acknowledgement cannot swallow a main row, a concurrent main turn cannot present or acknowledge an active branch grant, a no-op stale acknowledgement names the current presented wake's exact command, live-holder presentation contention stays bounded and retriable, and acknowledgement locking remains blocking. +The same suite pins the counted-equals-presentable invariant against `bin/fm-guard.sh` and `bin/fm-wake-drain.sh` together: a branch-held row raises the held advisory rather than the ordinary queued-wake warning for main, and is presented with its acknowledgement command - with the ordinary warning restored - as soon as the grant clears, and structurally unusable rows are retired by main alone while every remaining row stays presentable and acknowledgeable. `tests/fm-pi-branch-extension.test.sh` pins extension-side classification, claim publication and release, and the pre-drain recheck. ## Arm-layer cycle contract @@ -122,7 +131,7 @@ It also covers generation-claim single-flight, stuck-claim supersession, superse ## Active limits and verification -The goal is continuity without a Pi or OpenCode model-memory re-arm step. +The goal is continuity without a Pi, omp, or OpenCode model-memory re-arm step. No zero-latency guarantee is claimed because lock verification, watcher startup, and bounded retry delays remain deliberate safety work. OpenCode support targets persistent TUI sessions rather than headless `opencode run`. Claude depends on the Stop `asyncRewake` rewake, Copilot depends on tracked asynchronous shell completion notifications, Cursor depends on its awaited stop-hook park, Grok retains native background-completion notifications, and Codex retains bounded foreground checkpoints. diff --git a/skills/reconcile-firstmate-upstream/SKILL.md b/skills/reconcile-firstmate-upstream/SKILL.md index 4ea70000d69..5166592c77e 100644 --- a/skills/reconcile-firstmate-upstream/SKILL.md +++ b/skills/reconcile-firstmate-upstream/SKILL.md @@ -86,6 +86,16 @@ This step is complete when one prior upstream SHA is proven or the run has stopp ## 3. Choose the history path +Enable repository-local resolution reuse before introducing conflicts: + +```sh +git config --local rerere.enabled true +git config --local rerere.autoupdate false +``` + +Record the previous values and leave global Git settings unchanged. +Treat reused resolutions as review inputs, not approved changes: automatic index updates stay disabled. + When the frozen upstream SHA is already an ancestor of the fork SHA and content comparison shows no missing upstream change, report that the fork is already reconciled and stop without creating a PR. Use a normal merge when the prior upstream SHA is in both histories and the merge-base path shows that the previous synchronization preserved upstream ancestry. Create the feature branch from the frozen fork SHA and merge the frozen upstream SHA without committing: @@ -115,6 +125,8 @@ This step is complete when a normal merge records the frozen upstream SHA in `ME ## 4. Resolve conflicts from sources List every unresolved path with `git diff --name-only --diff-filter=U`. +After a three-way patch application, run `git rerere` to record the conflict preimages or apply a recorded resolution. +Inspect `git rerere status`, `git rerere diff`, and the resulting worktree before staging any reused resolution. For each path, inspect the base, fork, and upstream versions, then trace the commits and PRs that introduced both sides. Use `git log --all -- <path>`, blame, nearby tests, owned documentation, and hosted pull-request history rather than inferring intent from conflict markers alone. @@ -140,33 +152,29 @@ Add a newly required portability fix to the manifest before staging it rather th This step is complete when every conflict decision has a primary-source rationale and no unresolved or accidental conflict artifact remains. -## 5. Run observable local validation +## 5. Establish PR readiness with bounded local checks On Windows, run shell commands with Git for Windows `bash.exe`. Use the literal executable recorded during the freeze step instead of invoking ambient `bash`, which may select WSL. Keep tracked shell files as LF in the worktree before syntax and lint checks; a local line-ending repair is not a content change to commit. -The validation invariant is **focus**. -Select once, retry only failures, and compare only a reproducible failing case with frozen upstream. +The validation invariant is **routing**. +Local checks establish readiness to open an ordinary PR; GitHub Actions establishes broad regression evidence before merge. +Keep every mapped subject accounted for, but execute only fast gates and selected conflict-resolution cases locally. +Route broad portable coverage to the existing Linux lanes, real backend coverage to its dedicated lane, and native-platform contracts to their platform jobs. +Use `.github/workflows/ci.yml` and the runner's lane listings to identify actual owners rather than inventing a new complete matrix. +A full-inventory selection is a routing signal, never a reason to stop the reconciliation or start a full local suite. -### Select without a history explosion +### Inventory the exact tree -For a normal merge, list the changed-file-informed suite before starting it: +For a normal merge, list the changed-file-informed inventory: ```sh bin/fm-test-run.sh --list --changed --base <fork-sha> ``` -Count the listed scripts before executing them. -If the list expands to the complete inventory, stop before the run because the comparison is not focused. -Run the focused result only after that check: - -```sh -bin/fm-test-run.sh --changed --base <fork-sha> -``` - -For a reconstruction, do not run the changed selector directly from the upstream-based branch because its three-dot comparison can expand to the complete inventory. -Snapshot the resolved index into a temporary commit whose parent is the frozen fork, then run the existing mapper from a detached worktree: +For a reconstruction, avoid the upstream-based branch's three-dot history expansion. +Snapshot the resolved index into a temporary commit parented to the frozen fork and list from its detached worktree: ```sh reconciled_tree=$(git write-tree) @@ -182,62 +190,41 @@ git -C <outside-repository-test-worktree> diff --exit-code ) ``` -This gives the runner a fork-parented synthetic history while testing the exact staged tree. -Count the listed scripts and stop if the selection equals the complete inventory. -Run only the focused result: - -```sh -( - cd <outside-repository-test-worktree> - bin/fm-test-run.sh --changed --base <fork-sha> -) -``` - -Add only explicit tests for manually composed conflicts that the mapper does not select. -Keep the runner's automatic scheduler on Windows because it already serializes measured process-heavy scripts. -Do not force extra parallelism. -After any fix, remove the old synthetic worktree and rebuild it from the new index before the final changed run. - -Report the selected script count before the run. -Keep the runner output visible or retain its process handle and output stream when it must run asynchronously. -Use its `FM_TEST_BEGIN`, `FM_TEST_END`, `FM_TEST_SUMMARY`, and `FM_TEST_SLOWEST` markers to report the active script, completed count, failures, elapsed time, and slowest completed scripts. -A live process with advancing markers is progress; a process beyond the runner's per-script bound is a bounded failure to investigate rather than an opaque wait. - -Do not run `--all`, a complete lane matrix, or a `tests/*.test.sh` walk during this skill. -When a focused failure reveals another affected surface, add only its smallest relevant script to the focused run. -GitHub Actions owns the complete portable, Herdr, Windows, and macOS matrix; a separately requested full local regression is outside this skill. - -### Triage without restarting long scripts - -Rerun each failed script once with `--jobs 1` before editing code. -When the failure names a case and the script already has a case filter, use that filter for this retry instead of restarting the complete script. -A serial pass classifies the first result as contention and ends that investigation. - -When one named case in a long script fails, rerun that case through an existing filter. -When no filter exists, a temporary diagnostic selector may isolate the case, but remove it before staging. -After a fix, rerun the failed case and every case that shares the changed helper. -Rerun the complete script only when the production change or shared fixture makes its remaining cases relevant. - -Create a detached worktree at the literal upstream SHA only after one case is reproducibly red: - -```sh -git worktree add --detach <outside-repository-baseline-worktree> <upstream-sha> -``` - -Keep baseline implementation files byte-exact. -A selector-only test edit is allowed to expose the same case, and it must be restored before removing the worktree. -Run one baseline case per process with an external wall-clock bound that can terminate its complete process tree. -Do not batch several baseline cases behind one unbounded shell because the first Windows hang can hide every later result. - -If the frozen upstream case fails or hangs the same way, record it as a baseline incompatibility rather than changing reconciliation code. -Make a local test correction only when the fixture itself is non-portable and the correction preserves the asserted behavior. -Use `fm_test_make_symlink` for directory or file identity fixtures on Git for Windows, where plain `ln -s` may copy a directory. -Keep executable tool shims on their existing portable mechanism because native Windows symlinks are not reliable through a temporary `PATH`. -Normalize `\r` in line-oriented assertions instead of changing product output solely for CRLF. -Give positive asynchronous test events measured Windows polling headroom, but do not widen production deadlines or negative timeout assertions. -Do not redesign `bin/fm-timeout-lib.sh` unless the frozen-upstream differential proves the timeout implementation is the regression. - -Run every repository gate: +Verify its parent, tree, and clean state before running a case. +Record the inventory count and a local/CI disposition for every selected script. +Every CI disposition needs an existing job and a reason; an absent capable runner is an explicit coverage gap, not a pass. +Choose local cases from the actual composed conflicts, prioritizing launch/adapter routing, native command transport, ownership, rollback, and changed test-selection behavior. +Use a named case when a script is long. +`tests/lib.sh` owns the shared `FM_TEST_ONLY` and `FM_TEST_LIST_CASES` interface for suites registered with `fm_test_run_cases`. +For a suite without a case seam, defer its whole script to CI or add a supported case registry rather than repeatedly injecting temporary selectors. + +### Preflight, bound, and checkpoint + +Read the header or `--help` of `scripts/validate-local.mjs` before using it. +It is the orchestration owner for a local check plan, not a replacement for the behavior runner or timeout library. +Create its plan and state directory outside the repository. +Plan only explicit local commands; the changed inventory belongs in the routing evidence, not in an executable full-suite plan. +Declare each check's CI owner and prerequisite probes, including its interpreter, compiler, utilities, and harness package when relevant. +Probe relevant Git fixture policies too; preserve host signing, hooks, and `safe.bareRepository` safeguards rather than disabling them to make a fixture pass. +Use the script's preflight-only mode when establishing host capability. +An absent prerequisite is reported and routed to an equipped CI job; install a dependency locally only when local proof is necessary and the missing dependency has been demonstrated. +Respect optional live-test gates and keep credentialed harness work out of this external-maintainer workflow. + +Use the validator's default 20-minute command budget and two-timeout circuit breaker unless the user explicitly chooses another budget. +It invokes `fm-timeout-lib.sh` for real cancellation during a command, including preparation, rather than relying on the runner's post-run `--max-wall-ms` check. +Budget exhaustion or the timeout circuit breaker defers remaining local commands to their CI owners and advances toward PR creation; it neither certifies them nor ends the reconciliation. +Do not compensate by increasing concurrency, repeatedly resetting the budget, or widening production deadlines. +One serial representative-case retry and, when needed, one frozen-upstream differential share the same local budget. + +The validator writes results after each command, preserves per-invocation logs, and caches only successful checks that reported no skip. +Caching is opt-in: declare complete input files/directories and version probes, and use resume mode only after reviewing those dependencies. +The key includes frozen context, command, declared input bytes and modes, environment digest, host/tool versions, and validator/timeout implementation. +Keep caching off for external-state, git-history-dependent, or incompletely modeled checks. +Never reuse an interrupted, failed, skipped, or merely preflighted result as a pass. +After an edit, refresh the exact-tree snapshot and rerun only checks whose inputs changed; unrelated prose changes need not invalidate code checks with proven independent inputs. +After a forced interruption, remove a validator lock only after proving its recorded owner is dead. + +Include all four repository gates in the local plan, reusing only valid unchanged-input evidence: ```sh while IFS= read -r script; do /bin/bash -n "$script" || exit; done < <(bin/fm-lint.sh --list-files) @@ -247,12 +234,19 @@ bin/fm-test-run.sh --check-coverage ``` Record each command, result, duration, and explicit optional-tool skip. +Keep the validator's `FM_RECONCILE_*` markers and underlying runner output observable. +Review a local failure before editing production code. +Retry its named case once serially; a pass classifies contention, while a reproducible failure may justify a detached frozen-upstream worktree for that same bounded case. +Keep baseline implementation bytes exact and record any selector-only test adaptation. +Only a matching differential proves a platform baseline incompatibility; unresolved failures remain unresolved in PR evidence. +Keep Git-for-Windows identity fixtures on `fm_test_make_symlink`, normalize CRLF in assertions rather than product output, and preserve negative timeout assertions. +Do not redesign `bin/fm-timeout-lib.sh` without differential evidence. Clean only test-owned processes and temporary paths, never broad process classes. -Restore and remove the baseline worktree after its results are recorded. -Remove the synthetic test worktree after the final focused run. +Remove clean synthetic and baseline worktrees after retaining their evidence. Review the full branch diff after every test or lint fix. -This step is complete when every selected case either passes or has a frozen-upstream differential proving the same platform-specific failure, every reconciliation-specific case passes, all four gates pass, every skip is named, and no test-owned process, selector, debug artifact, baseline worktree, or synthetic test worktree remains. +This step is complete when the four local gates pass or have valid cached evidence, the focused cases have explicit outcomes, known reconciliation defects are fixed, every deferred or unresolved subject has a named CI owner or reported coverage blocker, and no local test process remains. +Broad CI completion is not a prerequisite for opening the PR. ## 6. Commit the frozen snapshot @@ -283,25 +277,29 @@ test "$(git show -s --format='%(trailers:key=Firstmate-Upstream-SHA,valueonly)' For a reconstruction, connect the frozen fork ancestry before opening the PR when the fork SHA is not already an ancestor: ```sh -validated_tree=$(git rev-parse 'HEAD^{tree}') +reviewed_tree=$(git rev-parse 'HEAD^{tree}') git merge -s ours --no-ff <fork-sha> \ -m "Anchor frozen fork base ancestry" \ -m "Firstmate-Upstream-SHA: <upstream-sha>" -test "$validated_tree" = "$(git rev-parse 'HEAD^{tree}')" +test "$reviewed_tree" = "$(git rev-parse 'HEAD^{tree}')" ``` Give the ancestry-only merge a non-interactive message containing the frozen upstream trailer. -Record both tree SHAs as proof that the anchor changed no validated bytes. +Record both tree SHAs as proof that the anchor changed no reviewed bytes or local-check inputs. Create this anchor before the PR so GitHub can calculate mergeability and start checks immediately. This step is complete when the branch is clean, the reconciliation commit exposes the exact trailer through Git's trailer formatter, any ancestry anchor is tree-identical to its first parent, and the full diff contains only intentional reconciliation changes. -## 7. Open an ordinary pull request +## 7. Open the PR early and report CI separately Push the feature branch to `origin`. Use the authenticated GitHub client selected by the user's standing tooling preference for pull-request creation, history, mergeability, and check status. Read that client's live help before constructing an exact invocation. Open a non-draft pull request against the fork's `main` and leave it unmerged. +Open it after Step 5's fast local gate rather than waiting for broad regression or exhaustive platform triage. +Read repository merge settings and request a merge commit for this reconciliation, preserving upstream ancestry. +State that requirement in the PR body; do not squash or rebase-merge a reconciliation, merge it yourself, enable auto-merge, or change repository-wide merge settings. +If merge commits are disabled, report the policy constraint for the user instead of silently promising preserved ancestry. The PR body must record: @@ -311,6 +309,7 @@ The PR body must record: - normal merge or reconstruction; - meaningful conflict decisions and retained fork behavior; - every local command, result, duration, and skip; +- cache provenance, interrupted attempts, and each deferred/unresolved subject's CI owner; - the diff between the PR's Actions workflow/test runner and the frozen upstream versions; - any frozen-upstream differential failure and any ancestry-anchor tree proof; - the statement that GitHub Actions owns the complete cross-platform matrix. @@ -319,8 +318,14 @@ Verify that the PR reports the full URL, is mergeable, and has started the expec Poll the forge's structured PR fields until mergeability is no longer unknown and at least one expected check is visible. A reconstruction reported as conflicting while `origin/main` still equals the frozen fork SHA indicates that the ancestry anchor is missing or invalid. If fork `main` moves, fetch only `origin/main`, reconcile that new base into the same branch, and keep the upstream SHA frozen. -An ancestry-only base merge is acceptable only after proving its resulting tree is byte-identical to the already reconciled and validated tree. -After a base update, rerun changed validation against the new literal base SHA plus every repository gate. +An ancestry-only base merge is acceptable only after proving its resulting tree is byte-identical to the already reconciled tree. +After a base update, refresh the routing inventory and bounded local plan against the new literal base; the changed context invalidates cached results. + +Distinguish PR creation from merge readiness in every handoff. +Report expected CI checks as pending, successful, failed, or absent for the exact PR head; a skipped or cancelled required lane is not proof. +Missing checks, unresolved regressions, or missing platform coverage block merge readiness, not the existence of the ordinary PR. +Use the same named-case and budget policy when investigating CI failures; keep their evidence attached to the PR rather than restarting the local inventory. +Never call the reconciliation fully validated until required checks for that head are successful and unresolved failures have been explained or fixed. This skill is complete when the ordinary PR is open, non-draft, unmerged, mergeable or has a precisely reported blocker, and its live checks are reported. Synchronize local `main` only after the PR has been merged and that follow-up is explicitly requested. diff --git a/skills/reconcile-firstmate-upstream/scripts/validate-local.mjs b/skills/reconcile-firstmate-upstream/scripts/validate-local.mjs new file mode 100644 index 00000000000..dfa81f1973e --- /dev/null +++ b/skills/reconcile-firstmate-upstream/scripts/validate-local.mjs @@ -0,0 +1,275 @@ +#!/usr/bin/env node +// Bounded reconciliation checks; fm-test-run.sh still owns test execution and +// fm-timeout-lib.sh still owns process-group deadlines. +// +// node validate-local.mjs --plan <json> --state-dir <outside-repo-dir> +// [--root <checkout>] [--bash <recorded-bash>] [--budget-seconds 1200] +// [--max-timeouts 2] [--resume] [--preflight-only] +// +// Plan: { context: { fork: "<sha>", upstream: "<sha>" }, checks: [ +// { id: "runner-case", kind: "test"|"gate", command: ["bash", "..."], +// ci: "owning CI job", requires: [["node", "--version"]], +// env: { FM_TEST_ONLY: "test_name" }, timeoutSeconds: 180, +// cache: true, inputs: ["bin", "tests", "AGENTS.md"] } +// ], deferred: [{ id: "broad-regression", ci: "CI lanes", reason: "..." }] } +// +// Commands are argv arrays, never implicitly shell-evaluated. Inputs are +// explicit repository-relative files/directories; cache is opt-in and requires +// a reviewed complete dependency set. External-state and git-history checks +// must leave cache off. Cache identity includes context, controller/timeout +// bytes, input bytes/modes, command, environment digest, host and tool probes. +// Each completion is persisted atomically to result.json; only successful, +// non-skipped checks enter cache.json. Logs and cache stay outside the checkout. +// A state directory admits one invocation; a stale lock requires proof that its +// recorded owner is dead before removal. Resume never treats interrupted work +// as passed. Exit 0=all local checks passed, 1=failure, 2=invalid input, +// 75=incomplete/deferred/preflight-only. None of these means merge-ready. +import fs from "node:fs"; +import path from "node:path"; +import os from "node:os"; +import { createHash } from "node:crypto"; +import { spawn } from "node:child_process"; +import { fileURLToPath } from "node:url"; + +const self = fileURLToPath(import.meta.url); +const digest = (value) => createHash("sha256").update(value).digest("hex"); +const inside = (root, target) => { + const relative = path.relative(root, target); + return relative === "" || (!relative.startsWith(`..${path.sep}`) && + relative !== ".." && !path.isAbsolute(relative)); +}; +const positive = (value, name) => { + const number = Number(value); + if (!Number.isSafeInteger(number) || number <= 0) throw new Error(`${name} must be a positive integer`); + return number; +}; +const argv = (value, name) => { + if (!Array.isArray(value) || value.length === 0 || + value.some((part) => typeof part !== "string" || part.includes("\0")) || !value[0]) { + throw new Error(`${name} must be a nonempty argv array`); + } +}; +const readJson = (file) => JSON.parse(fs.readFileSync(file, "utf8").replace(/^\uFEFF/, "")); +function atomicJson(file, value) { + const temporary = `${file}.${process.pid}.tmp`; + fs.writeFileSync(temporary, `${JSON.stringify(value, null, 2)}\n`, { flag: "wx", mode: 0o600 }); + fs.renameSync(temporary, file); +} +function options(args) { + const result = { root: process.cwd(), bash: "bash", budget: 1200, maxTimeouts: 2 }; + for (let index = 0; index < args.length; index++) { + const argument = args[index]; + if (argument === "--resume") result.resume = true; + else if (argument === "--preflight-only") result.preflightOnly = true; + else { + const keys = { "--plan": "plan", "--state-dir": "stateDir", "--root": "root", + "--bash": "bash", "--budget-seconds": "budget", "--max-timeouts": "maxTimeouts" }; + if (!keys[argument] || !args[index + 1]) throw new Error(`unknown or incomplete option: ${argument}`); + result[keys[argument]] = args[++index]; + } + } + if (!result.plan || !result.stateDir) throw new Error("--plan and --state-dir are required"); + result.budget = positive(result.budget, "--budget-seconds"); + result.maxTimeouts = positive(result.maxTimeouts, "--max-timeouts"); + if (process.platform === "win32" && !path.isAbsolute(result.bash)) { + throw new Error("Windows requires --bash with the recorded absolute Git-for-Windows bash.exe path"); + } + return result; +} +function inputDigest(root, inputs) { + const rows = []; + function visit(file, ancestors) { + const real = fs.realpathSync(file); + if (!inside(root, real)) throw new Error(`input escapes repository: ${file}`); + if (ancestors.has(real)) throw new Error(`cyclic input: ${file}`); + const info = fs.lstatSync(file); + const name = path.relative(root, file); + if (info.isSymbolicLink()) { + rows.push([name, "link", fs.readlinkSync(file)]); + visit(real, new Set([...ancestors, file])); + } else if (info.isDirectory()) { + rows.push([name, "directory", info.mode & 0o777]); + for (const child of fs.readdirSync(file).sort()) visit(path.join(file, child), new Set([...ancestors, real])); + } else if (info.isFile()) rows.push([name, info.mode & 0o777, digest(fs.readFileSync(file))]); + else throw new Error(`unsupported input type: ${file}`); + } + for (const input of [...inputs].sort()) { + if (typeof input !== "string" || !input || path.isAbsolute(input)) throw new Error("inputs must be repository-relative paths"); + const file = path.resolve(root, input); + if (!inside(root, file) || file === root) throw new Error(`unbounded input: ${input}`); + visit(file, new Set()); + } + return digest(JSON.stringify(rows)); +} +function validatePlan(plan) { + if (!plan.context || !["fork", "upstream"].every((key) => /^[a-f0-9]{40}$/.test(plan.context[key]))) { + throw new Error("context must name literal 40-character fork and upstream SHAs"); + } + if (!Array.isArray(plan.checks) || plan.checks.length === 0) throw new Error("checks must be nonempty"); + const ids = new Set(); + for (const check of plan.checks) { + if (!/^[a-z0-9][a-z0-9-]*$/.test(check.id) || ids.has(check.id)) throw new Error(`invalid or duplicate check id: ${check.id}`); + ids.add(check.id); + if (!["test", "gate"].includes(check.kind) || typeof check.ci !== "string" || !check.ci.trim()) { + throw new Error(`${check.id}: kind and an owning CI job are required`); + } + argv(check.command, `${check.id}.command`); + if (!Array.isArray(check.requires)) throw new Error(`${check.id}: requires must explicitly list prerequisite probes`); + for (const probe of check.requires) argv(probe, `${check.id}.requires`); + if (check.timeoutSeconds !== undefined) positive(check.timeoutSeconds, `${check.id}.timeoutSeconds`); + if (check.cache !== undefined && typeof check.cache !== "boolean") throw new Error(`${check.id}: cache must be boolean`); + if (check.cache && (!Array.isArray(check.inputs) || check.inputs.length === 0)) throw new Error(`${check.id}: caching requires explicit dependency inputs`); + if (check.env !== undefined && (!check.env || Array.isArray(check.env) || typeof check.env !== "object" || + Object.entries(check.env).some(([key, value]) => !/^[A-Za-z_][A-Za-z0-9_]*$/.test(key) || typeof value !== "string" || value.includes("\0")))) { + throw new Error(`${check.id}: env must contain string environment assignments`); + } + } + for (const deferred of plan.deferred ?? []) { + if (!deferred.id || !deferred.ci || !deferred.reason) throw new Error("every deferred item needs id, ci, and reason"); + } +} +async function main() { + const config = options(process.argv.slice(2)); + const root = fs.realpathSync(config.root); + const stateDir = path.resolve(config.stateDir); + if (inside(root, stateDir)) throw new Error("state directory must be outside the checkout"); + fs.mkdirSync(stateDir, { recursive: true }); + if (inside(root, fs.realpathSync(stateDir))) throw new Error("state directory resolves inside the checkout"); + const plan = readJson(config.plan); + validatePlan(plan); + const timeoutLibrary = path.join(root, "bin", "fm-timeout-lib.sh"); + const ownerHash = digest(Buffer.concat([fs.readFileSync(self), fs.readFileSync(timeoutLibrary)])); + const lock = path.join(stateDir, "validation.lock"); + fs.writeFileSync(lock, JSON.stringify({ pid: process.pid, started: new Date().toISOString() }), { flag: "wx", mode: 0o600 }); + try { + const cacheFile = path.join(stateDir, "cache.json"); + const cache = fs.existsSync(cacheFile) ? readJson(cacheFile) : { schema: 1, passed: {} }; + if (cache.schema !== 1 || !cache.passed || typeof cache.passed !== "object" || Array.isArray(cache.passed)) { + throw new Error("invalid validation cache"); + } + const started = Date.now(); + const deadline = started + config.budget * 1000; + const logDirectory = path.join(stateDir, "logs", `${started}-${process.pid}`); + fs.mkdirSync(logDirectory, { recursive: true }); + const results = []; + const probes = new Map(); + let timeouts = 0; + const remaining = () => Math.floor((deadline - Date.now()) / 1000); + const snapshot = () => atomicJson(path.join(stateDir, "result.json"), { + schema: 1, context: plan.context, started: new Date(started).toISOString(), + updated: new Date().toISOString(), budgetSeconds: config.budget, + elapsedMs: Date.now() - started, timeouts, results, deferred: plan.deferred ?? [], + }); + async function bounded(command, env, seconds, logName) { + if (seconds < 1) return { exit: 124, output: "", elapsedMs: 0 }; + const shellPath = (value) => process.platform === "win32" + ? value.replace(/^([A-Za-z]):[\\/]/, (_, drive) => `/${drive.toLowerCase()}/`).replaceAll("\\", "/") + : value; + const logFile = path.join(logDirectory, logName); + const descriptor = fs.openSync(logFile, "w", 0o600); + const begin = Date.now(); + let output = ""; + let skipped = false; + let skipTail = "\n"; + try { + const child = spawn(config.bash, ["-c", + 'set +e; . "$1"; shift; command -v "$2" >/dev/null 2>&1 || { printf "prerequisite or command not found: %s\\n" "$2" >&2; exit 127; }; fm_run_timed "$@"', + "_", shellPath(timeoutLibrary), String(seconds), shellPath(command[0]), ...command.slice(1)], + { cwd: root, env, windowsHide: true, stdio: ["ignore", "pipe", "pipe"] }); + const collect = (chunk) => { + fs.writeSync(descriptor, chunk); + process.stdout.write(chunk); + const text = chunk.toString(); + const scan = skipTail + text; + skipped ||= /\nskip:/.test(scan); + skipTail = scan.slice(-5); + if (output.length < 65536) output += text.slice(0, 65536 - output.length); + }; + child.stdout.on("data", collect); + child.stderr.on("data", collect); + const exit = await new Promise((resolve, reject) => { + child.once("error", reject); + child.once("close", (code, signal) => resolve(code ?? (signal ? 130 : 1))); + }); + return { exit, output, skipped, elapsedMs: Date.now() - begin, log: logFile }; + } finally { + fs.closeSync(descriptor); + } + } + snapshot(); + for (const check of plan.checks) { + const base = { id: check.id, kind: check.kind, ci: check.ci, command: check.command }; + if (remaining() < 1 || timeouts >= config.maxTimeouts) { + results.push({ ...base, status: "deferred", reason: remaining() < 1 ? "local budget exhausted" : "timeout circuit breaker" }); + snapshot(); + console.log(`FM_RECONCILE_DEFER ${check.id} ci=${check.ci}`); + continue; + } + const env = { ...process.env, ...check.env }; + const envHash = digest(JSON.stringify(Object.entries(env).sort(([a], [b]) => a.localeCompare(b)))); + const versions = []; + let missing = false; + for (let index = 0; index < check.requires.length; index++) { + const probe = check.requires[index]; + const key = digest(JSON.stringify([probe, envHash])); + if (!probes.has(key)) { + const result = await bounded(probe, env, Math.min(30, remaining()), `${check.id}.preflight-${index}.log`); + probes.set(key, result); + if (result.exit === 124) timeouts++; + } + const result = probes.get(key); + versions.push([probe, result.exit, digest(result.output)]); + if (result.exit !== 0) { + results.push({ ...base, status: "deferred", reason: "prerequisite probe failed", prerequisite: probe, exit: result.exit, log: result.log }); + console.log(`FM_RECONCILE_DEFER ${check.id} prerequisite=${JSON.stringify(probe)} ci=${check.ci}`); + missing = true; + break; + } + } + if (missing || config.preflightOnly) { + if (!missing) results.push({ ...base, status: "preflight-only" }); + snapshot(); + continue; + } + const fingerprint = digest(JSON.stringify({ context: plan.context, ownerHash, + host: [process.platform, process.arch, os.release(), process.version, config.bash], + command: check.command, envHash, versions, + inputs: check.cache ? inputDigest(root, check.inputs) : null })); + if (config.resume && check.cache && cache.passed[fingerprint]) { + results.push({ ...base, status: "cached", fingerprint, reused: cache.passed[fingerprint] }); + console.log(`FM_RECONCILE_CACHED ${check.id}`); + snapshot(); + continue; + } + console.log(`FM_RECONCILE_BEGIN ${check.id}`); + const result = await bounded(check.command, env, + Math.min(check.timeoutSeconds ?? config.budget, remaining()), `${check.id}.log`); + const skipped = result.skipped; + const status = result.exit === 124 ? "timeout" : result.exit !== 0 ? "failed" : skipped ? "deferred" : "passed"; + if (status === "timeout") timeouts++; + results.push({ ...base, status, fingerprint, exit: result.exit, elapsedMs: result.elapsedMs, log: result.log, + ...(skipped ? { reason: "command reported an optional-tool or live-policy skip" } : {}) }); + if (check.cache && status === "passed") { + cache.passed[fingerprint] = { id: check.id, completed: new Date().toISOString(), log: result.log }; + atomicJson(cacheFile, cache); + } + snapshot(); + console.log(`FM_RECONCILE_END ${check.id} status=${status} duration_ms=${result.elapsedMs}`); + } + if (results.some((result) => result.status === "failed")) return 1; + if (results.some((result) => !["passed", "cached"].includes(result.status))) return 75; + return 0; + } finally { + fs.unlinkSync(lock); + } +} +try { + if (process.argv.slice(2).length === 1 && process.argv[2] === "--help") { + const lines = fs.readFileSync(self, "utf8").split("\n"); + console.log(lines.slice(1, lines.findIndex((line) => line.startsWith("import "))) + .map((line) => line.replace(/^\/\/ ?/, "")).join("\n")); + } else process.exitCode = await main(); +} catch (error) { + console.error(`fm-reconcile-validation: ${error.message}`); + process.exitCode = 2; +} diff --git a/tests/fm-afk-launch.test.sh b/tests/fm-afk-launch.test.sh index 6d0c7bd9d1a..675805b6a1c 100755 --- a/tests/fm-afk-launch.test.sh +++ b/tests/fm-afk-launch.test.sh @@ -135,6 +135,7 @@ unit_fresh_vs_refresh() { lock="$st/state/.supervise-daemon.lock" mkdir -p "$lock" printf '%s' "$sleep_pid" > "$lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$sleep_pid" > "$lock/pid-identity" 2>/dev/null ) || true FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$START" >/dev/null 2>&1 if [ -e "$st/state/.subsuper-escalations" ] && [ -e "$st/state/.subsuper-inject-wedged" ]; then @@ -162,10 +163,12 @@ unit_stop_ordering() { trap "if [ -f \"$1/state/.afk\" ]; then echo present > \"$2\"; else echo absent > \"$2\"; fi; exit 0" TERM while :; do sleep 0.2; done ' _ "$st" "$marker" & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. daemon_pid=$! lock="$st/state/.supervise-daemon.lock" mkdir -p "$lock" printf '%s' "$daemon_pid" > "$lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$daemon_pid" > "$lock/pid-identity" 2>/dev/null ) || true printf 'none\t-\tnative\n' > "$st/state/.afk-daemon-terminal" FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$LAUNCH" stop >/dev/null 2>&1 @@ -196,6 +199,7 @@ unit_stop_rejects_reused_pid() { mkdir -p "$st/state" date '+%s' > "$st/state/.afk" sleep 600 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. sleeper_pid=$! lock="$st/state/.supervise-daemon.lock" mkdir -p "$lock" @@ -240,9 +244,13 @@ unit_concurrent_start_serialized() { TRACK_TMUX_SESSIONS="$TRACK_TMUX_SESSIONS $cap_session" cap_pane=$(tmux display-message -p -t "$cap_session" '#{pane_id}') FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_SUPERVISOR_TARGET="$cap_pane" \ - FM_SUPERVISOR_BACKEND=tmux FM_AFK_LAUNCH_ENTRY="$SLEEPER" "$LAUNCH" start >/dev/null 2>&1 & first=$! + FM_SUPERVISOR_BACKEND=tmux FM_AFK_LAUNCH_ENTRY="$SLEEPER" "$LAUNCH" start >/dev/null 2>&1 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. + first=$! FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_SUPERVISOR_TARGET="$cap_pane" \ - FM_SUPERVISOR_BACKEND=tmux FM_AFK_LAUNCH_ENTRY="$SLEEPER" "$LAUNCH" start >/dev/null 2>&1 & second=$! + FM_SUPERVISOR_BACKEND=tmux FM_AFK_LAUNCH_ENTRY="$SLEEPER" "$LAUNCH" start >/dev/null 2>&1 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. + second=$! wait "$first"; wait "$second" rec=$(cut -f2 "$st/state/.afk-daemon-terminal" 2>/dev/null || true) count=$(tmux list-sessions -F '#{session_name}' 2>/dev/null | awk -v expected="$rec" '$0 == expected {n++} END{print n+0}') @@ -266,6 +274,7 @@ unit_lock_initialization_grace() { sleep 0.15 if [ -d "$st/state/.afk-launch.lock" ]; then printf '%s' "$$" > "$st/state/.afk-launch.lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$$" > "$st/state/.afk-launch.lock/pid-identity" 2>/dev/null ) || true # shellcheck disable=SC2031 # The subshell writes the path value; it does not reassign the variable. : > "$marker" @@ -273,6 +282,7 @@ unit_lock_initialization_grace() { rm -rf "$st/state/.afk-launch.lock" fi ) & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. initializer=$! # shellcheck disable=SC2031 # The initializer communicates through this shared file path. if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" bash -c ' @@ -298,6 +308,7 @@ unit_signal_exits_with_lock_cleanup() { fm_afk_launch_main start : > "$2" ' _ "$LAUNCH" "$marker" & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. child=$! # Signal only once the lifecycle actually holds its lock. Killing before the # lock exists tests nothing, and on a loaded machine it used to race: the @@ -651,9 +662,12 @@ unit_stop_validates_before_signal() { mkdir -p "$st/state" : > "$st/state/.afk" printf 'tmux\tonly-two-fields\n' > "$st/state/.afk-daemon-terminal" - sleep 30 & sleeper_pid=$! + sleep 30 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. + sleeper_pid=$! mkdir -p "$st/state/.supervise-daemon.lock" printf '%s' "$sleeper_pid" > "$st/state/.supervise-daemon.lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$sleeper_pid" > "$st/state/.supervise-daemon.lock/pid-identity" ) FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" "$LAUNCH" stop >/dev/null 2>&1 || true if kill -0 "$sleeper_pid" 2>/dev/null && [ -e "$st/state/.afk" ]; then @@ -706,8 +720,10 @@ unit_stop_confirms_daemon_exit() { : > "$st/state/.afk" printf 'none\t-\tnative\n' > "$st/state/.afk-daemon-terminal" bash -c 'trap "" TERM; while :; do sleep 1; done' & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. daemon_pid=$! printf '%s' "$daemon_pid" > "$st/state/.supervise-daemon.lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$daemon_pid" > "$st/state/.supervise-daemon.lock/pid-identity" ) if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" bash -c ' . "$1" @@ -737,8 +753,11 @@ unit_refresh_validates_record() { st=$(mktemp -d "${TMPDIR:-/tmp}/fm-afk-refresh-record.XXXXXX") mkdir -p "$st/state/.supervise-daemon.lock" printf 'tmux\tonly-two-fields\n' > "$st/state/.afk-daemon-terminal" - sleep 30 & daemon_pid=$! + sleep 30 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. + daemon_pid=$! printf '%s' "$daemon_pid" > "$st/state/.supervise-daemon.lock/pid" + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_pid_identity "$daemon_pid" > "$st/state/.supervise-daemon.lock/pid-identity" ) if FM_HOME="$st" FM_STATE_OVERRIDE="$st/state" FM_SUPERVISOR_TARGET=unused \ FM_SUPERVISOR_BACKEND=tmux bash -c ' diff --git a/tests/fm-afk-pi-herdr-return-e2e.test.sh b/tests/fm-afk-pi-herdr-return-e2e.test.sh index 98c826085f7..7c2793dc530 100755 --- a/tests/fm-afk-pi-herdr-return-e2e.test.sh +++ b/tests/fm-afk-pi-herdr-return-e2e.test.sh @@ -2,8 +2,11 @@ # Real Pi/Herdr end-to-end regression for the 2026-07-14 two-owner incident. # # Opt-in because it launches a real interactive Pi primary, a real away daemon, -# and a real isolated Herdr lab session. Every explicit and production-adapter -# Herdr call is routed through fm-herdr-lab.sh. The scenario proves: +# and a real isolated Herdr lab session. Real model turns, real side effects, +# and heavyweight lab setup put it outside the token-free default-on class, so +# it does not run merely because its tools are installed. Every explicit and +# production-adapter Herdr call is routed through fm-herdr-lab.sh. The scenario +# proves: # - a live blocked status is classified and durably queued while away; # - a pending Pi composer refuses injection and receives no forced Enter; # - the existing wedge alarm remains observable and deduped; @@ -15,20 +18,14 @@ set -u # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_AFK_PI_HERDR_E2E herdr jq pi python3 + # shellcheck source=/dev/null . "$ROOT/bin/fm-supervise-daemon.sh" # shellcheck source=/dev/null . "$ROOT/bin/fm-backend.sh" -if [ "${FM_AFK_PI_HERDR_E2E:-0}" != 1 ]; then - echo "skip: set FM_AFK_PI_HERDR_E2E=1 to run the real Pi/Herdr away-return regression" - exit 0 -fi - -for tool in herdr jq pi python3; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done - LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} SESSION=$("$LAB_HELPER" name fm-afk-pi-return-e2e) TMP_ROOT=$(fm_test_tmproot fm-afk-pi-return-e2e) diff --git a/tests/fm-backend-herdr-presentation-e2e.test.sh b/tests/fm-backend-herdr-presentation-e2e.test.sh index b98176ff2ec..3491071f080 100755 --- a/tests/fm-backend-herdr-presentation-e2e.test.sh +++ b/tests/fm-backend-herdr-presentation-e2e.test.sh @@ -24,6 +24,7 @@ TMP_ROOT=$(mktemp -d "$(cd "${TMPDIR:-/tmp}" && pwd -P)/fm-herdr-presentation.XX FAKEBIN="$TMP_ROOT/fakebin" HERDR_CALL_LOG="$TMP_ROOT/herdr-calls.log" TREEHOUSE_CALL_LOG="$TMP_ROOT/treehouse-calls.log" +TREEHOUSE_LOCK_DIR="$TMP_ROOT/treehouse-call.lock" MOVE_CALL_LOG="$TMP_ROOT/workspace-move-calls.log" FOCUS_AUDIT_LOG="$TMP_ROOT/focus-audit.log" ACTIVE_SEEDED_CONTROL="$TMP_ROOT/active-seeded-control" @@ -34,7 +35,7 @@ mkdir -p "$FAKEBIN" : > "$MOVE_CALL_LOG" : > "$FOCUS_AUDIT_LOG" REAL_MOVER="$ROOT/bin/backends/herdr-workspace-move.py" -export REAL_HERDR REAL_TREEHOUSE REAL_MOVER HERDR_CALL_LOG TREEHOUSE_CALL_LOG MOVE_CALL_LOG FOCUS_AUDIT_LOG HERDR_ORIGINAL_PATH HERDR_LAB_HELPER +export REAL_HERDR REAL_TREEHOUSE REAL_MOVER HERDR_CALL_LOG TREEHOUSE_CALL_LOG TREEHOUSE_LOCK_DIR MOVE_CALL_LOG FOCUS_AUDIT_LOG HERDR_ORIGINAL_PATH HERDR_LAB_HELPER export ACTIVE_SEEDED_CONTROL POST_CREATE_ABORT_CONTROL TMP_ROOT # Log every production-adapter call, remove its already-validated trailing @@ -211,7 +212,17 @@ set -u if [ -d "$POST_CREATE_ABORT_CONTROL" ] && [ "${1:-}" = get ]; then exit 0 fi -exec "$REAL_TREEHOUSE" "$@" +# Treehouse's pool allocator is outside the Herdr concurrency contract under +# test. Serialize its calls so simultaneous recovery spawns cannot race for +# one pool slot before reaching the Herdr session lock exercised below. +while ! mkdir "$TREEHOUSE_LOCK_DIR" 2>/dev/null; do + sleep 0.01 +done +release_treehouse_lock() { rmdir "$TREEHOUSE_LOCK_DIR" 2>/dev/null || true; } +trap release_treehouse_lock EXIT +trap 'exit 1' HUP INT TERM +"$REAL_TREEHOUSE" "$@" +exit $? SH cat > "$FAKEBIN/herdr-workspace-mover" <<'SH' @@ -425,7 +436,7 @@ EOF spawn_task() { # <id> <home> <project> local id=$1 home=$2 project=$3 FM_GATE_REFUSE_BYPASS=1 FM_SPAWN_NO_GUARD=1 FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ - "$ROOT/bin/fm-spawn.sh" "$id" "$project" "sh -c 'sleep 120'" --mode no-mistakes --yolo off --backend herdr + "$ROOT/bin/fm-spawn.sh" "$id" "$project" "sh -c 'while :; do sleep 60; done'" --mode no-mistakes --yolo off --backend herdr } finish_concurrent_spawn() { # <id> <status> <stdout> <stderr> @@ -450,7 +461,7 @@ finish_concurrent_expected_abort() { # <id> <status> <stdout> <stderr> spawn_secondmate_task() { local id=$1 home=$2 FM_GATE_REFUSE_BYPASS=1 FM_SPAWN_NO_GUARD=1 FM_HOME="$HOME_DIR" FM_ROOT_OVERRIDE="$ROOT" \ - "$ROOT/bin/fm-spawn.sh" "$id" "$home" "sh -c 'sleep 120'" --secondmate --backend herdr + "$ROOT/bin/fm-spawn.sh" "$id" "$home" "sh -c 'while :; do sleep 60; done'" --secondmate --backend herdr } teardown_task() { # <id> <home> @@ -464,8 +475,10 @@ teardown_task() { # <id> <home> finish_concurrent_teardown() { # <id> <status> <stdout> <stderr> local id=$1 status=$2 out=$3 err=$4 [ "$status" -ne 0 ] || return 0 - grep -F "session presentation lock is contended" "$err" >/dev/null 2>&1 \ - || fail "projected teardown $id failed unexpectedly: $(cat "$err")" + if ! grep -F "session presentation lock is contended" "$err" >/dev/null 2>&1 \ + && ! grep -F "another Treehouse slot allocation or return is in progress" "$err" >/dev/null 2>&1; then + fail "projected teardown $id failed unexpectedly: $(cat "$err")" + fi teardown_task "$id" "$HOME_DIR" > "$out" 2> "$err" \ || fail "projected teardown $id retry failed after presentation cleanup completed: $(cat "$err")" } @@ -542,6 +555,7 @@ assert_no_projection_mutation_since() { # <line-count> <case-name> HOME_DIR="$TMP_ROOT/home" PROJECT_DIR="$TMP_ROOT/project" +RECOVERY_PROJECT_DIR="$TMP_ROOT/recovery-project" mkdir -p "$HOME_DIR/state" "$HOME_DIR/config" \ "$HOME_DIR/data/anchor" "$HOME_DIR/data/shape" \ "$HOME_DIR/data/order-a" "$HOME_DIR/data/order-b" \ @@ -569,6 +583,7 @@ write_ship_brief "$HOME_DIR" default-on 'Projection default-on fixture.' write_ship_brief "$HOME_DIR" active-cleanup-safe 'Projection active cleanup safe-focus fixture.' write_ship_brief "$HOME_DIR" active-cleanup-deferred 'Projection active cleanup deferred fixture.' make_project "$PROJECT_DIR" +make_project "$RECOVERY_PROJECT_DIR" # Keep one ordinary primary task live so the durable firstmate workspace is # first and remains present while disposable workers are projected around it. @@ -1052,9 +1067,12 @@ if wait "$ABORT_A_PID"; then ABORT_A_STATUS=0; else ABORT_A_STATUS=$?; fi if wait "$ABORT_B_PID"; then ABORT_B_STATUS=0; else ABORT_B_STATUS=$?; fi finish_concurrent_expected_abort abort-a "$ABORT_A_STATUS" "$TMP_ROOT/abort-a.out" "$TMP_ROOT/abort-a.err" finish_concurrent_expected_abort abort-b "$ABORT_B_STATUS" "$TMP_ROOT/abort-b.out" "$TMP_ROOT/abort-b.err" -grep -F "did not yield an isolated worktree" "$TMP_ROOT/abort-a.err" >/dev/null 2>&1 \ +# The forced foreground_cwd is a plain non-git directory, which the discovery +# poll now screens out on every read rather than adopting, so the armed failure +# arrives as the poll's own deadline refusal naming that path. +grep -F "did not enter an isolated worktree" "$TMP_ROOT/abort-a.err" >/dev/null 2>&1 \ || fail "post-create abort fixture A did not reach the armed validation failure: $(cat "$TMP_ROOT/abort-a.err")" -grep -F "did not yield an isolated worktree" "$TMP_ROOT/abort-b.err" >/dev/null 2>&1 \ +grep -F "did not enter an isolated worktree" "$TMP_ROOT/abort-b.err" >/dev/null 2>&1 \ || fail "post-create abort fixture B did not reach the armed validation failure: $(cat "$TMP_ROOT/abort-b.err")" ABORT_A_PANE=$(cat "$POST_CREATE_ABORT_CONTROL/abort-a/task-pane") ABORT_B_PANE=$(cat "$POST_CREATE_ABORT_CONTROL/abort-b/task-pane") @@ -1371,11 +1389,15 @@ teardown_task aflat "$SECOND_HOME_A" > "$TMP_ROOT/aflat-teardown.out" 2> "$TMP_R pass "real Herdr lab: session lock contention from a secondmate home falls back flat with no journal" # Same-identity recovery replaces only one exact agent-free husk in its -# original projected workspace. +# original projected workspace. These full-session restarts also stop the +# earlier multi-home workers whose restored panes are retained for the final +# exact-pane cleanup assertions. Keep the recovery fixtures in their own +# Treehouse pool so those intentionally retained records cannot claim a slot +# that a recovery fixture legitimately acquires after their processes stop. # Exercise both the leading fm- identity style seen in Hi Bit work and the # project-name identity style used by Wheelhouse work. for RESTART_ID in fm-hibit-resume-r1 wheelhouse-healing-r1; do - spawn_task "$RESTART_ID" "$HOME_DIR" "$PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-first.out" 2> "$TMP_ROOT/$RESTART_ID-first.err" \ + spawn_task "$RESTART_ID" "$HOME_DIR" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-first.out" 2> "$TMP_ROOT/$RESTART_ID-first.err" \ || fail "$RESTART_ID fixture's projected spawn failed: $(cat "$TMP_ROOT/$RESTART_ID-first.err")" RESTART_META="$HOME_DIR/state/$RESTART_ID.meta" OLD_RESTART_WT=$(remember_meta_worktree "$RESTART_META") @@ -1395,13 +1417,17 @@ for RESTART_ID in fm-hibit-resume-r1 wheelhouse-healing-r1; do PATH="$HERDR_ORIGINAL_PATH" \ "$HERDR_LAB_HELPER" provision "$HERDR_LAB_SESSION" \ || fail "could not reprovision the isolated session for $RESTART_ID validation" + # Stopping the whole Herdr session also ends the anchor's agent. Its restored + # shell remains useful as the durable layout anchor, but its task record no + # longer represents a live slot owner and must not poison later slot reuse. + rm -f "$ANCHOR_META" lab pane get "$OLD_RESTART_PANE" >/dev/null 2>&1 \ || fail "$RESTART_ID restart did not preserve the projected pane structurally" if lab agent get "$OLD_RESTART_PANE" >/dev/null 2>&1; then fail "$RESTART_ID restart fixture unexpectedly retained a registered agent" fi RECLAIM_FOCUS=$(focus_snapshot) - spawn_task "$RESTART_ID" "$HOME_DIR" "$PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-reclaim.out" 2> "$TMP_ROOT/$RESTART_ID-reclaim.err" \ + spawn_task "$RESTART_ID" "$HOME_DIR" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-reclaim.out" 2> "$TMP_ROOT/$RESTART_ID-reclaim.err" \ || fail "$RESTART_ID same-identity reclaim failed: $(cat "$TMP_ROOT/$RESTART_ID-reclaim.err")" NEW_RESTART_WT=$(remember_meta_worktree "$RESTART_META") NEW_RESTART_WSID=$(grep '^herdr_workspace_id=' "$RESTART_META" | cut -d= -f2-) @@ -1426,7 +1452,7 @@ for RESTART_ID in fm-hibit-resume-r1 wheelhouse-healing-r1; do || fail "could not reprovision the isolated session for idempotent reclaim" PRIOR_RESTART_WT=$NEW_RESTART_WT PRIOR_RESTART_PANE=$NEW_RESTART_PANE - spawn_task "$RESTART_ID" "$HOME_DIR" "$PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-idempotent.out" 2> "$TMP_ROOT/$RESTART_ID-idempotent.err" \ + spawn_task "$RESTART_ID" "$HOME_DIR" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/$RESTART_ID-idempotent.out" 2> "$TMP_ROOT/$RESTART_ID-idempotent.err" \ || fail "$RESTART_ID repeated reclaim failed: $(cat "$TMP_ROOT/$RESTART_ID-idempotent.err")" NEW_RESTART_WT=$(remember_meta_worktree "$RESTART_META") NEW_RESTART_WSID=$(grep '^herdr_workspace_id=' "$RESTART_META" | cut -d= -f2-) @@ -1435,7 +1461,9 @@ for RESTART_ID in fm-hibit-resume-r1 wheelhouse-healing-r1; do || fail "$RESTART_ID repeated reclaim changed workspace identity" [ "$NEW_RESTART_PANE" != "$PRIOR_RESTART_PANE" ] \ || fail "$RESTART_ID repeated reclaim reused the prior husk pane" - "$REAL_TREEHOUSE" return --force "$PRIOR_RESTART_WT" >/dev/null 2>&1 || true + if [ "$PRIOR_RESTART_WT" != "$NEW_RESTART_WT" ]; then + "$REAL_TREEHOUSE" return --force "$PRIOR_RESTART_WT" >/dev/null 2>&1 || true + fi fi teardown_task "$RESTART_ID" "$HOME_DIR" > "$TMP_ROOT/$RESTART_ID-teardown.out" 2> "$TMP_ROOT/$RESTART_ID-teardown.err" \ @@ -1451,7 +1479,7 @@ pass "real Herdr lab: Hi Bit and Wheelhouse-style same-identity restarts reclaim CROSS_RESTART_ID=wheel-child-resume mkdir -p "$SECOND_HOME_A/data/$CROSS_RESTART_ID" write_ship_brief "$SECOND_HOME_A" "$CROSS_RESTART_ID" 'Cross-home restart fixture.' -spawn_task "$CROSS_RESTART_ID" "$SECOND_HOME_A" "$PROJECT_DIR" > "$TMP_ROOT/cross-restart-first.out" 2> "$TMP_ROOT/cross-restart-first.err" \ +spawn_task "$CROSS_RESTART_ID" "$SECOND_HOME_A" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/cross-restart-first.out" 2> "$TMP_ROOT/cross-restart-first.err" \ || fail "cross-home restart fixture failed: $(cat "$TMP_ROOT/cross-restart-first.err")" CROSS_RESTART_META="$SECOND_HOME_A/state/$CROSS_RESTART_ID.meta" CROSS_OLD_WT=$(remember_meta_worktree "$CROSS_RESTART_META") @@ -1467,7 +1495,7 @@ PATH="$HERDR_ORIGINAL_PATH" "$HERDR_LAB_HELPER" stop "$HERDR_LAB_SESSION" >/dev/ || fail "could not stop the isolated session for cross-home restart" PATH="$HERDR_ORIGINAL_PATH" "$HERDR_LAB_HELPER" provision "$HERDR_LAB_SESSION" \ || fail "could not reprovision the isolated session for cross-home restart" -spawn_task "$CROSS_RESTART_ID" "$SECOND_HOME_A" "$PROJECT_DIR" > "$TMP_ROOT/cross-restart-resume.out" 2> "$TMP_ROOT/cross-restart-resume.err" \ +spawn_task "$CROSS_RESTART_ID" "$SECOND_HOME_A" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/cross-restart-resume.out" 2> "$TMP_ROOT/cross-restart-resume.err" \ || fail "cross-home same-identity reclaim failed: $(cat "$TMP_ROOT/cross-restart-resume.err")" CROSS_NEW_WT=$(remember_meta_worktree "$CROSS_RESTART_META") CROSS_NEW_WSID=$(grep '^herdr_workspace_id=' "$CROSS_RESTART_META" | cut -d= -f2-) @@ -1489,9 +1517,9 @@ BRAVO_WAVE_ID=resume-wave-bravo mkdir -p "$HOME_DIR/data/$PRIMARY_WAVE_ID" "$SECOND_HOME_B/data/$BRAVO_WAVE_ID" write_ship_brief "$HOME_DIR" "$PRIMARY_WAVE_ID" 'Concurrent primary recovery fixture.' write_ship_brief "$SECOND_HOME_B" "$BRAVO_WAVE_ID" 'Concurrent secondmate recovery fixture.' -spawn_task "$PRIMARY_WAVE_ID" "$HOME_DIR" "$PROJECT_DIR" > "$TMP_ROOT/primary-wave-first.out" 2> "$TMP_ROOT/primary-wave-first.err" \ +spawn_task "$PRIMARY_WAVE_ID" "$HOME_DIR" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/primary-wave-first.out" 2> "$TMP_ROOT/primary-wave-first.err" \ || fail "primary recovery-wave fixture failed: $(cat "$TMP_ROOT/primary-wave-first.err")" -spawn_task "$BRAVO_WAVE_ID" "$SECOND_HOME_B" "$PROJECT_DIR" > "$TMP_ROOT/bravo-wave-first.out" 2> "$TMP_ROOT/bravo-wave-first.err" \ +spawn_task "$BRAVO_WAVE_ID" "$SECOND_HOME_B" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/bravo-wave-first.out" 2> "$TMP_ROOT/bravo-wave-first.err" \ || fail "secondmate recovery-wave fixture failed: $(cat "$TMP_ROOT/bravo-wave-first.err")" PRIMARY_WAVE_META="$HOME_DIR/state/$PRIMARY_WAVE_ID.meta" BRAVO_WAVE_META="$SECOND_HOME_B/state/$BRAVO_WAVE_ID.meta" @@ -1506,9 +1534,9 @@ PATH="$HERDR_ORIGINAL_PATH" "$HERDR_LAB_HELPER" stop "$HERDR_LAB_SESSION" >/dev/ PATH="$HERDR_ORIGINAL_PATH" "$HERDR_LAB_HELPER" provision "$HERDR_LAB_SESSION" \ || fail "could not reprovision the isolated session for concurrent recovery" CONCURRENT_RECOVERY_FOCUS=$(focus_snapshot) -spawn_task "$PRIMARY_WAVE_ID" "$HOME_DIR" "$PROJECT_DIR" > "$TMP_ROOT/primary-wave-resume.out" 2> "$TMP_ROOT/primary-wave-resume.err" & +spawn_task "$PRIMARY_WAVE_ID" "$HOME_DIR" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/primary-wave-resume.out" 2> "$TMP_ROOT/primary-wave-resume.err" & PRIMARY_WAVE_PID=$! -spawn_task "$BRAVO_WAVE_ID" "$SECOND_HOME_B" "$PROJECT_DIR" > "$TMP_ROOT/bravo-wave-resume.out" 2> "$TMP_ROOT/bravo-wave-resume.err" & +spawn_task "$BRAVO_WAVE_ID" "$SECOND_HOME_B" "$RECOVERY_PROJECT_DIR" > "$TMP_ROOT/bravo-wave-resume.out" 2> "$TMP_ROOT/bravo-wave-resume.err" & BRAVO_WAVE_PID=$! wait "$PRIMARY_WAVE_PID" || fail "concurrent primary recovery failed: $(cat "$TMP_ROOT/primary-wave-resume.err")" wait "$BRAVO_WAVE_PID" || fail "concurrent secondmate recovery failed: $(cat "$TMP_ROOT/bravo-wave-resume.err")" @@ -1528,9 +1556,9 @@ if lab pane get "$PRIMARY_WAVE_OLD_PANE" >/dev/null 2>&1 \ fi assert_focus_is "$CONCURRENT_RECOVERY_FOCUS" "concurrent cross-home recovery" teardown_task "$PRIMARY_WAVE_ID" "$HOME_DIR" > "$TMP_ROOT/primary-wave-teardown.out" 2> "$TMP_ROOT/primary-wave-teardown.err" \ - || fail "concurrent primary recovery teardown failed" + || fail "concurrent primary recovery teardown failed: $(cat "$TMP_ROOT/primary-wave-teardown.err")" teardown_task "$BRAVO_WAVE_ID" "$SECOND_HOME_B" > "$TMP_ROOT/bravo-wave-teardown.out" 2> "$TMP_ROOT/bravo-wave-teardown.err" \ - || fail "concurrent secondmate recovery teardown failed" + || fail "concurrent secondmate recovery teardown failed: $(cat "$TMP_ROOT/bravo-wave-teardown.err")" "$REAL_TREEHOUSE" return --force "$PRIMARY_WAVE_OLD_WT" >/dev/null 2>&1 || true "$REAL_TREEHOUSE" return --force "$BRAVO_WAVE_OLD_WT" >/dev/null 2>&1 || true "$REAL_TREEHOUSE" return --force "$PRIMARY_WAVE_NEW_WT" >/dev/null 2>&1 || true diff --git a/tests/fm-backend-herdr-treehouse.test.sh b/tests/fm-backend-herdr-treehouse.test.sh index 9bb7a419072..7ed27a682f0 100755 --- a/tests/fm-backend-herdr-treehouse.test.sh +++ b/tests/fm-backend-herdr-treehouse.test.sh @@ -110,12 +110,11 @@ test_unpublished_lease_return_failure_is_visible() { pass "an unpublished lease return failure remains visible to spawn recovery" } -test_acquisition_mode_leases_only_on_native_windows -test_enter_worktree_command_quotes_literal_path -test_environment_command_quotes_literal_value -test_bash_script_command_quotes_literal_paths -test_acquire_lease_uses_project_and_holder -test_unpublished_lease_closes_endpoint_before_return -test_unpublished_lease_return_failure_is_visible - -echo "# all native Windows Herdr Treehouse tests passed" +fm_test_run_cases \ + test_acquisition_mode_leases_only_on_native_windows \ + test_enter_worktree_command_quotes_literal_path \ + test_environment_command_quotes_literal_value \ + test_bash_script_command_quotes_literal_paths \ + test_acquire_lease_uses_project_and_holder \ + test_unpublished_lease_closes_endpoint_before_return \ + test_unpublished_lease_return_failure_is_visible diff --git a/tests/fm-backend-herdr-windows-treehouse-live-e2e.test.sh b/tests/fm-backend-herdr-windows-treehouse-live-e2e.test.sh index f1bf41782cb..fa37fb808a3 100755 --- a/tests/fm-backend-herdr-windows-treehouse-live-e2e.test.sh +++ b/tests/fm-backend-herdr-windows-treehouse-live-e2e.test.sh @@ -62,7 +62,7 @@ git -C "$PROJECT" -c user.name='Firstmate Tests' -c user.email='tests@example.in git clone --quiet --bare "$PROJECT" "$TMP_ROOT/origin.git" git -C "$PROJECT" remote add origin "file://$TMP_ROOT/origin.git" -LAUNCH_COMMAND="FM_LAUNCH_VALUE=bridge-ok sh -c 'printf \"%s|%s\" \"\$FM_LAUNCH_VALUE\" \"\$GOTMPDIR\" > \"$LAUNCH_PROOF\"; sleep 120'" +LAUNCH_COMMAND="FM_LAUNCH_VALUE=bridge-ok sh -c 'printf \"%s|%s|%s\" \"\$FM_LAUNCH_VALUE\" \"\$GOTMPDIR\" \"\$FM_TASK_ID\" > \"$LAUNCH_PROOF\"; sleep 120'" FM_GATE_REFUSE_BYPASS=1 FM_SPAWN_NO_GUARD=1 FM_HOME="$HOME_DIR" FM_ROOT_OVERRIDE="$ROOT" \ HERDR_SESSION="$LAB" "$ROOT/bin/fm-spawn.sh" "$ID" "$PROJECT" \ "$LAUNCH_COMMAND" --mode no-mistakes --yolo off --backend herdr @@ -73,7 +73,7 @@ for _ in $(seq 1 100); do sleep 0.1 done LAUNCH_RESULT=$(cat "$LAUNCH_PROOF" 2>/dev/null || true) -if [ "$LAUNCH_RESULT" != "bridge-ok|/tmp/fm-$ID/gotmp" ]; then +if [ "$LAUNCH_RESULT" != "bridge-ok|/tmp/fm-$ID/gotmp|$ID" ]; then PANE=$(grep '^herdr_pane_id=' "$META" | cut -d= -f2-) echo "launch proof mismatch: got '$LAUNCH_RESULT'" >&2 "$HELPER" run "$LAB" pane read "$PANE" --source recent --lines 80 >&2 || true diff --git a/tests/fm-backend-orca.test.sh b/tests/fm-backend-orca.test.sh index 119c0e60f4f..60564719239 100755 --- a/tests/fm-backend-orca.test.sh +++ b/tests/fm-backend-orca.test.sh @@ -539,7 +539,7 @@ test_spawn_writes_orca_metadata_and_launches_harness() { "spawn should reuse the implicit terminal returned by Orca worktree creation" assert_contains "$(cat "$log")" $'orca\x1f''terminal'$'\x1f''send'$'\x1f''--terminal'$'\x1f''term-spawn'$'\x1f''--text'$'\x1f''export GOTMPDIR=/tmp/fm-orcaspawnz1/gotmp'$'\x1f''--enter'$'\x1f''--json' \ "spawn did not export GOTMPDIR through the Orca terminal" - assert_contains "$(cat "$log")" "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}'" \ + assert_contains "$(cat "$log")" "CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}'" \ "spawn did not send the selected harness launch command through Orca" rm -rf "/tmp/fm-$id" pass "fm-spawn.sh --backend orca: reuses implicit terminal, records metadata, launches harness" @@ -769,7 +769,7 @@ test_peek_send_and_crew_state_route_through_orca_meta() { [ "$body" = "hello orca" ] || fail "Orca task inbox did not preserve the send body, got '$body'" assert_not_contains "$(cat "$LOG")" $'--text\x1fhello orca\x1f' \ "send typed the payload instead of recording it" - assert_contains "$(cat "$LOG")" $'orca\x1f''terminal'$'\x1f''send'$'\x1f''--terminal'$'\x1f''term-io'$'\x1f''--text'$'\x1f''Firstmate instruction waiting:' \ + assert_contains "$(cat "$LOG")" $'orca\x1f''terminal'$'\x1f''send'$'\x1f''--terminal'$'\x1f''term-io'$'\x1f''--text'$'\x1f'': Firstmate instruction waiting:' \ "send did not ring the inbox doorbell through the recorded Orca terminal" assert_contains "$(cat "$LOG")" $'orca\x1f''terminal'$'\x1f''send'$'\x1f''--terminal'$'\x1f''term-io'$'\x1f''--text'$'\x1f\x1f''--enter'$'\x1f''--json' \ "send did not submit the doorbell through the recorded Orca terminal" diff --git a/tests/fm-backlog-atomicity.test.sh b/tests/fm-backlog-atomicity.test.sh index a9b9fddff69..e305951d8dc 100755 --- a/tests/fm-backlog-atomicity.test.sh +++ b/tests/fm-backlog-atomicity.test.sh @@ -224,6 +224,166 @@ SH chmod +x "$case_dir/fakebin/tasks-axi" } +# The commit's `start` reports success but never moves the row - the beads +# failure pattern - and interrupts the spawn, so the deferred-signal exit path +# must read the preserved state back before it claims anything about it. +# <repair> decides whether a LATER start (the error path's own read-back +# repair) can move the row, or fails too. +lie_start_then_interrupt() { # <case-dir> <repair: works|fails> + local case_dir=$1 repair=$2 real + real=$(command -v tasks-axi) + cat > "$case_dir/fakebin/tasks-axi" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = start ]; then + if [ ! -f "$case_dir/start-interrupted" ]; then + : > "$case_dir/start-interrupted" + spawn_pid=\$(ps -o ppid= -p "\$PPID" | tr -d ' ') + case "\$spawn_pid" in ''|*[!0-9]*) exit 1 ;; esac + kill -TERM "\$spawn_pid" + exit 0 + fi + if [ "$repair" = fails ]; then + echo 'error: "backlog is unwritable"' >&2 + exit 1 + fi +fi +exec "$real" "\$@" +SH + chmod +x "$case_dir/fakebin/tasks-axi" +} + +# The commit's `start` reports success but never moves the row - the beads +# failure pattern - and interrupts the spawn; every later `start` (the error +# path's own read-back repair) never answers. With FM_TASKS_AXI_TIMEOUT +# bounding each call, the verification must time out and exit with the honest +# attempted wording instead of holding the per-task meta lock open forever. +hang_start_after_first() { # <case-dir> + local case_dir=$1 real + real=$(command -v tasks-axi) + cat > "$case_dir/fakebin/tasks-axi" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = start ]; then + if [ ! -f "$case_dir/start-interrupted" ]; then + : > "$case_dir/start-interrupted" + spawn_pid=\$(ps -o ppid= -p "\$PPID" | tr -d ' ') + case "\$spawn_pid" in ''|*[!0-9]*) exit 1 ;; esac + kill -TERM "\$spawn_pid" + exit 0 + fi + sleep 300 +fi +exec "$real" "\$@" +SH + chmod +x "$case_dir/fakebin/tasks-axi" +} + +# --- fm_tasks_axi's own bound ------------------------------------------------- + +# A PATH with no timeout variant on it - the stock-macOS shape, where GNU +# timeout is absent and coreutils does not ship gtimeout. fm_tasks_axi must +# still bound the call, through its perl watchdog, instead of running it +# unbounded under the per-task meta lock. +make_fallback_bin() { # <case-dir> <tasks-axi-stub-script> + local case_dir=$1 stub=$2 fb="$1/fallbackbin" + mkdir -p "$fb" + ln -s "$(command -v perl)" "$fb/perl" + ln -s "$(command -v sleep)" "$fb/sleep" + printf '%s\n' "$stub" > "$fb/tasks-axi" + chmod +x "$fb/tasks-axi" + printf '%s\n' "$fb" +} + +run_bounded_fm_tasks_axi() { # <fallback-bin> <bound> [args...] + local fb=$1 bound=$2 out rc=0 saved_path=$PATH + shift 2 + # The fallback shape itself: a PATH with no timeout variant on it. Set and + # restored here, never in a subshell, so the change cannot leak into other + # tests. + PATH="$fb" + out=$( + . "$ROOT/bin/fm-backlog-transition-lib.sh" + FM_TASKS_AXI_TIMEOUT="$bound" fm_tasks_axi "$@" 2>&1 + ) || rc=$? + PATH=$saved_path + printf '%s' "$out" + return "$rc" +} + +test_fm_tasks_axi_fallback_bounds_the_call_without_a_timeout_binary() { + local case_dir fb out rc=0 started + case_dir=$(make_home fm-tasks-axi-fallback) + fb=$(make_fallback_bin "$case_dir" '#!/bin/bash +exec sleep 300') + started=$SECONDS + out=$(run_bounded_fm_tasks_axi "$fb" 2 show never-answers) || rc=$? + [ "$rc" -eq 124 ] \ + || fail "the perl watchdog fallback did not report the call as timed out (rc=$rc, out=$out)" + [ $((SECONDS - started)) -ge 2 ] \ + || fail "the perl watchdog fallback fired before the bound elapsed" + [ $((SECONDS - started)) -lt 20 ] \ + || fail "the perl watchdog fallback did not bound the call (${SECONDS}s)" + pass "fm_tasks_axi bounds the call through its perl watchdog when no timeout binary exists" +} + +test_fm_tasks_axi_fallback_passes_the_child_status_and_output_through() { + local case_dir fb out rc=0 + case_dir=$(make_home fm-tasks-axi-passthrough) + fb=$(make_fallback_bin "$case_dir" '#!/bin/bash +echo "stub failed" +exit 7') + out=$(run_bounded_fm_tasks_axi "$fb" 5 show x) || rc=$? + [ "$rc" -eq 7 ] \ + || fail "the perl watchdog fallback did not pass the child status through (rc=$rc)" + assert_contains "$out" "stub failed" "the perl watchdog fallback lost the child's output" + pass "fm_tasks_axi's perl watchdog passes the child status and output through unchanged" +} + +test_fm_tasks_axi_fails_closed_when_nothing_can_bound_the_call() { + local case_dir fb out rc=0 + case_dir=$(make_home fm-tasks-axi-unboundable) + fb="$case_dir/unboundablebin" + mkdir -p "$fb" + printf '#!/bin/bash\nexit 0\n' > "$fb/tasks-axi" + chmod +x "$fb/tasks-axi" + out=$(run_bounded_fm_tasks_axi "$fb" 5 show x) || rc=$? + [ "$rc" -eq 127 ] \ + || fail "fm_tasks_axi ran the call although nothing could bound it (rc=$rc)" + assert_contains "$out" "cannot bound tasks-axi" \ + "the fail-closed diagnostic did not say why the call was refused" + pass "fm_tasks_axi fails closed rather than running unbounded when no bounding mechanism exists" +} + +test_fm_tasks_axi_gnu_timeout_forces_termination_of_a_sigterm_ignoring_child() { + local case_dir fb out rc=0 started + if ! command -v timeout >/dev/null 2>&1; then + pass "fm_tasks_axi's GNU timeout forces termination (skipped: no timeout binary on this host)" + return 0 + fi + case_dir=$(make_home fm-tasks-axi-kill-after) + # A real GNU timeout on the PATH, and a tasks-axi that ignores SIGTERM + # (the ignored disposition survives exec into sleep). timeout alone would + # wait forever for such a child; only its kill-after stops it, so this + # test fails on an unforced bound and passes once TERM is followed by + # KILL at one further bound. + fb="$case_dir/gnubin" + mkdir -p "$fb" + ln -s "$(command -v timeout)" "$fb/timeout" + ln -s "$(command -v sleep)" "$fb/sleep" + printf '#!/bin/bash\ntrap "" TERM\nexec sleep 300\n' > "$fb/tasks-axi" + chmod +x "$fb/tasks-axi" + started=$SECONDS + out=$(run_bounded_fm_tasks_axi "$fb" 2 show never-answers) || rc=$? + case $rc in + 124 | 137) ;; + *) fail "the GNU timeout path did not report the TERM-ignoring child as timed out (rc=$rc, out=$out)" ;; + esac + [ $((SECONDS - started)) -ge 2 ] \ + || fail "the GNU timeout path fired before the bound elapsed" + [ $((SECONDS - started)) -lt 20 ] \ + || fail "the GNU timeout path did not force-terminate the TERM-ignoring child (${SECONDS}s)" + pass "fm_tasks_axi's GNU timeout kills a child that ignores SIGTERM after one further bound" +} + change_row_on_second_show() { # <case-dir> <done|rm> local case_dir=$1 action=$2 real real=$(command -v tasks-axi) @@ -480,6 +640,73 @@ test_dispatch_omits_the_file_for_a_beads_show() { pass "dispatch omits the markdown file when probing a Beads backlog" } +test_completion_omits_the_file_for_a_beads_done() { + local case_dir home id out + id=atomic-completion-beads-b1 + case_dir=$(make_home completion-beads "$id") + home=$(home_of "$case_dir") + printf '%s\n' 'backend = "beads"' '[beads]' 'path = ".beads"' \ + 'prefix = "atomic"' > "$home/.tasks.toml" + # A Beads home keeps no markdown backlog at all: the transition gate, the + # row probe, and the close must all address the configured backend without + # requiring or overriding a markdown file. + rm -f "$home/data/backlog.md" + cat > "$case_dir/fakebin/tasks-axi" <<SH +#!/usr/bin/env bash +printf '%s\n' "\$*" >> "$case_dir/tasks-axi-calls" +case "\${1:-}" in + --version) printf '%s\n' '0.2.5' ;; + update) + [ "\${2:-}" = --help ] || exit 1 + printf '%s\n' '--archive-body' + ;; + mv) + [ "\${2:-}" = --help ] || exit 1 + printf '%s\n' 'usage: tasks-axi mv [<id>...]' + ;; + show) + [ "\${2:-}" = "$id" ] || exit 1 + case " \$* " in + *" --file "*) + printf '%s\n' 'error: beads show received a markdown file override' >&2 + exit 1 + ;; + esac + printf '%s\n' 'task:' + printf ' id: %s\n' "$id" + printf '%s\n' ' state: in_flight' ' held: no' ' blocked: no' + ;; + done) + [ "\${2:-}" = "$id" ] || exit 1 + case " \$* " in + *" --file "*) + printf '%s\n' 'error: beads done received a markdown file override' >&2 + exit 1 + ;; + esac + printf 'ok: done %s\n' "$id" + ;; + *) exit 1 ;; +esac +SH + chmod +x "$case_dir/fakebin/tasks-axi" + write_task_meta "$case_dir" "$id" ship local-only "spawn_gen=spawn-beads-done" + + out=$(run_teardown "$case_dir" "$id") \ + || fail "Beads completion teardown failed without a markdown backlog: $out" + assert_grep "done $id" "$case_dir/tasks-axi-calls" \ + "the Beads close never ran" + assert_no_grep "done $id --file" "$case_dir/tasks-axi-calls" \ + "the Beads close passed the markdown file to done" + assert_absent "$home/state/$id.meta" \ + "the Beads completion teardown left the task record behind" + assert_absent "$home/state/$id.backlog-close" \ + "the Beads completion teardown left its pending-close record behind" + assert_not_contains "$out" "backlog.md" \ + "the Beads teardown reported the close against a markdown file this home has not got" + pass "completion applies and closes a Beads backlog without any markdown file" +} + test_dispatch_refuses_a_pending_authoritative_close() { local case_dir id marker out rc=0 id=atomic-dispatch-pending-close-b1 @@ -995,8 +1222,8 @@ test_dispatch_defers_interruption_across_backlog_commit() { rc=0 out=$(run_ship_spawn "$case_dir" "$id") || rc=$? [ "$rc" -ne 0 ] || fail "a $timing-commit interruption was reported as success" - assert_contains "$out" "paired task record and In-flight backlog state were preserved" \ - "a $timing-commit interruption did not report its atomic outcome" + assert_contains "$out" "verified preserved: its paired task record is present and its backlog item is In flight" \ + "a $timing-commit interruption did not report its verified atomic outcome" [ "$(row_state "$case_dir" "$id")" = in_flight ] \ || fail "a $timing-commit interruption left the backlog row queued" assert_present "$(home_of "$case_dir")/state/$id.meta" \ @@ -1005,6 +1232,79 @@ test_dispatch_defers_interruption_across_backlog_commit() { pass "dispatch retries interrupted transitions before honoring termination" } +test_deferred_signal_reads_back_preserved_state() { + local case_dir id out rc=0 + id=atomic-dispatch-signal-readback-b5 + case_dir=$(make_home dispatch-signal-readback "$id") + add_item "$case_dir" "$id" + lie_start_then_interrupt "$case_dir" works + + out=$(run_ship_spawn "$case_dir" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "an interrupted spawn reported success" + assert_contains "$out" "moved to In flight now and verified" \ + "a signal-deferred spawn did not read the row back and repair it" + [ "$(row_state "$case_dir" "$id")" = in_flight ] \ + || fail "the read-back repair left the backlog row at $(row_state "$case_dir" "$id")" + assert_present "$(home_of "$case_dir")/state/$id.meta" \ + "the read-back repair lost the paired task record" + pass "a signal-deferred spawn verifies its preserved state and repairs a row that did not move" +} + +test_deferred_signal_never_claims_unverified_preservation() { + local case_dir id out rc=0 + id=atomic-dispatch-signal-unverified-b5 + case_dir=$(make_home dispatch-signal-unverified "$id") + add_item "$case_dir" "$id" + lie_start_then_interrupt "$case_dir" fails + + out=$(run_ship_spawn "$case_dir" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "an interrupted spawn reported success" + assert_contains "$out" "preservation could not be verified" \ + "an unverifiable preservation was not reported as attempted, not verified" + case "$out" in + *"In-flight backlog state were preserved"*) \ + fail "the interrupted spawn still claimed a preservation it never verified" ;; + esac + [ "$(row_state "$case_dir" "$id")" = queued ] \ + || fail "the failed repair left the backlog row at $(row_state "$case_dir" "$id")" + assert_present "$(home_of "$case_dir")/state/$id.meta" \ + "the failed repair removed the paired task record" + pass "a signal-deferred spawn reports attempted preservation as attempted when it cannot be verified" +} + +test_deferred_signal_verification_outlives_an_unresponsive_tasks_axi() { + local case_dir id out rc=0 + id=atomic-dispatch-signal-hang-b5 + case_dir=$(make_home dispatch-signal-hang "$id") + add_item "$case_dir" "$id" + hang_start_after_first "$case_dir" + + # The read-back's own `start` never answers, so the spawn must bound it + # (FM_TASKS_AXI_TIMEOUT=3), print the attempted wording naming the timeout, + # and exit - the outer `timeout -k 5 30` only turns a regression back into + # the lock-held-forever hang it exists to catch. + mkdir -p "$case_dir/user-home" + out=$(FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$(home_of "$case_dir")" \ + HOME="$case_dir/user-home" FM_SPAWN_NO_GUARD=1 \ + FM_FAKE_PANE_PATH="$case_dir/wt" TMUX="fake,1,0" CLAUDE_CONFIG_DIR='' \ + FM_TASKS_AXI_TIMEOUT=3 PATH="$case_dir/fakebin:$PATH" \ + timeout -k 5 30 "$SPAWN" "$id" "$case_dir/project" \ + --mode no-mistakes --yolo off 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "an interrupted spawn reported success" + case "$rc" in + 124|137) fail "the verification hung on the unresponsive start instead of timing out: $out" ;; + esac + assert_contains "$out" "preservation could not be verified" \ + "a timed-out verification did not report the preservation as attempted, not verified" + assert_contains "$out" "did not finish within 3s" \ + "the attempted-preservation wording did not name the timeout as its reason" + [ "$(row_state "$case_dir" "$id")" = queued ] \ + || fail "the timed-out repair left the backlog row at $(row_state "$case_dir" "$id")" + assert_present "$(home_of "$case_dir")/state/$id.meta" \ + "the timed-out repair removed the paired task record" + pass "a signal-deferred spawn bounds its verification so an unresponsive tasks-axi cannot hold the meta lock forever" +} + test_dispatch_interruption_during_kimi_readiness_fails_before_commit() { local case_dir home id out rc=0 id=atomic-dispatch-kimi-readiness-signal-b5 @@ -2079,6 +2379,13 @@ test_bootstrap_refuses_a_symlinked_state_directory_before_reconciliation() { test_bootstrap_stops_when_data_disappears_before_reconciliation() { local case_dir id saved out rc=0 + # The data-removal fault is injected by a fake stat on PATH; on Darwin the + # budget link-count helper now calls /usr/bin/stat directly, so the fake can + # never fire there. Skip the Darwin run of this case. + if [ "$(uname)" = Darwin ]; then + pass "bootstrap data-disappears fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat" + return + fi id=atomic-bootstrap-data-race-b11 case_dir=$(make_home bootstrap-data-race) add_item "$case_dir" "$id" @@ -2325,6 +2632,7 @@ test_a_persistent_secondmate_is_never_a_backlog_item() { test_dispatch_moves_the_item_in_flight_in_the_same_run test_dispatch_omits_the_file_for_a_beads_show +test_completion_omits_the_file_for_a_beads_done test_dispatch_refuses_a_pending_authoritative_close test_dispatch_refuses_a_held_row_before_creating_resources test_dispatch_refuses_a_blocked_row_before_creating_resources @@ -2347,6 +2655,13 @@ test_dispatch_reports_an_incomplete_record_rollback test_dispatch_reports_an_incomplete_busy_rollback test_dispatch_rolls_back_before_a_failed_launch_delivery test_dispatch_defers_interruption_across_backlog_commit +test_deferred_signal_reads_back_preserved_state +test_deferred_signal_never_claims_unverified_preservation +test_deferred_signal_verification_outlives_an_unresponsive_tasks_axi +test_fm_tasks_axi_fallback_bounds_the_call_without_a_timeout_binary +test_fm_tasks_axi_fallback_passes_the_child_status_and_output_through +test_fm_tasks_axi_fails_closed_when_nothing_can_bound_the_call +test_fm_tasks_axi_gnu_timeout_forces_termination_of_a_sigterm_ignoring_child test_dispatch_interruption_during_kimi_readiness_fails_before_commit test_dispatch_does_not_resurrect_a_row_closed_after_preflight test_dispatch_fails_when_its_row_vanishes_after_preflight diff --git a/tests/fm-bearings-board-lavish-live-e2e.test.sh b/tests/fm-bearings-board-lavish-live-e2e.test.sh new file mode 100755 index 00000000000..a413e27c3a0 --- /dev/null +++ b/tests/fm-bearings-board-lavish-live-e2e.test.sh @@ -0,0 +1,120 @@ +#!/usr/bin/env bash +# tests/fm-bearings-board-lavish-live-e2e.test.sh - live drift guard proving +# the real lavish-axi still behaves the way bin/fm-bearings-board.sh's session +# liveness check is written against. +# +# Why this file exists: the build's "is this board actually live" verdict comes +# from what lavish-axi emits, which is a surface the vendor controls and changes +# without notice. The defect this guards was exactly that - opening a session +# the captain had ended from the browser EXITS 0 while refusing to reopen, so a +# build that trusted the exit status armed a poll against a dead session and the +# board read "not listening" with nobody watching it. A stubbed lavish-axi can +# only confirm the assumption already written into the stub, so the assumption +# itself needs a run against the real tool. +# +# The captain-ended state is reached through the same server route the browser's +# End session button calls, so no browser is needed and nothing here depends on +# a human. The artifact is a scratch page in a temporary directory, and the +# session it opens is ended again before the guard returns. +# +# Standard CI has no lavish-axi, so this reports a capability skip there. The +# portable counterpart in tests/fm-bearings-board.test.sh pins the build's logic +# in CI against a stub that reproduces these shapes. Run this guard after a +# lavish-axi upgrade and before trusting refreshed evidence. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" + +fm_live_gate default-on FM_BEARINGS_LAVISH_LIVE lavish-axi jq curl + +pass() { printf 'ok - %s\n' "$1"; } +note() { printf '# %s\n' "$1"; } + +LAB='' +cleanup() { + [ -z "$LAB" ] || { + [ ! -f "$LAB/.lavish/bearings-board.html" ] \ + || lavish-axi end "$LAB/.lavish/bearings-board.html" >/dev/null 2>&1 || true + rm -rf "$LAB" + } +} +fail() { printf 'not ok - %s\n' "$1" >&2; cleanup; exit 1; } +trap cleanup EXIT + +VERSION=$(lavish-axi --version 2>/dev/null | tr -d '[:space:]') +note "lavish-axi ${VERSION:-version-unknown}" + +LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-bearings-lavish-live.XXXXXX") || fail "cannot create the guard lab" +LAB=$(cd -P -- "$LAB" && pwd -P) +mkdir -p "$LAB/state" "$LAB/data" + +cat > "$LAB/payload.json" <<'JSON' +{ + "schema": "fm-bearings-board.v1", + "home": "lavish-live-guard", + "generated": "2026-01-01T00:00Z", + "prs_live": false, + "captains_call": [ + { + "key": "sample-live-guard-call", + "type": "decision", + "repo": "sample", + "title": "Guard placeholder", + "options": [{ "value": "yes", "label": "Yes" }] + } + ], + "underway": [], + "landed": [], + "charted": [] +} +JSON + +run_board() { + FM_HOME="$LAB" FM_STATE_OVERRIDE="$LAB/state" FM_DATA_OVERRIDE="$LAB/data" \ + FM_PROCEVENT_CLAIM_ROOT="$LAB/procevent-claims" \ + "$ROOT/bin/fm-bearings-board.sh" "$@" +} + +BOARD="$LAB/.lavish/bearings-board.html" +run_board build "$LAB/payload.json" >/dev/null 2>&1 || fail "the guard board did not build" +[ -f "$BOARD" ] || fail "the guard board was not published" + +url=$(lavish-axi "$BOARD" | sed -n 's/^[[:space:]]*url:[[:space:]]*//p' | head -1 | tr -d '"') +case "$url" in + http://*/session/*) ;; + *) fail "could not read the guard board session url: $url" ;; +esac +key=${url##*/} +base=${url%/session/*} + +# End it exactly as the browser's End session button does. +curl -fsS -X POST "$base/api/$key/end" >/dev/null 2>&1 \ + || fail "could not end the guard board session as the captain" + +# ASSUMPTION UNDER GUARD: this exits 0 while reporting the session is not live. +set +e +ended_out=$(lavish-axi "$BOARD" 2>&1) +ended_rc=$? +set -e +[ "$ended_rc" -eq 0 ] \ + || fail "lavish-axi ${VERSION:-version-unknown} now exits $ended_rc on a captain-ended session; the board build's liveness check must be revisited" +ended_status=$(printf '%s\n' "$ended_out" | sed -n 's/^[[:space:]]*status:[[:space:]]*//p' | head -1 | tr -d '"') +[ "$ended_status" != opened ] \ + || fail "lavish-axi ${VERSION:-version-unknown} silently reopened a captain-ended session; the board build's liveness check must be revisited" +lavish-axi 2>/dev/null | grep -F "$BOARD," | grep -q ',open,' \ + && fail "lavish-axi ${VERSION:-version-unknown} still lists a captain-ended session as open; the board build's liveness check must be revisited" +pass "lavish-axi ${VERSION:-version-unknown} reports a captain-ended session without reopening it and without failing" + +# THE BEHAVIOR UNDER GUARD: the build must not accept that, and must recover. +out=$(run_board build "$LAB/payload.json" 2>&1) \ + || fail "the board build refused a recoverable captain-ended session: $out" +case "$out" in + *"session: reopened"*) ;; + *) fail "the board build did not reopen the captain-ended session: $out" ;; +esac +lavish-axi 2>/dev/null | grep -F "$BOARD," | grep -q ',open,' \ + || fail "the board build reported success while the session was still not live" +pass "the board build reopens a captain-ended session against real lavish-axi instead of arming a dead one" diff --git a/tests/fm-bearings-board-render.test.sh b/tests/fm-bearings-board-render.test.sh index afa6b9350cc..cf26fd31428 100755 --- a/tests/fm-bearings-board-render.test.sh +++ b/tests/fm-bearings-board-render.test.sh @@ -20,9 +20,39 @@ command -v node >/dev/null 2>&1 || { echo "skip: node not found"; exit 0; } make_home() { # <name> local home="$TMP_ROOT/$1" fakebin + # A build starts a listener for the board it publishes. Registered with + # tests/lib.sh, not with a shell array: make_home runs inside a command + # substitution, where an array append never reaches the caller. + fm_test_track_procevent_home "$home" "$home/procevent-claims" mkdir -p "$home/state" "$home/data" fakebin=$(fm_fakebin "$home") - fm_fake_exit0 "$fakebin" lavish-axi + # The build proves the board session is live before it arms anything, so the + # stub reports the opened shape the real lavish-axi emits. This suite is about + # what the template renders, not about session liveness, which + # tests/fm-bearings-board.test.sh owns. + cat > "$fakebin/lavish-axi" <<'SH' +#!/usr/bin/env bash +case "${1-}" in + --version) printf '0.1.61\n' ;; + '') + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + [ ! -s "$FM_HOME/lavish-open" ] \ + || printf ' %s,open,"http://127.0.0.1/session/render",0\n' "$(cat "$FM_HOME/lavish-open")" + ;; + poll) + # Bounded, so a listener that escapes its test stops on its own. + while [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ]; do sleep 1; done + exit 75 + ;; + *) + real=$(cd "$(dirname "$1")" && pwd -P)/$(basename "$1") + printf '%s\n' "$real" > "$FM_HOME/lavish-open" + printf 'session:\n status: opened\n' + ;; +esac +exit 0 +SH + chmod +x "$fakebin/lavish-axi" printf '%s\n' "$home" } diff --git a/tests/fm-bearings-board.test.sh b/tests/fm-bearings-board.test.sh index d87acb652a2..b59010036e9 100644 --- a/tests/fm-bearings-board.test.sh +++ b/tests/fm-bearings-board.test.sh @@ -13,20 +13,100 @@ TMP_ROOT=$(fm_test_tmproot fm-bearings-board) command -v jq >/dev/null 2>&1 || { echo "skip: jq not found"; exit 0; } +# A lavish-axi stub that reproduces the shapes verified against the real +# lavish-axi 0.1.61, because the build's liveness verdict is read from what the +# vendor emits. The load-bearing shape is the refusal: opening a session the +# captain ended from the browser EXITS 0 while reporting `status: user-ended`, +# and that session is absent from the server's listing. `--reopen` restores it. +# Markers under lavish-state drive the fixture: `user-ended` makes the next +# plain open refuse, and `refuse-reopen` makes even --reopen leave it dead. make_home() { # <name> local home="$TMP_ROOT/$1" fakebin - mkdir -p "$home/state" "$home/data" + # Registered with tests/lib.sh, not with a shell array: make_home is called + # inside a command substitution, so an array append here never reaches the + # caller and every listener this suite started used to survive the run. + fm_test_track_procevent_home "$home" "$home/procevent-claims" + mkdir -p "$home/state" "$home/data" "$home/lavish-state" fakebin=$(fm_fakebin "$home") - fm_fake_exit0 "$fakebin" lavish-axi + cat > "$fakebin/lavish-axi" <<'SH' +#!/usr/bin/env bash +set -u +state=${LAVISH_FAKE_STATE:?} +emit() { # <canonical-file> <status> + printf 'session:\n' + printf ' file: %s\n' "$1" + printf ' url: "http://127.0.0.1:4387/session/deadbeef"\n' + printf ' status: %s\n' "$2" +} +case "${1-}" in + --version) printf '0.1.61\n'; exit 0 ;; + poll) + # A real blocking listener: it returns only when the trigger appears, so a + # live owner in these tests is a live process rather than a timing artifact. + # Both waits are bounded, so a listener that escapes its test cannot keep + # spawning processes for as long as the host stays up. + limit=${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120} + while [ ! -e "$state/poll-trigger" ]; do + [ "$SECONDS" -lt "$limit" ] || exit 75 + sleep 0.05 + done + printf 'session:\n status: ended\n' + if [ -e "$state/hold-after-terminal" ]; then + : > "$state/terminal-emitted" + while [ -e "$state/hold-after-terminal" ]; do + [ "$SECONDS" -lt "$limit" ] || exit 75 + sleep 0.05 + done + fi + exit 0 + ;; + '') + if [ -e "$state/end-before-next-list" ]; then + : > "$state/open" + rm -f "$state/end-before-next-list" + fi + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + if [ -s "$state/open" ]; then + while IFS= read -r listed; do + [ -n "$listed" ] || continue + printf ' %s,open,"http://127.0.0.1:4387/session/deadbeef",0\n' "$listed" + done < "$state/open" + fi + exit 0 + ;; + end) : > "$state/open"; printf 'session:\n status: ended\n'; exit 0 ;; +esac +file=$1 +shift +reopen=0 +for arg in "$@"; do [ "$arg" != --reopen ] || reopen=1; done +real=$(cd "$(dirname "$file")" && pwd -P)/$(basename "$file") +if [ -e "$state/user-ended" ] && [ "$reopen" = 0 ]; then + emit "$real" user-ended + exit 0 +fi +if [ -e "$state/refuse-reopen" ]; then + emit "$real" user-ended + exit 0 +fi +rm -f -- "$state/user-ended" +printf '%s\n' "$real" > "$state/open" +emit "$real" opened +exit 0 +SH + chmod +x "$fakebin/lavish-axi" printf '%s\n' "$home" } +end_session_as_captain() { : > "$1/lavish-state/user-ended"; : > "$1/lavish-state/open"; } + run_board() { # <home> <args...> local home=$1 shift PATH="$home/fakebin:$PATH" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ FM_PROCEVENT_CLAIM_ROOT="$home/procevent-claims" \ + LAVISH_FAKE_STATE="$home/lavish-state" \ "$BOARD" "$@" } @@ -154,6 +234,12 @@ test_build_refuses_malformed_payloads_before_touching_the_board() { set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e [ "$rc" -ne 0 ] || fail "a negative omitted-warning count was accepted" + write_valid_payload "$data" + jq '.captains_call[0].subject = {"artifact":"quota-axi","version":"0.1"}' "$data" > "$data.tmp" \ + && mv "$data.tmp" "$data" + set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e + [ "$rc" -ne 0 ] || fail "an invalid structured version subject was accepted" + write_valid_payload "$data" jq '.captains_call[0].type = "verdict"' "$data" > "$data.tmp" && mv "$data.tmp" "$data" set +e; out=$(run_board "$home" build "$data" 2>&1); rc=$?; set -e @@ -219,13 +305,19 @@ test_build_injects_binds_then_arms() { assert_contains "$out" "armed: " "the first build did not arm the board source: $out" assert_present "$board" "build reported success without a board" - # Round-trip: the payload extracted from the built page is byte-for-byte the - # same JSON document, and the escaped </script> string can no longer - # terminate the data block. + # Round-trip: apart from the reconcile choice the build adds to every + # decision card, the payload extracted from the built page is the same JSON + # document, and the escaped </script> string can no longer terminate the + # data block. extract_payload "$board" | jq -S . > "$home/extracted.json" \ || fail "the built board does not carry parseable payload JSON" - jq -S . "$data" > "$home/expected.json" - diff -u "$home/expected.json" "$home/extracted.json" >/dev/null \ + jq -S '.captains_call = [.captains_call[] + | .options = [.options[] | select(.value != "reconcile")]]' \ + "$home/extracted.json" > "$home/stripped.json" + jq -S '.captains_call = [.captains_call[] + | .options = [.options[] | select(.value != "reconcile")]]' \ + "$data" > "$home/expected.json" + diff -u "$home/expected.json" "$home/stripped.json" >/dev/null \ || fail "the injected payload does not round-trip to the input document" grep -qF '</script><b>' "$board" \ && fail "a payload string embedded a live closing script tag in the page" @@ -285,7 +377,16 @@ SH chmod +x "$runtime/bin/fm-procevent-lavish.sh" cat > "$home/fakebin/lavish-axi" <<'SH' #!/usr/bin/env bash +if [ -z "${1:-}" ]; then + printf 'sessions[1]{file,status,url,pending_prompts}:\n' + [ ! -s "$FM_HOME/order-open" ] \ + || printf ' %s,open,"http://127.0.0.1/session/order",0\n' "$(cat "$FM_HOME/order-open")" + exit 0 +fi if [ "${1:-}" != poll ]; then + real=$(cd "$(dirname "$1")" && pwd -P)/$(basename "$1") + printf '%s\n' "$real" > "$FM_HOME/order-open" + printf 'session:\n status: opened\n' exit 0 fi cat <<EOF @@ -293,7 +394,7 @@ session: status: feedback session_ended: false prompts[1]{uid,prompt,selector,tag,text}: - "2","Order proof: yes\\n\\nContext data:\\n{\\n \\"question\\": \\"$ORDER_PROOF_HOLD\\",\\n \\"answer\\": \\"yes\\"\\n}","form",choice,"Order proof: yes" + "2","Order proof: yes\\n\\nContext data:\\n{\\n \\"schema\\": \\"fm-bearings-answer.v1\\",\\n \\"question\\": \\"$ORDER_PROOF_HOLD\\",\\n \\"selection\\": \\"yes\\",\\n \\"note\\": \\"\\"\\n}","form",choice,"Order proof: yes" EOF SH chmod +x "$home/fakebin/lavish-axi" @@ -404,6 +505,264 @@ test_charted_kind_is_optional_and_accepts_both_values() { pass "charted kind is optional and accepts queued and warning" } + +# --- part 1: never arm a poll on an ended session --------------------------- + +test_build_reopens_a_session_the_captain_ended() { + local home data board out sid claim old_pid old_token new_pid new_token + home=$(make_home ended-session) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the first build failed" + sid=$(run_lavish_source_id "$home" "$board") + claim="$home/procevent-claims/$sid.claim" + old_pid=$(sed -n '2p' "$claim") + old_token=$(sed -n '3p' "$claim") + + # The reported case: the captain ends the board from the browser, so opening + # it again keeps the same session id, reports it ended, and EXITS 0. A build + # that trusts the exit status arms a poll nothing can ever attach to. + : > "$home/lavish-state/hold-after-terminal" + : > "$home/lavish-state/poll-trigger" + for _ in $(seq 1 100); do + [ -e "$home/lavish-state/terminal-emitted" ] && break + sleep 0.05 + done + [ -e "$home/lavish-state/terminal-emitted" ] \ + || fail "the old listener did not receive its terminal result" + rm -f "$home/lavish-state/poll-trigger" + end_session_as_captain "$home" + out=$(run_board "$home" build "$data") || fail "the rebuild refused a recoverable ended session" + rm -f "$home/lavish-state/hold-after-terminal" + assert_contains "$out" "session: reopened" \ + "the rebuild did not reopen the ended session: $out" + [ ! -e "$home/lavish-state/user-ended" ] \ + || fail "the rebuild reported success while the session was still ended" + new_pid=$(sed -n '2p' "$claim") + new_token=$(sed -n '3p' "$claim") + [ "$new_pid" != "$old_pid" ] || [ "$new_token" != "$old_token" ] \ + || fail "the rebuild accepted the pre-reopen source generation" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the reopened board has no live listener" + pass "a board build reopens a session the captain ended instead of arming a dead one" +} + +test_build_reopens_when_an_opened_session_ends_before_listing() { + local home data out board sid + home=$(make_home establish-list-race) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + : > "$home/lavish-state/end-before-next-list" + out=$(run_board "$home" build "$data") || fail "the raced session build failed: $out" + assert_contains "$out" "session: reopened" \ + "the build trusted an opened response after the server no longer listed it: $out" + sid=$(run_lavish_source_id "$home" "$board") + [ -s "$home/lavish-state/open" ] || fail "the raced session was not live before arming" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the replacement session did not receive a live listener" + pass "build reopens a session that ends between establish and listing" +} + +test_build_refuses_to_arm_when_the_session_stays_ended() { + local home data rc out sid + home=$(make_home dead-session) + data="$home/payload.json" + write_valid_payload "$data" + # An ended session that will not come back: the build must stop rather than + # register a poll against it. + : > "$home/lavish-state/refuse-reopen" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "build armed a poll on a session that stayed ended: $out" + assert_contains "$out" "ended session" "the refusal did not say why: $out" + sid=$(run_lavish_source_id "$home" "$home/.lavish/bearings-board.html") + ! run_decisions "$home" binding "$sid" >/dev/null 2>&1 \ + || fail "build bound the board to a session that stayed ended" + ! run_procevent "$home" list | awk 'NR > 1 { print $1 }' | grep -Fxq "$sid" \ + || fail "build armed the board against a session that stayed ended" + pass "build refuses to arm a poll on a session that stays ended" +} + +test_build_starts_a_listener_for_an_already_armed_board() { + local home data board out sid claim + home=$(make_home relisten) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the first build failed" + sid=$(run_lavish_source_id "$home" "$board") + + # Registered is not listening: drop the listener the way a crashed generation + # would, then rebuild. `already-armed` must not be the end of the story. + claim="$home/procevent-claims/$sid.claim" + assert_present "$claim" "the first build left no listener to lose" + kill -KILL -"$(sed -n '2p' "$claim")" 2>/dev/null || true + kill -KILL "$(sed -n '2p' "$claim")" 2>/dev/null || true + sleep 1 + + out=$(run_board "$home" build "$data") || fail "the rebuild failed" + assert_contains "$out" "already-armed: $sid" "the rebuild re-registered the source: $out" + [ "$(run_procevent "$home" list | awk -v id="$sid" 'NR > 1 && $1 == id { print $3 }')" = live ] \ + || fail "the rebuilt board is registered but nothing is listening" + pass "a rebuild starts a listener when an already-armed board has none" +} + +# --- part 2: a landed subject is not a live call ---------------------------- + +test_build_drops_decision_cards_whose_subject_already_landed() { + local home data board out + home=$(make_home landed-cards) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + jq '.captains_call = [ + {"key":"landed-by-task","type":"decision","repo":"sample","title":"Already shipped", + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"timeout-reattach","type":"decision","repo":"sample","title":"Already merged", + "pr_url":"https://github.com/sample/sample/pull/7", + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"quota-version","type":"decision","repo":"sample","title":"Old quota release", + "subject":{"artifact":"quota-axi","version":"0.1.37"}, + "options":[{"value":"yes","label":"Yes"}]}, + {"key":"still-open","type":"decision","repo":"sample","title":"Genuinely open", + "subject":{"artifact":"quota-axi","version":"0.2.0"}, + "options":[{"value":"yes","label":"Yes"}]} + ] + | .landed = [ + {"id":"landed-by-task","repo":"sample","what":"shipped it","owner":"crew"}, + {"id":"some-other-task","repo":"sample","what":"merged timeout reattach","owner":"crew", + "pr_url":"https://github.com/sample/sample/pull/7"}, + {"id":"quota-release","repo":"sample","what":"published quota-axi","owner":"crew", + "subject":{"artifact":"quota-axi","version":"0.1.38"}}, + {"id":"unrelated\nstill-open","repo":"sample","what":"unrelated multiline identity","owner":"crew"} + ]' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + + out=$(run_board "$home" build "$data" 2>&1) || fail "the hygiene build failed: $out" + assert_contains "$out" "dropped-landed-card: landed-by-task" \ + "the build did not report dropping the landed work item card: $out" + assert_contains "$out" "dropped-landed-card: timeout-reattach" \ + "the build did not report dropping the merged timeout/reattach card: $out" + assert_contains "$out" "dropped-landed-card: quota-version" \ + "the build did not report dropping the superseded quota-axi version card: $out" + extract_payload "$board" | jq -e '[.captains_call[].key] == ["still-open"]' >/dev/null \ + || fail "the board dropped an open card or kept one whose subject already landed" + pass "build drops decision cards whose subject already landed and keeps open ones" +} + +test_build_keeps_a_decision_absent_from_the_main_backlog() { + local home data board out + home=$(make_home remote-decision-card) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + cp "$ROOT/.tasks.toml" "$home/.tasks.toml" + cat > "$home/data/backlog.md" <<'EOF' +## In flight + +## Queued + +## Done +EOF + write_valid_payload "$data" + jq '.captains_call = [{ + "key":"remote-mate-call","type":"decision","repo":"sample", + "title":"Remote secondmate decision", + "options":[{"value":"yes","label":"Yes"}] + }] + | .landed = []' "$data" > "$data.tmp" && mv "$data.tmp" "$data" + + out=$(run_board "$home" build "$data" 2>&1) || fail "the remote-card build failed: $out" + assert_not_contains "$out" "dropped-landed-card: remote-mate-call" \ + "an absent remote card was reported as landed: $out" + extract_payload "$board" | jq -e ' + [.captains_call[] | select(.key == "remote-mate-call")] | length == 1 + ' >/dev/null || fail "the hygiene check dropped a decision absent from the main backlog" + pass "build keeps remote decisions absent from the main backlog" +} + +# --- part 3: every decision card offers reconcile --------------------------- + +test_build_fails_when_reconcile_cannot_establish_a_listener() { + local home data out rc sid + home=$(make_home no-listener) + data="$home/payload.json" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "could not establish the listener fixture" + sid=$(run_lavish_source_id "$home" "$home/.lavish/bearings-board.html") + cat > "$home/fakebin/ps" <<'SH' +#!/usr/bin/env bash +exit 1 +SH + chmod +x "$home/fakebin/ps" + set +e + out=$(FM_PROC_ROOT_OVERRIDE="$home/no-proc" run_board "$home" build "$data" 2>&1) + rc=$? + set -e + rm -f "$home/fakebin/ps" + [ "$rc" -ne 0 ] || fail "a build with an uncertain listener reported success: $out" + assert_contains "$out" "source $sid is not listening after reconcile" \ + "the refusal did not name the source: $out" + assert_contains "$out" "observed owner: uncertain" \ + "the refusal did not name the observed owner: $out" + pass "build fails when reconcile cannot prove a live listener" +} + +test_every_decision_card_carries_the_reconcile_choice() { + local home data board + home=$(make_home reconcile-option) + data="$home/payload.json" + board="$home/.lavish/bearings-board.html" + write_valid_payload "$data" + run_board "$home" build "$data" >/dev/null || fail "the reconcile-option build failed" + extract_payload "$board" | jq -e ' + ([.captains_call[] | select(.type == "decision")] | length) > 0 + and ([.captains_call[] + | select(.type == "decision") + | ([.options[] | select(.value == "reconcile")] | length) == 1 + and ([.options[] | select(.value == "reconcile") | .label | length > 0] | all)] | all) + ' >/dev/null || fail "a decision card was published without the reconcile choice" + extract_payload "$board" | jq -e ' + ([.captains_call[] | select(.type != "decision") + | .options[] | select(.value == "reconcile")] | length) == 0 + ' >/dev/null || fail "reconcile was injected into a non-decision card" + pass "every decision card carries exactly one reconcile choice" +} + +test_build_refuses_a_payload_that_occupies_the_reconcile_value() { + local home data rc out + home=$(make_home reconcile-reserved) + data="$home/payload.json" + write_valid_payload "$data" + jq '.captains_call[0].options += [{"value":"reconcile","label":"Something else"}]' \ + "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a payload occupying the reserved reconcile value was accepted" + assert_absent "$home/.lavish/bearings-board.html" "a refused payload still produced a board" + pass "build refuses a payload that occupies the reserved reconcile value" +} + +test_build_refuses_a_nondecision_reconcile_value() { + local home data rc out + home=$(make_home merge-reconcile-reserved) + data="$home/payload.json" + write_valid_payload "$data" + jq '.captains_call[1].options += [{"value":"reconcile","label":"Merge action"}]' \ + "$data" > "$data.tmp" && mv "$data.tmp" "$data" + set +e + out=$(run_board "$home" build "$data" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a merge card occupying the reconcile value was accepted" + assert_absent "$home/.lavish/bearings-board.html" "a refused merge card still produced a board" + pass "build reserves reconcile across non-decision cards" +} + test_path_is_stable_and_home_scoped test_build_refuses_malformed_payloads_before_touching_the_board test_charted_kind_is_optional_and_accepts_both_values @@ -412,3 +771,13 @@ test_registration_cannot_consume_before_any_origin_binding test_build_does_not_bind_or_arm_when_session_start_fails test_rebuild_is_idempotent_and_does_not_double_arm test_build_refuses_a_template_without_exactly_one_slot +test_build_reopens_a_session_the_captain_ended +test_build_reopens_when_an_opened_session_ends_before_listing +test_build_refuses_to_arm_when_the_session_stays_ended +test_build_starts_a_listener_for_an_already_armed_board +test_build_drops_decision_cards_whose_subject_already_landed +test_build_keeps_a_decision_absent_from_the_main_backlog +test_build_fails_when_reconcile_cannot_establish_a_listener +test_every_decision_card_carries_the_reconcile_choice +test_build_refuses_a_payload_that_occupies_the_reconcile_value +test_build_refuses_a_nondecision_reconcile_value diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index 3423cc4acfc..b25d9c69899 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -519,7 +519,7 @@ test_orca_backend_gates_orca_tool_only_when_selected() { printf '%s\n' manual > "$case_dir/home/config/backlog-backend" printf '%s\n' orca > "$case_dir/home/config/backend" fakebin=$(make_fake_toolchain "$case_dir") - out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" orca)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") [ "$out" = "$missing_orca" ] || fail "backend=orca should require only the Orca-specific missing tool, got: $out" @@ -901,17 +901,17 @@ exit 1 SH chmod +x "$fakebin/gh" - all_out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + all_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 "$ROOT/bin/fm-bootstrap.sh") assert_contains "$all_out" "MISSING: node (install:" "the unsplit run lost its local diagnostic" assert_contains "$all_out" "NEEDS_GH_AUTH" "the unsplit run lost its network diagnostic" - skip_out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + skip_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=skip "$ROOT/bin/fm-bootstrap.sh") assert_contains "$skip_out" "MISSING: node (install:" "the local half lost its own diagnostic" assert_not_contains "$skip_out" "NEEDS_GH_AUTH" "the local half still made a network call" - only_out=$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + only_out=$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=only "$ROOT/bin/fm-bootstrap.sh") assert_contains "$only_out" "NEEDS_GH_AUTH" "the network half lost its own diagnostic" assert_not_contains "$only_out" "MISSING: node" "the network half repeated the local half's work" @@ -922,7 +922,7 @@ SH # A typo must never silently drop a safety sweep, so anything unrecognized # resolves to the complete run. - [ "$(PATH="$fakebin:$BASE_PATH" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ + [ "$(PATH="$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)" FM_HOME="$case_dir/home" FM_ROOT_OVERRIDE="$case_dir/home" \ FM_FAKE_TREEHOUSE_LEASE_HELP=1 FM_BOOTSTRAP_NETWORK=sikp "$ROOT/bin/fm-bootstrap.sh")" = "$all_out" ] \ || fail "an unrecognized FM_BOOTSTRAP_NETWORK value did not fall back to the complete run" pass "bootstrap: FM_BOOTSTRAP_NETWORK partitions one run into local and network halves" diff --git a/tests/fm-brief.test.sh b/tests/fm-brief.test.sh index 2840d99f3a0..5467b5cbdca 100755 --- a/tests/fm-brief.test.sh +++ b/tests/fm-brief.test.sh @@ -846,6 +846,29 @@ test_scout_and_secondmate_scaffold() { pass "fm-brief: scout and secondmate code paths still scaffold well-formed briefs" } +test_worker_role_scope() { + local kind home brief + home="$TMP_ROOT/worker-role" + for kind in no-mistakes direct-PR local-only scout; do + if [ "$kind" = scout ]; then + FM_HOME="$home" "$ROOT/bin/fm-brief.sh" "$kind" arbitrary-project-name --scout >/dev/null || fail "scout scaffold failed" + else + FM_HOME="$home" "$ROOT/bin/fm-brief.sh" "$kind" arbitrary-project-name --mode "$kind" >/dev/null || fail "$kind scaffold failed" + fi + brief="$home/data/$kind/brief.md" + assert_no_grep '# Current worker role contract' "$brief" "$kind scaffolded a second owner of the role scope fm-spawn.sh delivers" + done + FM_HOME="$home" FM_SECONDMATE_CHARTER='Supervise assigned work.' \ + "$ROOT/bin/fm-brief.sh" supervisor --secondmate --no-projects >/dev/null || fail "secondmate scaffold failed" + brief="$home/data/supervisor/brief.md" + assert_no_grep '# Current worker role contract' "$brief" "secondmate received the worker exception" + assert_no_grep 'do not adopt the supervisor identity' "$brief" "secondmate received the worker exception" + assert_grep "The local \`AGENTS.md\` is your job description" "$brief" "secondmate lost its supervisor contract" + assert_grep 'That file is your parent channel' "$brief" "secondmate lost its parent channel" + pass "fm-brief: scaffolds leave the worker role scope to the launch boundary and keep the secondmate contract" +} + +test_worker_role_scope test_script_parses test_no_heredoc_in_command_substitution test_help_includes_entire_header diff --git a/tests/fm-calm-pi-extension.test.sh b/tests/fm-calm-pi-extension.test.sh index 0cec12f290e..b66240c25f7 100755 --- a/tests/fm-calm-pi-extension.test.sh +++ b/tests/fm-calm-pi-extension.test.sh @@ -71,6 +71,82 @@ find_chrome() { return 1 } +# Render an exported session in real Chrome and leave the DOM in <out_file>. +# +# Rendering is a vendor-tool step, not a Calm guarantee: the DOM assertions the +# caller runs afterwards are what protect the contract. Headless Chrome start-up +# is the part that fails intermittently on a loaded CI runner - it can exit +# before writing any DOM at all - and the original single unattended attempt +# discarded both Chrome's stderr and its exit status, so a CI break surfaced as +# a bare "could not render" with nothing in the log to tell a Chrome start-up +# crash apart from a real change in Pi's export shape. +# +# So: retry the render a bounded number of times on a fresh profile, and when +# every attempt fails, print the Chrome binary, its version, the installed Pi +# version, and each attempt's exit status, stderr tail, and whether the helper +# timed the attempt out - when it did, the exit status is only this helper's own +# kill signal. The extra flags remove Chrome's background-network and /dev/shm +# dependencies, which are the start-up surfaces that fail on a runner; neither +# changes the rendered DOM of a local file. +render_export_dom() { + local chrome=$1 source_file=$2 out_file=$3 pi_version=$4 + local attempt pid status wait_count wait_limit reap_wait log profile report timed_out + report="$TMP_ROOT/chrome-render-report.txt" + wait_limit=${FM_CHROME_RENDER_WAIT_TICKS:-300} + : >"$report" + for attempt in 1 2 3; do + log="$TMP_ROOT/chrome-render-$attempt.err" + profile="$TMP_ROOT/chrome-profile-$attempt" + rm -rf "$profile" + : >"$out_file" + "$chrome" \ + --headless=new \ + --disable-gpu \ + --no-sandbox \ + --disable-dev-shm-usage \ + --disable-background-networking \ + --user-data-dir="$profile" \ + --virtual-time-budget=2000 \ + --dump-dom \ + "file://$source_file" >"$out_file" 2>"$log" & + pid=$! + # Check the DOM before Chrome's liveness, so an attempt that writes the + # complete dump and exits immediately is still read as a success. + wait_count=0 + while [ "$wait_count" -lt "$wait_limit" ]; do + grep -Fq '</html>' "$out_file" 2>/dev/null && break + kill -0 "$pid" 2>/dev/null || break + sleep 0.1 + wait_count=$((wait_count + 1)) + done + timed_out=no + if [ "$wait_count" -ge "$wait_limit" ]; then + timed_out=yes + fi + kill "$pid" 2>/dev/null || true + # Chrome can retain --headless=new after --dump-dom completes and ignore TERM, + # so an unbounded wait can hang after the complete DOM has been captured. + reap_wait=0 + while kill -0 "$pid" 2>/dev/null && [ "$reap_wait" -lt 20 ]; do + sleep 0.1 + reap_wait=$((reap_wait + 1)) + done + if kill -0 "$pid" 2>/dev/null; then + kill -9 "$pid" 2>/dev/null || true + fi + status=0 + wait "$pid" 2>/dev/null || status=$? + grep -Fq '</html>' "$out_file" 2>/dev/null && return 0 + printf 'attempt %s: exit=%s timed_out=%s bytes=%s stderr=%s\n' \ + "$attempt" "$status" "$timed_out" "$(wc -c <"$out_file" | tr -d ' ')" \ + "$(tail -c 400 "$log" 2>/dev/null | tr '\n' ' ')" >>"$report" + done + printf 'chrome=%s chrome_version=%s pi=%s; %s' \ + "$chrome" "$("$chrome" --version 2>&1 | head -1)" "$pi_version" \ + "$(tr '\n' ' ' <"$report")" + return 1 +} + test_home_resolution() { local fixture out status version if ! command -v node >/dev/null 2>&1 || ! command -v npm >/dev/null 2>&1; then @@ -704,7 +780,7 @@ import { fileURLToPath, pathToFileURL } from "node:url"; const extPath = fileURLToPath(pathToFileURL(process.env.EXT).href); const packageRoot = process.env.PI_PACKAGE_DIR; -const [{ AssistantMessageComponent }, { CustomEntryComponent }, { ToolExecutionComponent }, { UserMessageComponent }, { InteractiveMode }, { initTheme, theme }, { Text, getKeybindings, setCapabilities }, { createToolHtmlRenderer }] = await Promise.all([ +const [{ AssistantMessageComponent }, { CustomEntryComponent }, { ToolExecutionComponent }, { UserMessageComponent }, { InteractiveMode }, { initTheme, theme }, { Text, getKeybindings, setCapabilities }, { createToolHtmlRenderer }, { createReadToolDefinition, createBashToolDefinition, createEditToolDefinition, createWriteToolDefinition, createGrepToolDefinition, createFindToolDefinition, createLsToolDefinition }] = await Promise.all([ import(pathToFileURL(`${packageRoot}/dist/modes/interactive/components/assistant-message.js`).href), import(pathToFileURL(`${packageRoot}/dist/modes/interactive/components/custom-entry.js`).href), import(pathToFileURL(`${packageRoot}/dist/modes/interactive/components/tool-execution.js`).href), @@ -713,7 +789,22 @@ const [{ AssistantMessageComponent }, { CustomEntryComponent }, { ToolExecutionC import(pathToFileURL(`${packageRoot}/dist/modes/interactive/theme/theme.js`).href), import(pathToFileURL(`${packageRoot}/node_modules/@earendil-works/pi-tui/dist/index.js`).href), import(pathToFileURL(`${packageRoot}/dist/core/export-html/tool-renderer.js`).href), + // The calm-off equivalence baseline needs each built-in's REAL stock renderers. + // Pi 0.84 and older silently substituted the built-in definition when a + // ToolExecutionComponent was constructed without one, so a definition-less + // baseline used to read as stock; Pi 0.85 removed that substitution and the + // definition-less row now renders the generic text fallback instead. + import(pathToFileURL(`${packageRoot}/dist/core/tools/index.js`).href), ]); +const stockDefinitions = { + read: createReadToolDefinition, + bash: createBashToolDefinition, + edit: createEditToolDefinition, + write: createWriteToolDefinition, + grep: createGrepToolDefinition, + find: createFindToolDefinition, + ls: createLsToolDefinition, +}; initTheme("dark"); setCapabilities({ images: null, trueColor: true, hyperlinks: false }); @@ -889,7 +980,7 @@ const renderUi = { requestRender() {} }; const rows = []; for (const [name, args, result] of cases) { const wrapped = tools.find((tool) => tool.name === name); - const baseline = new ToolExecutionComponent(name, `baseline-${name}`, args, { showImages: false }, undefined, renderUi, process.cwd()); + const baseline = new ToolExecutionComponent(name, `baseline-${name}`, args, { showImages: false }, stockDefinitions[name](process.cwd()), renderUi, process.cwd()); const actual = new ToolExecutionComponent(name, `wrapped-${name}`, args, { showImages: false }, wrapped, renderUi, process.cwd()); for (const row of [baseline, actual]) { row.markExecutionStarted(); @@ -1341,7 +1432,6 @@ for (const reason of ["startup", "new", "resume", "fork", "reload"]) { await calmCommand.handler("", commandContext); const readWrapper = tools.find((tool) => tool.name === "read"); -const { createReadToolDefinition } = await import(pathToFileURL(`${packageRoot}/dist/index.js`).href); const originalRead = createReadToolDefinition(process.cwd()); const executeContext = { cwd: process.cwd() }; const [originalResult, wrappedResult] = await Promise.all([ @@ -1819,7 +1909,18 @@ TS fail "Pi follow-up $label case did not process the monitoring notification" fi - pane=$(tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" -S - 2>/dev/null || true) + # The session file is written before the TUI repaints, so wait for the + # rendered rows themselves instead of capturing the pane right away. + i=0 + while [ "$i" -lt 240 ]; do + pane=$(tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" -S - 2>/dev/null || true) + if printf '%s\n' "$pane" | grep -Fq "CAPTAIN_ANSWER_$label" && + printf '%s\n' "$pane" | grep -Fq "MONITOR_HANDLED_${label}_ONE"; then + break + fi + sleep 0.05 + i=$((i + 1)) + done [ "$(printf '%s\n' "$pane" | grep -Fc "CAPTAIN_ANSWER_$label" || true)" -eq 1 ] \ || fail "Pi follow-up $label case rendered a duplicate captain answer" assert_contains "$pane" "CAPTAIN_PROMPT_$label" "Pi follow-up $label case hid the genuine captain prompt" @@ -2137,7 +2238,9 @@ TS i=0 while [ "$i" -lt 120 ]; do capture_geometry_viewport "$snapshot" - tail -12 "$snapshot" | grep -Fq "Working..." || break + # Pi <=0.84 rendered a "Working..." transcript row; Pi >=0.85 embeds the + # indicator in the editor border as "Working". Match either spelling. + tail -12 "$snapshot" | grep -Eq "Working(\\.\\.\\.)?([[:space:]]|─|$)" || break sleep 0.05 i=$((i + 1)) done @@ -3081,8 +3184,105 @@ JS pass "Pi Calm working ship moves on a slow independent cadence over faster fixed-cell blue water, paints the complete boat standard yellow with balanced resets, keeps ANSI-stripped width exact, flips the directional sail on the exact bounce at both edges and every width, clamps visible and hidden resizes, falls back deterministically when narrow, freezes and resumes column/direction across settle/start without hidden-time jumps or duplicate timers, resets only on a fresh session, and installs and removes one scheduler-owning widget across starts, settle, abort, failure, shutdown, reload, replacement, and Calm toggles while leaving Calm-off visibility untouched" } +# The rendered-DOM assertions below depend on a real browser, so the render step +# itself is the part that fails for reasons that have nothing to do with Calm. +# This pins that guard with real processes and no browser: one clean render, one +# that only succeeds after Chrome's start-up flake, and one that never renders +# and must report enough to tell a Chrome failure apart from a Pi export change. +test_export_dom_render_guard() { + local dir source_file out_file report + + dir="$TMP_ROOT/render-guard" + mkdir -p "$dir" + source_file="$dir/export.html" + out_file="$dir/dom.html" + printf '<html><body>export</body></html>\n' >"$source_file" + + cat >"$dir/chrome-ok" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +printf '<html><head></head><body>export</body></html>\n' +SH + cat >"$dir/chrome-flaky" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +if [ "$(wc -l <"$FM_FAKE_CHROME_ATTEMPTS")" -lt 3 ]; then + echo "fake chrome start-up crashed" >&2 + exit 1 +fi +printf '<html><head></head><body>export</body></html>\n' +SH + cat >"$dir/chrome-broken" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +echo "FAKE_CHROME_STARTUP_MARKER" >&2 +exit 9 +SH + cat >"$dir/chrome-hang" <<'SH' +#!/bin/sh +case "${1:-}" in --version) echo "FakeChrome 1.2.3"; exit 0 ;; esac +echo attempt >>"$FM_FAKE_CHROME_ATTEMPTS" +printf '<html><head></head><body>export' +exec sleep 30 +SH + chmod +x "$dir/chrome-ok" "$dir/chrome-flaky" "$dir/chrome-broken" "$dir/chrome-hang" + + : >"$dir/attempts-ok" + FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-ok" \ + render_export_dom "$dir/chrome-ok" "$source_file" "$out_file" 9.9.9 >"$dir/report-ok" \ + || fail "render_export_dom rejected a Chrome that dumped a complete DOM" + grep -Fq '</html>' "$out_file" || fail "render_export_dom did not leave the rendered DOM behind" + [ "$(wc -l <"$dir/attempts-ok")" -eq 1 ] \ + || fail "render_export_dom retried a Chrome that had already rendered the DOM" + [ ! -s "$dir/report-ok" ] || fail "render_export_dom reported a diagnostic for a successful render" + + : >"$dir/attempts-flaky" + : >"$out_file" + FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-flaky" \ + render_export_dom "$dir/chrome-flaky" "$source_file" "$out_file" 9.9.9 >"$dir/report-flaky" \ + || fail "render_export_dom gave up on a Chrome that renders after a start-up failure" + grep -Fq '</html>' "$out_file" || fail "a retried render left no DOM behind" + [ "$(wc -l <"$dir/attempts-flaky")" -eq 3 ] \ + || fail "render_export_dom did not retry the failed Chrome start-ups exactly" + + : >"$dir/attempts-broken" + : >"$out_file" + if FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-broken" \ + render_export_dom "$dir/chrome-broken" "$source_file" "$out_file" 9.9.9 >"$dir/report-broken" + then + fail "render_export_dom accepted a Chrome that never rendered the DOM" + fi + [ "$(wc -l <"$dir/attempts-broken")" -eq 3 ] \ + || fail "render_export_dom did not exhaust its bounded retries before failing" + report=$(cat "$dir/report-broken") + assert_contains "$report" "$dir/chrome-broken" "the render failure did not name the Chrome binary it used" + assert_contains "$report" "FakeChrome 1.2.3" "the render failure did not name the Chrome version it used" + assert_contains "$report" "pi=9.9.9" "the render failure did not name the installed Pi version" + assert_contains "$report" "exit=9" "the render failure did not report Chrome's exit status" + assert_contains "$report" "timed_out=no" "the render failure did not report that Chrome exited on its own" + assert_contains "$report" "FAKE_CHROME_STARTUP_MARKER" "the render failure discarded Chrome's own diagnostic" + + : >"$dir/attempts-hang" + : >"$out_file" + if FM_FAKE_CHROME_ATTEMPTS="$dir/attempts-hang" FM_CHROME_RENDER_WAIT_TICKS=3 \ + render_export_dom "$dir/chrome-hang" "$source_file" "$out_file" 9.9.9 >"$dir/report-hang" + then + fail "render_export_dom accepted a Chrome that never finished the DOM" + fi + [ "$(wc -l <"$dir/attempts-hang")" -eq 3 ] \ + || fail "render_export_dom did not exhaust its bounded retries on a Chrome that never finished" + report=$(cat "$dir/report-hang") + assert_contains "$report" "timed_out=yes" \ + "the render failure reported its own kill signal without saying the attempt was timed out" + + pass "the rendered-export-DOM guard renders in one pass, retries a bounded number of Chrome start-up failures, and reports the Chrome binary, Chrome version, Pi version, exit status, and Chrome diagnostic when every attempt fails" +} + test_interactive_terminal_e2e() { - local project config home session_file export_file export_dom default_snapshot expanded_snapshot hidden_snapshot active_before_snapshot active_hidden_snapshot export_snapshot export_settled_snapshot restored_snapshot working_snapshot working_response_snapshot restarted_snapshot resumed_restored_snapshot hash_before hash_after now version chrome chrome_pid chrome_wait chrome_reap_wait active_wait active_screen_wait boat_frame_one boat_frame_two boat_resized_snapshot boat_focus_snapshot boat_cleared_snapshot boat_hull_line boat_sail_line boat_column_one boat_column_two boat_line boat_color_snapshot boat_color_line boat_water_snapshot boat_water_line boat_water_first boat_water_changed boat_narrow_snapshot boat_narrow_sails boat_freeze_snapshot boat_resume_snapshot boat_freeze_column boat_freeze_sail boat_resume_column boat_resume_sail + local project config home session_file export_file export_dom default_snapshot expanded_snapshot hidden_snapshot active_before_snapshot active_hidden_snapshot export_snapshot export_settled_snapshot restored_snapshot working_snapshot working_response_snapshot restarted_snapshot resumed_restored_snapshot hash_before hash_after now version chrome chrome_report active_wait active_screen_wait boat_frame_one boat_frame_two boat_resized_snapshot boat_focus_snapshot boat_cleared_snapshot boat_hull_line boat_sail_line boat_column_one boat_column_two boat_line boat_color_snapshot boat_color_line boat_water_snapshot boat_water_line boat_water_first boat_water_changed boat_narrow_snapshot boat_narrow_sails boat_freeze_snapshot boat_resume_snapshot boat_freeze_column boat_freeze_sail boat_resume_column boat_resume_sail if ! command -v pi >/dev/null 2>&1 || ! command -v tmux >/dev/null 2>&1; then echo "skip: pi or tmux not found for Pi calm interactive E2E" return 0 @@ -3529,36 +3729,10 @@ if (!serialized.includes("firstmate-synthetic-input") || !serialized.includes("/ const synthetic = entries.find((entry) => entry.type === "custom_message" && entry.customType === "firstmate-synthetic-input"); if (!synthetic || synthetic.display) process.exit(1); JS - chrome=$(find_chrome) || fail "Chrome or Chromium is required for rendered export DOM assertions" - "$chrome" \ - --headless=new \ - --disable-gpu \ - --no-sandbox \ - --user-data-dir="$TMP_ROOT/chrome-profile" \ - --virtual-time-budget=2000 \ - --dump-dom \ - "file://$export_file" >"$export_dom" 2>/dev/null & - chrome_pid=$! - chrome_wait=0 - while kill -0 "$chrome_pid" 2>/dev/null && [ "$chrome_wait" -lt 100 ]; do - grep -Fq '</html>' "$export_dom" 2>/dev/null && break - sleep 0.1 - chrome_wait=$((chrome_wait + 1)) - done - kill "$chrome_pid" 2>/dev/null || true - # Chrome can retain --headless=new after --dump-dom completes and ignore TERM, - # so an unbounded wait can hang after the complete DOM has been captured. - chrome_reap_wait=0 - while kill -0 "$chrome_pid" 2>/dev/null && [ "$chrome_reap_wait" -lt 20 ]; do - sleep 0.1 - chrome_reap_wait=$((chrome_reap_wait + 1)) - done - if kill -0 "$chrome_pid" 2>/dev/null; then - kill -9 "$chrome_pid" 2>/dev/null || true - fi - wait "$chrome_pid" 2>/dev/null || true - grep -Fq '</html>' "$export_dom" 2>/dev/null \ - || fail "could not render calm-mode HTML export DOM" + chrome=$(find_chrome) \ + || fail "Chrome or Chromium is required for rendered export DOM assertions; set FM_CHROME_BIN to one" + chrome_report=$(render_export_dom "$chrome" "$export_file" "$export_dom" "$version") \ + || fail "could not render calm-mode HTML export DOM: $chrome_report" node - "$export_dom" <<'JS' || fail "rendered export DOM violated the Calm conversation boundary" const dom = require("node:fs").readFileSync(process.argv[2], "utf8"); const messages = dom.match(/<div id="messages">([\s\S]*?)<\/main>/)?.[1]; @@ -3664,7 +3838,7 @@ JS done cp "$working_snapshot" "$boat_frame_one" assert_contains "$(cat "$boat_frame_one")" '\__/' "Calm did not show the working ship during a real provider wait" - assert_not_contains "$(cat "$boat_frame_one")" "Working..." "Calm left Pi's stock working row visible while the ship was shown" + assert_not_contains "$(cat "$boat_frame_one")" "Working" "Calm left Pi's stock working row visible while the ship was shown" assert_not_contains "$(cat "$boat_frame_one")" "calm transcript" "the real provider wait showed a persistent Calm status row" assert_not_contains "$(cat "$boat_frame_one")" "FIRSTMATE WATCHER WAKE: signal: /tmp/probe.status" "the real provider wait restored a hidden operational row" boat_hull_line=$(grep -F '\__/' "$boat_frame_one" | head -1) @@ -3828,13 +4002,19 @@ JS || fail "freeze frame never left the left edge (column '${boat_freeze_column:-empty}')" # Escape aborts the run, and the abort path removes the ship with no residue. + # Escape can land while the just-started run is not yet abortable, so retry + # until pi records the abort instead of assuming one keypress sufficed. tmux -L "$TMUX_SOCKET" send-keys -t "$TMUX_SESSION" Escape active_screen_wait=0 while [ "$active_screen_wait" -lt 200 ]; do - tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" >"$boat_cleared_snapshot" - if ! grep -Fq '\__/' "$boat_cleared_snapshot"; then + tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" -S -600 >"$boat_cleared_snapshot" + if ! grep -Fq '\__/' "$boat_cleared_snapshot" && + [ "$(grep -Fc 'Operation aborted' "$boat_cleared_snapshot" || true)" -ge 1 ]; then break fi + if [ "$((active_screen_wait % 20))" -eq 19 ]; then + tmux -L "$TMUX_SOCKET" send-keys -t "$TMUX_SESSION" Escape + fi sleep 0.05 active_screen_wait=$((active_screen_wait + 1)) done @@ -3870,17 +4050,23 @@ JS || fail "the second working period reset the boat from column $boat_freeze_column to $boat_resume_column instead of resuming" [ "$boat_resume_sail" = "$boat_freeze_sail" ] \ || fail "the second working period changed sail from $boat_freeze_sail to $boat_resume_sail" - assert_not_contains "$(cat "$boat_resume_snapshot")" "Working..." \ + assert_not_contains "$(cat "$boat_resume_snapshot")" "Working" \ "the second working period left Pi's stock working row visible" - # Clear the resumed run before the Calm-off stock-row probe. + # Clear the resumed run before the Calm-off stock-row probe, with the same + # abort-recorded retry as the first boat so a swallowed Escape cannot leave + # the long-delay run occupying the agent. tmux -L "$TMUX_SOCKET" send-keys -t "$TMUX_SESSION" Escape active_screen_wait=0 while [ "$active_screen_wait" -lt 200 ]; do - tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" >"$boat_cleared_snapshot" - if ! grep -Fq '\__/' "$boat_cleared_snapshot"; then + tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" -S -600 >"$boat_cleared_snapshot" + if ! grep -Fq '\__/' "$boat_cleared_snapshot" && + [ "$(grep -Fc 'Operation aborted' "$boat_cleared_snapshot" || true)" -ge 2 ]; then break fi + if [ "$((active_screen_wait % 20))" -eq 19 ]; then + tmux -L "$TMUX_SOCKET" send-keys -t "$TMUX_SESSION" Escape + fi sleep 0.05 active_screen_wait=$((active_screen_wait + 1)) done @@ -3903,13 +4089,13 @@ JS active_screen_wait=0 while [ "$active_screen_wait" -lt 200 ]; do tmux -L "$TMUX_SOCKET" capture-pane -p -t "$TMUX_SESSION" >"$working_snapshot" - if grep -Fq "Working..." "$working_snapshot"; then + if grep -Eq "Working(\\.\\.\\.)?([[:space:]]|─|$)" "$working_snapshot"; then break fi sleep 0.025 active_screen_wait=$((active_screen_wait + 1)) done - assert_contains "$(cat "$working_snapshot")" "Working..." "Calm off did not keep Pi's stock working row" + assert_contains "$(cat "$working_snapshot")" "Working" "Calm off did not keep Pi's stock working row" assert_not_contains "$(cat "$working_snapshot")" '\__/' "Calm off showed the working ship" wait_for_text "$working_response_snapshot" "CALM_WORKING_E2E_RESPONSE" \ || fail "the deterministic provider did not settle after proving Pi's stock working row" @@ -3980,4 +4166,5 @@ test_calm_mid_turn_working_notes test_operational_followup_turn_e2e test_hidden_block_geometry_e2e test_working_ship_geometry_and_lifecycle +test_export_dom_render_guard test_interactive_terminal_e2e diff --git a/tests/fm-captain-hold-lifecycle.test.sh b/tests/fm-captain-hold-lifecycle.test.sh index 87b391f5305..99e02ea4178 100755 --- a/tests/fm-captain-hold-lifecycle.test.sh +++ b/tests/fm-captain-hold-lifecycle.test.sh @@ -72,6 +72,15 @@ run_captain() { # <home> <command args...> FM_CONFIG_OVERRIDE="$home/config" "$ROOT/bin/fm-captain-hold.sh" "$@" } +request_reconciles() { # <home> <source-id> <task-id>... + local home=$1 source_id=$2 id + shift 2 + run_captain "$home" bind "$source_id" >/dev/null || return 1 + for id in "$@"; do printf '%s\n' "$id"; done \ + | run_captain "$home" reconcile-requests --source-id "$source_id" \ + --source "captured board result" >/dev/null +} + sha256_file() { # <path> if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}' @@ -104,6 +113,432 @@ write_origin_meta() { # <home> <id> [kind] "spawn_gen=fixture-$id" } +# --- markdown-to-beads migration resolution ---------------------------------- +# +# A home on the Beads backend no longer carries the legacy markdown ids a scout +# report attested: fm-hold-migration rehomed each held row under a prefixed fm- +# id and recorded its markdown identity in the row's notes as the exact line +# "migrated from data/backlog.md id <legacy id>". The fixture graph is driven +# through bd directly where possible because the npm-published tasks-axi ships +# the markdown backend only; the hold itself needs a beads-capable tasks-axi, +# so that family probes once and skips itself with an explicit reason on +# markdown-only installs, mirroring tests/fm-control-relaunch.test.sh. + +# Build a fixture home whose configured backend is a scratch Beads graph. +# Echoes "<home>|<graph-beads-dir>". The graph repo dir is named "fm" because +# bd derives the row-id prefix from the repo directory name. +make_beads_home() { # <name> + local name=$1 case_dir home graph fb + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + graph="$case_dir/fm" + mkdir -p "$home/data" "$home/config" "$home/projects" "$graph" + (umask 077; mkdir -p "$home/state") + git -C "$graph" init -q + if ! (cd "$graph" && bd init >"$case_dir/bd-init.log" 2>&1); then + cat "$case_dir/bd-init.log" >&2 + fail "fixture bd init failed on $graph" + fi + cat > "$home/.tasks.toml" <<EOF +backend = "beads" + +[beads] +path = "$graph/.beads" +binary = "bd" +prefix = "fm" + +[markdown] +path = "data/backlog.md" +archive = "data/done-archive.md" +done_keep = 10 +EOF + fb=$(fm_fakebin "$home") + fm_fake_exit0 "$fb" tmux treehouse no-mistakes gh gh-axi + printf '%s\n' "$home|$graph/.beads" +} + +bdrow() { # <beads-dir> <args...> + BEADS_DIR="$1" bd "${@:2}" +} + +# One capability probe for the migration family: can the installed tasks-axi +# operate on a beads-backed home? The npm-published tasks-axi cannot, and the +# hold fixture needs its beads backend; those installs skip with this reason. +probe_tasks_axi_beads() { + local probe_home="$TMP_ROOT/.probe" probe_graph="$TMP_ROOT/.probe-fm" + rm -rf "$probe_home" "$probe_graph" + mkdir -p "$probe_home/data" "$probe_graph" + git -C "$probe_graph" init -q + (cd "$probe_graph" && bd init >/dev/null 2>&1) || return 1 + cat > "$probe_home/.tasks.toml" <<PROBEEOF +backend = "beads" + +[beads] +path = "$probe_graph/.beads" +binary = "bd" +prefix = "fm" +PROBEEOF + (cd "$probe_home" && tasks-axi list) >/dev/null 2>&1 +} +TASKS_AXI_BEADS_OK=0 +if bd --version >/dev/null 2>&1 && jq --version >/dev/null 2>&1 \ + && probe_tasks_axi_beads; then + TASKS_AXI_BEADS_OK=1 +fi + +require_tasks_axi_beads() { # <what> + [ "$TASKS_AXI_BEADS_OK" = 1 ] && return 0 + pass "skipped on markdown-only tasks-axi: $1" + return 1 +} + +write_scout_with_attested_inventory() { # <home> <scout-id> <keys> + local home=$1 scout=$2 keys=$3 + mkdir -p "$home/data/$scout" + fm_write_meta "$home/state/$scout.meta" \ + "window=firstmate:fm-$scout" \ + "worktree=$home/projects/missing-$scout" \ + "project=$home/projects/sample" \ + "harness=codex" \ + "kind=scout" \ + "spawn_gen=fixture-$scout" \ + "decisions_reviewed=1" \ + "decision_keys=$keys" + printf 'done: report complete\n' > "$home/state/$scout.status" + printf '# Report\n\nThe investigation finished.\n' > "$home/data/$scout/report.md" +} + +test_verify_resolves_a_hold_migrated_to_beads_notes() { + local fixture home beads scout + require_tasks_axi_beads "verify against a beads-migrated hold" || return 0 + fixture=$(make_beads_home migrated-notes) + home=${fixture%%|*} + beads=${fixture##*|} + scout=sample-beads-scout + (cd "$home" && BEADS_ACTOR=fixture tasks-axi add fm-herald-github-delete \ + "Delete the herald repo" --repo herald) >/dev/null 2>&1 \ + || fail "could not create the migrated hold fixture" + (cd "$home" && BEADS_ACTOR=fixture tasks-axi hold fm-herald-github-delete \ + --kind captain --reason "captain must confirm the delete") >/dev/null 2>&1 \ + || fail "could not hold the migrated fixture row" + bdrow "$beads" note fm-herald-github-delete \ + "Origin: herald-retire +Decision key: github-delete +State: awaiting captain decision. + +migrated from data/backlog.md id herald-retire-decision-github-delete on 2026-09-04" \ + >/dev/null 2>&1 || fail "could not record the migration marker note" + write_scout_with_attested_inventory "$home" "$scout" \ + herald-retire-decision-github-delete + + run_captain "$home" verify "$scout" >/dev/null \ + || fail "verify did not resolve the attested legacy id through its migrated beads row" + pass "verify resolves a captain hold migrated to a beads row with a marker note" +} + +# A tasks-axi stub that knows ONLY the row ids it is given. The real +# beads-capable tasks-axi resolves a bare legacy id onto its prefixed row +# itself, answering before any migration lookup runs; against this stub the +# migrated-hold resolution order is what has to answer. +write_known_rows_stub() { # <fakebin> <row-id...> + local fb=$1 known + shift + known=$(printf '%s|' "$@") + cat > "$fb/tasks-axi" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + --version) printf '%s\n' '0.2.5' ;; + update) + [ "${2:-}" = --help ] || exit 1 + printf '%s\n' '--archive-body' + ;; + mv) + [ "${2:-}" = --help ] || exit 1 + printf '%s\n' 'usage: tasks-axi mv [<id>...]' + ;; + hold) + [ "${2:-}" = --help ] || exit 1 + printf '%s\n' ' --kind captain' + ;; + show) + case "${2:-}" in + @KNOWN@) ;; + *) printf 'error: no task %s in this backlog\n' "${2:-}" >&2; exit 1 ;; + esac + printf '%s\n' 'task:' + printf ' id: %s\n' "$2" + printf '%s\n' ' state: queued' ' held: yes' ' blocked: no' \ + ' hold_kind: captain' ' body: ""' + ;; + *) exit 1 ;; +esac +SH + sed -i.bak "s%@KNOWN@%${known%|}%" "$fb/tasks-axi" + rm -f "$fb/tasks-axi.bak" + chmod +x "$fb/tasks-axi" +} + +test_verify_resolves_a_hold_migrated_under_the_configured_prefix() { + local fixture home scout out + require_tasks_axi_beads "verify against a prefix-migrated hold" || return 0 + fixture=$(make_beads_home migrated-prefix) + home=${fixture%%|*} + scout=sample-prefix-scout + (cd "$home" && BEADS_ACTOR=fixture tasks-axi add fm-other-legacy-row \ + "Second migrated hold" --repo sample) >/dev/null 2>&1 \ + || fail "could not create the prefix-migrated fixture row" + (cd "$home" && BEADS_ACTOR=fixture tasks-axi hold fm-other-legacy-row \ + --kind captain --reason "captain must decide") >/dev/null 2>&1 \ + || fail "could not hold the prefix-migrated fixture row" + # No row in this graph carries a marker note, so the narrowed prefix guess is + # the only resolution left for the attested legacy id. + write_known_rows_stub "$(fm_fakebin "$home")" fm-other-legacy-row + write_scout_with_attested_inventory "$home" "$scout" other-legacy-row + + run_captain "$home" verify "$scout" >/dev/null \ + || fail "verify did not resolve the legacy id under the configured prefix" + out=$(run_captain "$home" complete "$scout" other-legacy-row) \ + || fail "the completion gate refused the prefix-resolved hold" + assert_contains "$out" "other-legacy-row=fm-other-legacy-row" \ + "completion did not name the row the prefix guess attested against" + pass "verify resolves a captain hold whose id is the legacy id under the configured prefix" +} + +# The prefix guess is a name, not evidence: when a row actually carries the +# migration marker, that row is the one attested even though an unrelated +# captain-held task occupies the <prefix>-<legacy id> name. +test_marker_noted_row_wins_over_a_prefix_namesake() { + local fixture home beads scout row out + require_tasks_axi_beads "prefer a marker-noted row over a prefix namesake" || return 0 + fixture=$(make_beads_home migrated-marker-wins) + home=${fixture%%|*} + beads=${fixture##*|} + scout=sample-marker-wins-scout + for row in fm-dual-row fm-marked-dual-row fm-solo-row; do + (cd "$home" && BEADS_ACTOR=fixture tasks-axi add "$row" "Captain call $row" \ + --repo sample) >/dev/null 2>&1 || fail "could not create the fixture row $row" + (cd "$home" && BEADS_ACTOR=fixture tasks-axi hold "$row" --kind captain \ + --reason "captain must decide") >/dev/null 2>&1 \ + || fail "could not hold the fixture row $row" + done + bdrow "$beads" note fm-marked-dual-row \ + "migrated from data/backlog.md id dual-row on 2026-09-04" \ + >/dev/null 2>&1 || fail "could not record the migration marker note" + write_known_rows_stub "$(fm_fakebin "$home")" \ + fm-dual-row fm-marked-dual-row fm-solo-row + write_scout_with_attested_inventory "$home" "$scout" "dual-row,solo-row" + + out=$(run_captain "$home" complete "$scout" dual-row solo-row) \ + || fail "the completion gate refused an inventory carrying both migrated shapes" + assert_not_contains "$out" "dual-row=fm-dual-row" \ + "the bare prefix namesake shadowed the row carrying the migration marker" + assert_contains "$out" "solo-row=fm-solo-row" \ + "completion did not name the row the prefix guess attested against" + run_captain "$home" verify "$scout" >/dev/null \ + || fail "verify did not re-resolve both migrated shapes after completion" + pass "the marker-noted row wins over an unrelated row holding the prefix namesake" +} + +test_complete_accepts_a_migrated_inventory_on_beads() { + local fixture home scout + require_tasks_axi_beads "complete against a beads-migrated hold" || return 0 + fixture=$(make_beads_home migrated-complete) + home=${fixture%%|*} + beads=${fixture##*|} + scout=sample-complete-scout + (cd "$home" && BEADS_ACTOR=fixture tasks-axi add fm-herald-github-delete \ + "Delete the herald repo" --repo herald) >/dev/null 2>&1 \ + || fail "could not create the migrated hold fixture" + (cd "$home" && BEADS_ACTOR=fixture tasks-axi hold fm-herald-github-delete \ + --kind captain --reason "captain must confirm the delete") >/dev/null 2>&1 \ + || fail "could not hold the migrated fixture row" + bdrow "$beads" note fm-herald-github-delete \ + "migrated from data/backlog.md id herald-retire-decision-github-delete on 2026-09-04" \ + >/dev/null 2>&1 || fail "could not record the migration marker note" + fm_write_meta "$home/state/$scout.meta" \ + "window=firstmate:fm-$scout" \ + "worktree=$home/projects/missing-$scout" \ + "project=$home/projects/sample" \ + "harness=codex" \ + "kind=scout" \ + "spawn_gen=fixture-$scout" + printf 'done: report complete\n' > "$home/state/$scout.status" + mkdir -p "$home/data/$scout" + printf '# Report\n\nThe investigation finished.\n' > "$home/data/$scout/report.md" + + run_captain "$home" complete "$scout" herald-retire-decision-github-delete >/dev/null \ + || fail "the completion gate refused an inventory resolved through a migrated beads row" + assert_grep "decision_keys=herald-retire-decision-github-delete" \ + "$home/state/$scout.meta" \ + "the attestation did not record the attested legacy id" + run_captain "$home" verify "$scout" >/dev/null \ + || fail "verify did not re-resolve the attested legacy id after completion" + pass "the completion gate attests an inventory resolved through a migrated beads row" +} + +test_verify_names_the_unresolvable_legacy_id_once() { + local fixture home scout err rc + require_tasks_axi_beads "verify an unresolvable beads legacy id" || return 0 + fixture=$(make_beads_home migrated-absent) + home=${fixture%%|*} + scout=sample-absent-scout + write_scout_with_attested_inventory "$home" "$scout" ghost-legacy-id + + rc=0 + err=$(run_captain "$home" verify "$scout" 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "verify accepted an attested id that resolves to nothing" + assert_contains "$err" "ghost-legacy-id" \ + "the refusal did not name the id it could not resolve" + [ "$(printf '%s\n' "$err" | grep -c '^fm-captain-hold:')" = 1 ] \ + || fail "the refusal emitted more than one line: $err" + pass "an unresolvable legacy id is refused once, naming the id" +} + +test_verify_resolves_a_pre_collapse_key_through_its_derived_marker() { + local fixture home beads scout + require_tasks_axi_beads "verify a derived pre-collapse key" || return 0 + fixture=$(make_beads_home migrated-derived) + home=${fixture%%|*} + beads=${fixture##*|} + scout=sample-derived-scout + # The row was migrated under the DERIVED pre-collapse identity, keyed by a + # bare decision key the origin's old metadata attests. + (cd "$home" && BEADS_ACTOR=fixture tasks-axi add fm-dec-call \ + "Migrated pre-collapse call" --repo sample) >/dev/null 2>&1 \ + || fail "could not create the derived-marker fixture row" + (cd "$home" && BEADS_ACTOR=fixture tasks-axi hold fm-dec-call \ + --kind captain --reason "captain must decide") >/dev/null 2>&1 \ + || fail "could not hold the derived-marker fixture row" + bdrow "$beads" note fm-dec-call \ + "migrated from data/backlog.md id $scout-decision-github-delete on 2026-09-04" \ + >/dev/null 2>&1 || fail "could not record the derived-id marker note" + write_scout_with_attested_inventory "$home" "$scout" github-delete + + run_captain "$home" verify "$scout" >/dev/null \ + || fail "verify did not probe the derived pre-collapse identity for its migrated row" + pass "a pre-collapse key resolves through its derived identity's migration marker" +} + +# The captain-hold mutation wrapper must address the configured backend like +# the transition library does: on a beads-configured home its hold/answer/done +# calls reach tasks-axi with no markdown file override. Fully portable - the +# stubbed tasks-axi fakes the beads backend, so no bd or beads-capable install +# is needed. +test_captain_hold_mutations_address_the_beads_backend() { + local home id fb log + home="$TMP_ROOT/captain-stub-beads/home" + mkdir -p "$home/data" "$home/config" "$home/projects" "$home/state" + cat > "$home/.tasks.toml" <<'EOF' +backend = "beads" + +[beads] +path = "graph/.beads" +binary = "bd" +prefix = "fm" +EOF + id=fm-stub-held-row + fb=$(fm_fakebin "$home") + log="$home/tasks-axi-calls" + cat > "$fb/tasks-axi" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$*" >> "@LOG@" +case "${1:-}" in + --version) printf '%s\n' '0.2.5' ;; + update) + if [ "${2:-}" = --help ]; then + printf '%s\n' '--archive-body' + exit 0 + fi + case " $* " in + *" --file "*) + printf '%s\n' 'error: beads update received a markdown file override' >&2 + exit 1 + ;; + esac + stub_prev= + stub_path= + for stub_arg in "$@"; do + if [ "$stub_prev" = --body-file ]; then stub_path=$stub_arg; fi + stub_prev=$stub_arg + done + [ -n "$stub_path" ] && cp -- "$stub_path" "@HOME@/last-body" + printf 'ok: update %s\n' "${2:-}" + ;; + mv) + [ "${2:-}" = --help ] || exit 1 + printf '%s\n' 'usage: tasks-axi mv [<id>...]' + ;; + hold) + case " $* " in + *" --help "*) printf '%s\n' ' --kind captain' ; exit 0 ;; + *" --file "*) + printf '%s\n' 'error: beads hold received a markdown file override' >&2 + exit 1 + ;; + esac + printf 'ok: hold %s\n' "${2:-}" + ;; + done) + [ "${2:-}" = "@ID@" ] || exit 1 + case " $* " in + *" --file "*) + printf '%s\n' 'error: beads done received a markdown file override' >&2 + exit 1 + ;; + esac + printf 'ok: done %s\n' "${2:-}" + ;; + show) + [ "${2:-}" = "@ID@" ] || exit 1 + case " $* " in + *" --file "*) + printf '%s\n' 'error: beads show received a markdown file override' >&2 + exit 1 + ;; + esac + printf '%s\n' 'task:' + printf ' id: %s\n' "@ID@" + printf '%s\n' ' state: queued' ' held: yes' ' blocked: no' \ + ' hold_kind: captain' + if [ -f "@HOME@/last-body" ]; then + printf '%s' ' body: ' + perl -MJSON::PP -e 'local $/; print encode_json(<STDIN>)' < "@HOME@/last-body" + printf '\n' + else + printf '%s\n' ' body: ""' + fi + ;; + *) exit 1 ;; +esac +SH + sed -i.bak "s|@HOME@|$home|g; s|@ID@|$id|g; s|@LOG@|$log|g" "$fb/tasks-axi" + rm -f "$fb/tasks-axi.bak" + chmod +x "$fb/tasks-axi" + + PATH="$fb:$PATH" REAL_TASKS_AXI="$TASKS_AXI_BIN" \ + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + FM_DATA_OVERRIDE="$home/data" FM_CONFIG_OVERRIDE="$home/config" \ + "$ROOT/bin/fm-captain-hold.sh" hold "$id" --reason "captain must decide" >/dev/null \ + || fail "holding on a beads-configured home failed without a markdown backlog" + assert_grep "hold $id" "$log" \ + "the captain-hold mutation never reached the configured backend" + + decision="$home/captain-decision.txt" + printf 'Ship the gold-only plan.\n' > "$decision" + PATH="$fb:$PATH" REAL_TASKS_AXI="$TASKS_AXI_BIN" \ + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + FM_DATA_OVERRIDE="$home/data" FM_CONFIG_OVERRIDE="$home/config" \ + "$ROOT/bin/fm-captain-hold.sh" answer "$id" --decision-file "$decision" >/dev/null \ + || fail "answering on a beads-configured home failed without a markdown backlog" + assert_grep "update $id --body-file" "$log" \ + "the captain answer never reached the configured backend" + assert_grep "done $id" "$log" \ + "the captain answer close never reached the configured backend" + assert_no_grep " --file " "$log" \ + "a captain-hold mutation passed a markdown file override to a beads home" + pass "captain-hold mutations address the beads backend without a markdown override" +} + # Reproduces the loss exactly with privacy-safe synthetic names: the investigation # and visual review have ended, the only genuine unresolved captain call is report # prose, no held backlog item or open status exists, and the authoritative @@ -893,6 +1328,76 @@ EOF pass "a secondmate home publishes each hold occurrence and its answer on the parent channel" } +test_secondmate_reconcile_publishes_before_request_retirement() { + local parent mate channel evidence out show rc request + parent=$(make_home reconcile-parent-channel) + mate=$(make_home reconcile-channel-mate) + printf 'reconcile-channel-mate\n' > "$mate/.fm-secondmate-home" + printf 'schema=fm-secondmate-parent.v1\nroute=local\nparent_home=%s\n' "$parent" \ + > "$mate/.fm-secondmate-parent" + channel="$parent/state/reconcile-channel-mate.status" + evidence="$mate/reconcile-evidence.txt" + + tasks_in "$mate" add reconcile-channel-call "Verify the mate call" --kind ship --repo sample >/dev/null \ + || fail "could not create the reconcile channel call" + run_captain "$mate" hold reconcile-channel-call --reason "verify current release state" >/dev/null \ + || fail "could not hold the reconcile channel call" + request_reconciles "$mate" reconcile-board reconcile-channel-call \ + || fail "could not request the channel reconciliation" + printf 'The release has already landed.\n' > "$evidence" + request="$mate/state/reconcile-requests/reconcile-channel-call.request" + + chmod 0500 "$mate/state/reconcile-requests" + set +e + out=$(run_captain "$mate" reconcile close reconcile-channel-call \ + --evidence-file "$evidence" 2>&1) + rc=$? + set -e + chmod 0700 "$mate/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "failed reconcile request retirement reported success" + assert_contains "$out" "reconcile-channel-call" \ + "the reconcile retirement failure did not name its task: $out" + show=$(tasks_in "$mate" show reconcile-channel-call --full) + assert_contains "$show" "state: done" "request retirement failure reversed the reconciled close" + assert_contains "$show" "Resolution mode: reconciled" \ + "request retirement failure lost the reconciled resolution mode" + [ -f "$request" ] || fail "the request retired despite its forced retirement failure" + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + || fail "the parent resolution was not published before retirement failed: $(cat "$channel")" + + run_captain "$mate" reconcile close reconcile-channel-call --evidence-file "$evidence" >/dev/null \ + || fail "the closed reconciliation could not finish publication and retirement" + [ ! -e "$request" ] || fail "the retry did not retire the published reconcile request" + [ "$(grep -c 'resolved \[key=captain-hold-reconcile-channel-call-1\]: captain hold reconcile-channel-call: reconciled' "$channel")" -eq 1 ] \ + || fail "the reconciliation retry duplicated or changed its parent resolution: $(cat "$channel")" + tasks_in "$mate" add answer-channel-call "Answer the mate call" --kind ship --repo sample >/dev/null \ + || fail "could not create the normal-answer channel call" + run_captain "$mate" hold answer-channel-call --reason "captain answer needed" >/dev/null \ + || fail "could not hold the normal-answer channel call" + request_reconciles "$mate" reconcile-board answer-channel-call \ + || fail "could not create the normal-answer retry trigger" + printf 'Proceed with the release.\n' > "$mate/answer.txt" + request="$mate/state/reconcile-requests/answer-channel-call.request" + chmod 0500 "$mate/state/reconcile-requests" + set +e + out=$(run_captain "$mate" answer answer-channel-call --decision-file "$mate/answer.txt" 2>&1) + rc=$? + set -e + chmod 0700 "$mate/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "failed normal-answer request retirement reported success" + show=$(tasks_in "$mate" show answer-channel-call --full) + assert_contains "$show" "state: done" "request retirement failure reversed the captain answer" + [ -f "$request" ] || fail "the normal-answer retry trigger retired after its forced failure" + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + || fail "the normal answer did not publish before retirement failed: $(cat "$channel")" + run_captain "$mate" answer answer-channel-call --decision-file "$mate/answer.txt" >/dev/null \ + || fail "the normal-answer retry could not finish request retirement" + [ ! -e "$request" ] || fail "the normal-answer retry left its request pending" + [ "$(grep -c 'resolved \[key=captain-hold-answer-channel-call-1\]: captain hold answer-channel-call: answered' "$channel")" -eq 1 ] \ + || fail "the normal-answer retry duplicated its parent resolution: $(cat "$channel")" + pass "secondmate resolutions publish before retiring durable retry triggers" +} + # The one keyed-answer intake, fed through the real process-event runner by a # fixture channel that knows nothing about captain holds: task-id keys close at # answer time, a card-declared release mode frees held work, freeform prose can @@ -915,12 +1420,23 @@ test_bound_channel_answers_close_at_answer_time() { --reason "captain forged choice pending" --repo sample --origin "$id" >/dev/null run_captain "$home" hold sample-invalid-close-call --title "Captain call: invalid close" \ --reason "captain close mode validation pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-source-reconcile --title "Captain call: reconcile" \ + --reason "captain re-check pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-bare-reconcile --title "Captain call: bare reconcile" \ + --reason "captain bare re-check pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-shape --title "Captain call: old board shape" \ + --reason "captain old board pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-reconcile --title "Captain call: old bare reconcile" \ + --reason "captain old bare reconcile pending" --repo sample --origin "$id" >/dev/null + run_captain "$home" hold sample-old-reconcile-note --title "Captain call: old annotated reconcile" \ + --reason "captain old annotated reconcile pending" --repo sample --origin "$id" >/dev/null tasks_in "$home" add sample-gated-work "Gated sample work" --kind ship --repo sample \ --body 'Gated work plan.' >/dev/null run_captain "$home" hold sample-gated-work --reason "captain go needed" >/dev/null run_captain "$home" complete "$id" \ sample-membership-call sample-headline-call sample-forged-call sample-invalid-close-call \ - sample-gated-work >/dev/null \ + sample-source-reconcile sample-bare-reconcile sample-old-shape sample-old-reconcile \ + sample-old-reconcile-note sample-gated-work >/dev/null \ || fail "completion failed for the deck's inventoried calls" artifact="$home/data/$id/review.html" @@ -941,25 +1457,44 @@ session: status: feedback session_ended: true ended_by: user -prompts[6]{uid,prompt,selector,tag,text}: - "2","Membership: gold-only\n\nContext data:\n{\n \"question\": \"sample-membership-call\",\n \"answer\": \"gold-only\"\n}","section#call > form:nth-of-type(1)",choice,"Membership: gold-only" - "3","Headline: f1-when-fp-gold\n\nContext data:\n{\n \"question\": \"sample-headline-call\",\n \"answer\": \"f1-when-fp-gold\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" - "4","Gated work: go\n\nContext data:\n{\n \"question\": \"sample-gated-work\",\n \"answer\": \"go\",\n \"close\": \"release\"\n}","section#call > form:nth-of-type(3)",choice,"Gated work: go" - "5","Absent call: yes\n\nContext data:\n{\n \"question\": \"sample-nonexistent-call\",\n \"answer\": \"yes\"\n}","section#call > form:nth-of-type(4)",choice,"Absent call: yes" +prompts[13]{uid,prompt,selector,tag,text}: + "1","Reconcile first\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-source-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(6)",choice,"Reconcile" + "2","Membership: gold-only - captain detail\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-membership-call\",\n \"selection\": \"gold-only\",\n \"note\": \"captain detail\"\n}","section#call > form:nth-of-type(1)",choice,"Membership: gold-only - captain detail" + "3","Headline: f1-when-fp-gold\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-headline-call\",\n \"selection\": \"f1-when-fp-gold\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" + "4","Gated work: go\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-gated-work\",\n \"selection\": \"go\",\n \"note\": \"\",\n \"close\": \"release\"\n}","section#call > form:nth-of-type(3)",choice,"Gated work: go" + "5","Absent call: yes\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-nonexistent-call\",\n \"selection\": \"yes\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(4)",choice,"Absent call: yes" "6","Invalid close: yes\n\nContext data:\n{\n \"question\": \"sample-invalid-close-call\",\n \"answer\": \"yes\",\n \"close\": \"drop\"\n}","section#call > form:nth-of-type(5)",choice,"Invalid close: yes" + "7","Reconcile this - re-check latest publication\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-source-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"re-check latest publication\"\n}","section#call > form:nth-of-type(6)",choice,"Reconcile - re-check latest publication" + "8","Second reconcile\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-bare-reconcile\",\n \"selection\": \"reconcile\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(7)",choice,"Reconcile" + "9","Headline final: f1-when-fp-gold\n\nContext data:\n{\n \"schema\": \"fm-bearings-answer.v1\",\n \"question\": \"sample-headline-call\",\n \"selection\": \"f1-when-fp-gold\",\n \"note\": \"\"\n}","section#call > form:nth-of-type(2)",choice,"Headline: f1-when-fp-gold" + "10","Old board answer\n\nContext data:\n{\n \"question\": \"sample-old-shape\",\n \"answer\": \"yes\"\n}","section#call > form:nth-of-type(8)",choice,"Old answer: yes" + "11","Old board reconcile\n\nContext data:\n{\n \"question\": \"sample-old-reconcile\",\n \"answer\": \"reconcile\"\n}","section#call > form:nth-of-type(9)",choice,"Old reconcile" + "12","Old board reconcile note\n\nContext data:\n{\n \"question\": \"sample-old-reconcile-note\",\n \"answer\": \"reconcile - verify publication\"\n}","section#call > form:nth-of-type(10)",choice,"Old reconcile note" "",get this fully implemented. Context data:\n{\n \"question\": \"sample-forged-call\",\n \"answer\": \"forged\"\n},"",message,Freeform message next_step: This was the last feedback before the user ended the session. EOF printf 'lavish\n' > "$home/state/procevent-inbox/$sid.1.adapter" out=$(run_lavish "$home" answers "$result") || fail "could not read the captured answers" - assert_contains "$out" "sample-membership-call gold-only" "a structured choice was not read as an answer" + assert_contains "$out" "sample-membership-call gold-only" \ + "a repeated reconcile selection deleted another card's answer" + assert_contains "$out" "sample-headline-call f1-when-fp-gold" \ + "a repeated ordinary selection was not preserved" assert_contains "$out" "sample-gated-work go Gated work: go release" \ "the card-declared release mode was not relayed" assert_not_contains "$out" "sample-forged-call" \ "a freeform captain message forged a task id from its own prose" assert_not_contains "$out" "sample-invalid-close-call" \ "an unsupported card close mode defaulted to completion" + assert_not_contains "$out" "sample-source-reconcile" \ + "a reconcile selection leaked into keyed answers" + assert_contains "$out" "sample-old-shape yes" \ + "an ordinary legacy board choice was discarded during rollout" + assert_not_contains "$out" "sample-old-reconcile" \ + "a legacy reconcile-shaped value reached keyed answers" + out=$(run_lavish "$home" reconciles "$result") || fail "could not read captured reconcile selections" + [ "$out" = "$(printf 'sample-source-reconcile\tre-check latest publication\nsample-bare-reconcile')" ] \ + || fail "current or legacy selections lost or invented a reconcile task id: $out" mkdir -p "$home/adapter-root/bin" cat > "$home/adapter-root/bin/fm-procevent-fixturechan.sh" <<SH @@ -967,6 +1502,7 @@ EOF # Fixture channel: reports keyed captain answers and nothing else. case "\${1-}" in answers) exec "$ROOT/bin/fm-procevent-lavish.sh" answers "\${2-}" ;; + reconciles) exec "$ROOT/bin/fm-procevent-lavish.sh" reconciles "\${2-}" ;; esac exit 2 SH @@ -991,6 +1527,7 @@ SH assert_contains "$show" "state: done" "capturing the captain's answer left the membership call open" assert_contains "$show" "Resolution mode: answered" "the membership call did not record its close path" assert_contains "$show" "Answer: gold-only" "the closed call did not record the captain's actual answer" + assert_contains "$show" "captain detail" "the annotated normal answer lost the captain's note" show=$(tasks_in "$home" show sample-gated-work --full) assert_contains "$show" "state: queued" "the released work item did not stay queued" assert_contains "$show" "held: no" "the card-declared release did not lift the hold" @@ -1001,6 +1538,33 @@ SH show=$(tasks_in "$home" show sample-invalid-close-call --full) assert_contains "$show" "state: queued" "an unsupported card close mode closed a captain call" assert_contains "$show" "held: yes" "an unsupported card close mode released a captain call" + out=$(run_captain "$home" reconcile list) + assert_contains "$out" "sample-source-reconcile" \ + "the bound captured reconcile selection did not create a request" + assert_contains "$out" "captain note: re-check latest publication" \ + "the annotated reconcile selection lost its note provenance" + show=$(tasks_in "$home" show sample-old-shape --full) + assert_contains "$show" "state: done" "an ordinary legacy board choice did not close its task" + assert_contains "$show" "Resolution mode: answered" \ + "an ordinary legacy board choice did not use the keyed-answer intake" + show=$(tasks_in "$home" show sample-old-reconcile --full) + assert_contains "$show" "state: queued" "a bare legacy reconcile value closed its task" + assert_contains "$show" "held: yes" "a bare legacy reconcile value released its task" + show=$(tasks_in "$home" show sample-old-reconcile-note --full) + assert_contains "$show" "state: queued" "an annotated legacy reconcile value closed its task" + assert_contains "$show" "held: yes" "an annotated legacy reconcile value released its task" + assert_not_contains "$out" "sample-old-reconcile" \ + "a legacy reconcile value created a generationless request" + show=$(tasks_in "$home" show sample-bare-reconcile --full) + assert_contains "$show" "state: queued" "a bare captured reconcile selection closed its task" + assert_contains "$show" "held: yes" "a bare captured reconcile selection released its task" + printf 'The captured call is moot.\n' > "$home/source-reconcile-evidence.txt" + run_captain "$home" reconcile close sample-source-reconcile \ + --evidence-file "$home/source-reconcile-evidence.txt" >/dev/null \ + || fail "the annotated captured request did not authorize evidence-backed closure" + run_captain "$home" reconcile close sample-bare-reconcile \ + --evidence-file "$home/source-reconcile-evidence.txt" >/dev/null \ + || fail "the bare captured request did not authorize evidence-backed closure" # Replaying the same capture is a no-op, not a rejected different decision. A # run that could not close every answered key still reports nonzero. @@ -1023,6 +1587,10 @@ SH printf 'Captain answered the invalid-close call directly.\n' > "$home/invalid-close.txt" run_captain "$home" answer sample-invalid-close-call --decision-file "$home/invalid-close.txt" >/dev/null \ || fail "could not close the invalid-close call through the answer path" + run_captain "$home" answer sample-old-reconcile --decision-file "$home/invalid-close.txt" >/dev/null \ + || fail "could not deliberately close the bare legacy reconcile call" + run_captain "$home" answer sample-old-reconcile-note --decision-file "$home/invalid-close.txt" >/dev/null \ + || fail "could not deliberately close the annotated legacy reconcile call" run_captain "$home" verify "$id" >/dev/null \ || fail "answered calls did not satisfy the completion gate" pass "a bound channel's captured answers close their captain-held tasks at answer time" @@ -1030,6 +1598,331 @@ SH # Answer-time closure is opt-in per source. A channel with no binding must behave # exactly as it always did: capture, announce, close nothing. +# A reconcile is "go re-check reality", never the captain's answer. The value is +# reserved at the one keyed-answer intake, so no channel and no card-declared +# close mode can turn it into a close or a release, and the obligation to verify +# survives as a durable request instead of evaporating with the wake. +test_reconcile_never_closes_through_the_keyed_answer_intake() { + local home out rc show list + home=$(make_home reconcile-intake) + tasks_in "$home" add sample-reconcile-call "Captain call: still current?" --repo sample >/dev/null \ + || fail "could not create the reconcile call" + tasks_in "$home" add sample-reconcile-gated "Gated work" --repo sample >/dev/null \ + || fail "could not create the gated work item" + run_captain "$home" hold sample-reconcile-call --reason "is this still current?" >/dev/null \ + || fail "could not hold the reconcile call" + run_captain "$home" hold sample-reconcile-gated --reason "waiting on the captain" >/dev/null \ + || fail "could not hold the gated work item" + + set +e + out=$(printf 'sample-reconcile-call\treconcile\tReconcile\n%s\n' \ + "$(printf 'sample-reconcile-gated\treconcile\tReconcile\trelease')" \ + | run_captain "$home" answers --source "captain chat" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the shared answer intake accepted reconcile as an answer" + assert_contains "$out" "refused: sample-reconcile-call" \ + "the shared intake did not visibly refuse reconcile: $out" + case "$out" in + *"closed: sample-reconcile"*) fail "a reconcile row closed a captain call: $out" ;; + esac + + show=$(tasks_in "$home" show sample-reconcile-call --full) + assert_contains "$show" "state: queued" "a reconcile row completed a captain call" + assert_contains "$show" "held: yes" "a reconcile row released a captain call" + case "$show" in + *"Resolution recorded by"*) fail "a reconcile row wrote a resolution record" ;; + esac + show=$(tasks_in "$home" show sample-reconcile-gated --full) + assert_contains "$show" "held: yes" "a release-mode reconcile row lifted a captain hold" + + list=$(run_captain "$home" reconcile list) || fail "could not list the reconcile requests" + assert_contains "$list" "reconcile-requests: 0" \ + "the shared answer intake created a reconcile request: $list" + set +e + out=$(printf 'sample-reconcile-call\n' \ + | run_captain "$home" reconcile-requests --source-id unbound-src --source "captured board" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "an unbound captured source created a reconcile request" + request_reconciles "$home" board-src sample-reconcile-call sample-reconcile-gated \ + || fail "the bound captured source did not create reconcile requests" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "sample-reconcile-call" "the reconcile obligation was not recorded durably: $list" + assert_contains "$list" "reconcile-requests: 2" "the reconcile requests were not both recorded: $list" + + request_reconciles "$home" board-src sample-reconcile-call \ + || fail "replaying a captured reconcile selection failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 2" "a replayed reconcile selection duplicated the obligation: $list" + pass "only a bound captured source creates reconcile requests" +} + +test_normal_answers_retire_pending_reconcile_requests() { + local home list id + home=$(make_home reconcile-normal-answer) + for id in sample-direct-close sample-direct-release sample-keyed-close; do + tasks_in "$home" add "$id" "Captain call $id" --repo sample >/dev/null + run_captain "$home" hold "$id" --reason "waiting for the captain" >/dev/null + done + request_reconciles "$home" board-src sample-direct-close sample-direct-release sample-keyed-close \ + || fail "could not create reconcile requests before normal answers" + + printf 'Captain said close.\n' > "$home/close.txt" + printf 'Captain said release.\n' > "$home/release.txt" + run_captain "$home" answer sample-direct-close --decision-file "$home/close.txt" >/dev/null \ + || fail "a direct close answer failed" + run_captain "$home" answer sample-direct-release --decision-file "$home/release.txt" --release >/dev/null \ + || fail "a direct release answer failed" + printf 'sample-keyed-close\tyes\tYes\n' \ + | run_captain "$home" answers --source "board sequence 2" >/dev/null \ + || fail "a keyed normal answer failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" \ + "normal answers left stranded reconcile requests: $list" + + run_captain "$home" answer sample-direct-close --decision-file "$home/close.txt" >/dev/null \ + || fail "a direct close replay failed" + run_captain "$home" answer sample-direct-release --decision-file "$home/release.txt" --release >/dev/null \ + || fail "a direct release replay failed" + printf 'sample-keyed-close\tyes\tYes\n' \ + | run_captain "$home" answers --source "board sequence 2" >/dev/null \ + || fail "a keyed answer replay failed" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" \ + "an idempotent normal-answer replay restored a reconcile request: $list" + pass "normal answers and their replays retire reconcile requests" +} + +# The two verification outcomes, and the honesty of the record each writes. +test_reconcile_closes_with_evidence_or_keeps_the_call_open() { + local home show list rc out + home=$(make_home reconcile-outcomes) + tasks_in "$home" add sample-moot-call "Captain call: ship 0.1.37?" --repo sample >/dev/null + tasks_in "$home" add sample-active-call "Captain call: which admission order?" --repo sample >/dev/null + tasks_in "$home" add sample-mode-call "Captain call: verify replay mode?" --repo sample >/dev/null + run_captain "$home" hold sample-moot-call --reason "ship 0.1.37?" >/dev/null + run_captain "$home" hold sample-active-call --reason "which admission order?" >/dev/null + run_captain "$home" hold sample-mode-call --reason "verify replay mode?" >/dev/null + printf '0.1.38 was published on 2026-09-05, so the 0.1.37 question is moot.\n' > "$home/evidence.txt" + printf 'Still open: nothing has shipped and the choice is unchanged.\n' > "$home/note.txt" + + set +e + out=$(run_captain "$home" reconcile close sample-moot-call --evidence-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "reconcile closed a call without a pending board request" + assert_contains "$out" "no pending board-created reconcile request" \ + "the ungated close refusal did not name the missing board request: $out" + set +e + out=$(run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "reconcile annotated a call without a pending board request" + assert_contains "$out" "no pending board-created reconcile request" \ + "the ungated note refusal did not name the missing board request: $out" + + request_reconciles "$home" board-src sample-moot-call sample-active-call sample-mode-call \ + || fail "could not file the reconcile requests" + + set +e + out=$(run_captain "$home" reconcile close sample-moot-call 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a reconcile close was accepted with no evidence" + assert_contains "$out" "evidence" "the refusal did not name the missing evidence: $out" + + cp "$home/state/reconcile-requests/sample-mode-call.request" "$home/mode-request.backup" + run_captain "$home" answer sample-mode-call --decision-file "$home/evidence.txt" >/dev/null \ + || fail "could not record the normal answer for the mode fixture" + cp "$home/mode-request.backup" "$home/state/reconcile-requests/sample-mode-call.request" + set +e + out=$(run_captain "$home" reconcile close sample-mode-call --evidence-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a normal captain answer replayed as a reconciliation" + assert_contains "$out" "was not closed by reconciliation" \ + "the reconcile replay refusal did not identify the incompatible resolution mode: $out" + run_captain "$home" answer sample-mode-call --decision-file "$home/evidence.txt" >/dev/null \ + || fail "the normal answer replay did not retire its restored pending request" + + run_captain "$home" reconcile close sample-moot-call --evidence-file "$home/evidence.txt" >/dev/null \ + || fail "could not close the moot call with evidence" + show=$(tasks_in "$home" show sample-moot-call --full) + assert_contains "$show" "state: done" "the moot call did not close" + assert_contains "$show" "Resolution mode: reconciled" "the moot call did not record how it closed" + assert_contains "$show" "Reconciliation evidence:" "the moot call did not record the evidence" + assert_contains "$show" "0.1.38 was published" "the recorded evidence was lost" + case "$show" in + *"Captain decision:"*) fail "a reconciled close was recorded as the captain's own words" ;; + esac + set +e + out=$(run_captain "$home" answer sample-moot-call --decision-file "$home/evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a reconciled resolution replayed as a captain answer" + assert_contains "$out" "not a captain-answer replay" \ + "the answer replay refusal did not identify the incompatible resolution mode: $out" + + run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" >/dev/null \ + || fail "could not annotate the still-active call" + show=$(tasks_in "$home" show sample-active-call --full) + assert_contains "$show" "state: queued" "annotating a still-active call closed it" + assert_contains "$show" "held: yes" "annotating a still-active call released it" + assert_contains "$show" "Captain hold reconciled:" "the re-check left no dated note" + assert_contains "$show" "Still open: nothing has shipped" "the note body was lost" + case "$show" in + *"Resolution recorded by"*) fail "annotating a still-active call wrote a resolution record" ;; + esac + set +e + out=$(run_captain "$home" reconcile note sample-active-call --note-file "$home/note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a retired reconcile request appended a duplicate note" + assert_contains "$out" "no pending board-created reconcile request" \ + "the duplicate-note refusal did not name the retired request: $out" + + printf 'sample-active-call\n' \ + | FM_CAPTAIN_HOLD_NOW=2026-09-07T06:00:00Z run_captain "$home" reconcile-requests \ + --source-id board-src --source "captured board result sequence 2" >/dev/null \ + || fail "could not create the second reconcile request" + FM_CAPTAIN_HOLD_NOW=2026-09-07T06:01:00Z run_captain "$home" reconcile note sample-active-call \ + --note-file "$home/note.txt" >/dev/null \ + || fail "the second request with the same finding was not recorded" + show=$(tasks_in "$home" show sample-active-call --full) + [ "$(printf '%s\n' "$show" | grep -o 'Captain hold reconciled:' | wc -l | tr -d ' ')" -eq 2 ] \ + || fail "a later reconcile request with the same note did not append its own record" + assert_contains "$show" "Captain hold reconciled: 2026-09-07T06:01:00Z" \ + "the second reconcile request lost its own dated note" + + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "the verified requests were not retired: $list" + pass "reconcile closes a moot call with evidence and keeps an active one open with a note" +} + +test_reconcile_outcomes_retry_partial_failures_once() { + local home out show list rc + home=$(make_home reconcile-partial-retry) + tasks_in "$home" add sample-reconcile-close-retry "Close retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-note-retry "Note retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-retire-retry "Retire retry" --repo sample >/dev/null + tasks_in "$home" add sample-reconcile-close-retire "Close retire" --repo sample >/dev/null + tasks_in "$home" add sample-answer-retire "Answer retire" --repo sample >/dev/null + run_captain "$home" hold sample-reconcile-close-retry --reason "verify close" >/dev/null + run_captain "$home" hold sample-reconcile-note-retry --reason "verify note" >/dev/null + run_captain "$home" hold sample-reconcile-retire-retry --reason "verify retire" >/dev/null + run_captain "$home" hold sample-reconcile-close-retire --reason "verify close retirement" >/dev/null + run_captain "$home" hold sample-answer-retire --reason "verify answer retirement" >/dev/null + request_reconciles "$home" board-src sample-reconcile-close-retry sample-reconcile-note-retry \ + sample-reconcile-retire-retry sample-reconcile-close-retire sample-answer-retire \ + || fail "could not create partial-retry requests" + printf 'Verified moot.\n' > "$home/retry-evidence.txt" + printf 'Verified active.\n' > "$home/retry-note.txt" + printf 'Verified retirement retry.\n' > "$home/retry-retire-note.txt" + printf 'Verified close retirement.\n' > "$home/close-retire-evidence.txt" + printf 'Captain answered despite retirement failure.\n' > "$home/answer-retire.txt" + cat > "$home/fakebin/tasks-axi" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = done ] && [ "${2:-}" = sample-reconcile-close-retry ] \ + && [ ! -e "$FM_HOME/reconcile-close-failed" ]; then + : > "$FM_HOME/reconcile-close-failed" + exit 92 +fi +if [ "${1:-}" = update ] && [ "${2:-}" = sample-reconcile-note-retry ]; then + "$REAL_TASKS_AXI" "$@" || exit $? + : > "$FM_HOME/reconcile-note-updated" + exit 0 +fi +if [ "${1:-}" = show ] && [ "${2:-}" = sample-reconcile-note-retry ] \ + && [ -e "$FM_HOME/reconcile-note-updated" ] && [ ! -e "$FM_HOME/reconcile-note-show-failed" ]; then + : > "$FM_HOME/reconcile-note-show-failed" + exit 93 +fi +exec "$REAL_TASKS_AXI" "$@" +SH + chmod +x "$home/fakebin/tasks-axi" + + set +e + out=$(run_captain "$home" reconcile close sample-reconcile-close-retry \ + --evidence-file "$home/retry-evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the forced reconcile close failure reported success" + run_captain "$home" reconcile close sample-reconcile-close-retry \ + --evidence-file "$home/retry-evidence.txt" >/dev/null \ + || fail "reconcile close did not recover from its partial failure" + show=$(tasks_in "$home" show sample-reconcile-close-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Resolution mode: reconciled')" -eq 1 ] \ + || fail "reconcile close duplicated its resolution record on retry" + + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-note-retry \ + --note-file "$home/retry-note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "the forced post-note probe failure reported success" + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-note-retry \ + --note-file "$home/retry-note.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a note retry succeeded after the request was retired" + show=$(tasks_in "$home" show sample-reconcile-note-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Captain hold reconciled:')" -eq 1 ] \ + || fail "reconcile note duplicated its durable annotation" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" reconcile close sample-reconcile-close-retire \ + --evidence-file "$home/close-retire-evidence.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a failed close request retirement reported success" + assert_contains "$out" "sample-reconcile-close-retire" \ + "the close retirement failure did not name its task: $out" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "sample-reconcile-close-retire" \ + "the failed close retirement hid its pending request" + chmod 0700 "$home/state/reconcile-requests" + run_captain "$home" reconcile close sample-reconcile-close-retire \ + --evidence-file "$home/close-retire-evidence.txt" >/dev/null \ + || fail "the closed reconciliation could not finish request retirement" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" answer sample-answer-retire \ + --decision-file "$home/answer-retire.txt" 2>&1) + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "a failed answer-boundary retirement reported success" + show=$(tasks_in "$home" show sample-answer-retire --full) + assert_contains "$show" "state: done" "retirement failure reversed the durable captain answer" + assert_contains "$show" "Captain answered despite retirement failure" \ + "retirement failure lost the durable captain answer" + chmod 0700 "$home/state/reconcile-requests" + run_captain "$home" answer sample-answer-retire --decision-file "$home/answer-retire.txt" >/dev/null \ + || fail "the answer replay could not finish request retirement" + + chmod 0500 "$home/state/reconcile-requests" + set +e + out=$(run_captain "$home" reconcile note sample-reconcile-retire-retry \ + --note-file "$home/retry-retire-note.txt" 2>&1) + rc=$? + set -e + chmod 0700 "$home/state/reconcile-requests" + [ "$rc" -ne 0 ] || fail "a failed request retirement reported note success" + assert_not_contains "$out" "still-open:" "failed retirement reported a successful outcome" + run_captain "$home" reconcile note sample-reconcile-retire-retry \ + --note-file "$home/retry-retire-note.txt" >/dev/null \ + || fail "the applied note could not finish request retirement on retry" + show=$(tasks_in "$home" show sample-reconcile-retire-retry --full) + [ "$(printf '%s\n' "$show" | grep -c 'Captain hold reconciled:')" -eq 1 ] \ + || fail "failed request retirement duplicated the reconcile note" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "partial retries left a reconcile request pending" + pass "reconcile outcomes apply durable mutations once across partial failures" +} + test_unbound_source_closes_no_hold() { local home id sid artifact result out show rc home=$(make_home lavish-unbound) @@ -1201,7 +2094,7 @@ test_legacy_identities_keep_working() { # The intake is channel-agnostic, so chat must reach it the same way a captured # review does - for a task-id key, and for a legacy composed identity. test_chat_channel_feeds_the_same_keyed_answer_intake() { - local home id fb show + local home id fb show list home=$(make_home chat-channel) id=sample-chat-review mkdir -p "$home/data/$id" @@ -1216,7 +2109,11 @@ test_chat_channel_feeds_the_same_keyed_answer_intake() { run_captain "$home" hold sample-chat-followup --title "Choose the chat follow-up" \ --reason "captain follow-up choice pending" --repo sample >/dev/null \ || fail "could not register the task-id chat call" - run_captain "$home" complete "$id" "$id-decision-chat-choice" sample-chat-followup >/dev/null \ + run_captain "$home" hold sample-chat-reconcile --title "Reconcile from chat" \ + --reason "captain chat reconcile pending" --repo sample >/dev/null \ + || fail "could not register the chat reconcile call" + run_captain "$home" complete "$id" "$id-decision-chat-choice" sample-chat-followup \ + sample-chat-reconcile >/dev/null \ || fail "completion failed for the chat calls" grep -F 'captain-held [key=chat-choice]' "$home/state/$id.status" >/dev/null \ || fail "precondition: completion did not transfer the decision to its durable owner" @@ -1276,6 +2173,25 @@ SH assert_contains "$show" "Answer: take the second option" "the chat-answered call lost the captain answer" assert_contains "$show" "answer sent to $id" "the chat-answered call lost its channel provenance" + : > "$home/send.log" + set +e + env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$home" FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_SEND_LOG="$home/send.log" FM_SEND_SETTLE=0 \ + "$ROOT/bin/fm-send.sh" "$id" --resolve-key sample-chat-reconcile reconcile >/dev/null 2>&1 + set -e + show=$(tasks_in "$home" show sample-chat-reconcile --full) + assert_contains "$show" "state: queued" "a chat reconcile answer closed the call" + list=$(run_captain "$home" reconcile list) + assert_contains "$list" "reconcile-requests: 0" "a chat reconcile answer created a board request" + printf 'Chat cannot authorize this closure.\n' > "$home/chat-reconcile.txt" + if run_captain "$home" reconcile close sample-chat-reconcile \ + --evidence-file "$home/chat-reconcile.txt" >/dev/null 2>&1; then + fail "a chat reconcile answer authorized evidence-backed closure" + fi + run_captain "$home" answer sample-chat-reconcile --decision-file "$home/chat-reconcile.txt" >/dev/null \ + || fail "could not close the chat reconcile fixture normally" + if env PATH="$fb:$PATH" FM_ROOT_OVERRIDE="$home" FM_HOME="$home" \ FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ FM_SEND_LOG="$home/send.log" FM_SEND_SETTLE=0 \ @@ -1544,6 +2460,7 @@ test_interrupted_cleanup_keeps_the_captain_call_recoverable() { id=sample-held-cleanup-failure wt="$home/projects/$id" mkdir -p "$home/data/$id" "$wt" "$home/projects/sample" + git -C "$home/projects/sample" init -q || fail "could not initialize cleanup-failure project fixture" tasks_in "$home" add "$id" "Investigate failed sample cleanup" --kind scout \ --repo sample --start >/dev/null || fail "could not create the cleanup-failure fixture" fm_write_meta "$home/state/$id.meta" \ @@ -1708,7 +2625,12 @@ test_none_inventory_and_resolved_prose_do_not_create_holds test_terminal_single_owner_status_decision_does_not_block_empty_inventory test_secondmate_hold_stays_in_authoritative_home test_secondmate_home_publishes_holds_and_answers +test_secondmate_reconcile_publishes_before_request_retirement test_bound_channel_answers_close_at_answer_time +test_reconcile_never_closes_through_the_keyed_answer_intake +test_normal_answers_retire_pending_reconcile_requests +test_reconcile_closes_with_evidence_or_keeps_the_call_open +test_reconcile_outcomes_retry_partial_failures_once test_unbound_source_closes_no_hold test_legacy_identities_keep_working test_chat_channel_feeds_the_same_keyed_answer_intake @@ -1719,3 +2641,10 @@ test_teardown_never_closes_a_captain_held_task test_interrupted_cleanup_keeps_the_captain_call_recoverable test_teardown_retains_captain_calls_in_a_relocated_backlog test_teardown_refuses_a_ship_when_the_captain_hold_cannot_be_read +test_verify_resolves_a_hold_migrated_to_beads_notes +test_verify_resolves_a_hold_migrated_under_the_configured_prefix +test_marker_noted_row_wins_over_a_prefix_namesake +test_complete_accepts_a_migrated_inventory_on_beads +test_verify_names_the_unresolvable_legacy_id_once +test_verify_resolves_a_pre_collapse_key_through_its_derived_marker +test_captain_hold_mutations_address_the_beads_backend diff --git a/tests/fm-claude-stop-autoarm-live-e2e.test.sh b/tests/fm-claude-stop-autoarm-live-e2e.test.sh index 60012667981..ae14d9f3af5 100755 --- a/tests/fm-claude-stop-autoarm-live-e2e.test.sh +++ b/tests/fm-claude-stop-autoarm-live-e2e.test.sh @@ -12,10 +12,10 @@ # shellcheck disable=SC2016 # the model, not this test shell, reads the prompt text set -u -if [ "${FM_CLAUDE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CLAUDE_LIVE_E2E=1 to run the Claude Stop auto-arm regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_CLAUDE_LIVE_E2E claude ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -24,8 +24,6 @@ fail() { exit 1 } -command -v claude >/dev/null 2>&1 || fail "claude not found" - LAB="$ROOT/.claude-autoarm-live-e2e.$$" PROJECT="$LAB/project" HOME_DIR="$LAB/fmhome" diff --git a/tests/fm-claude-stop-autoarm.test.sh b/tests/fm-claude-stop-autoarm.test.sh index 042d04ba947..bbae5f16574 100755 --- a/tests/fm-claude-stop-autoarm.test.sh +++ b/tests/fm-claude-stop-autoarm.test.sh @@ -174,6 +174,15 @@ echo "$$" >> "$FM_HOME/state/arm-ran" printf 'watcher: started pid=%s (beacon fresh)\n' "$$" printf 'stale: fixture-win actionable\n' exit 0 +SH + ;; + records-grace) + cat > "$dir/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +echo "$$" >> "$FM_HOME/state/arm-ran" +printf '%s\n' "${FM_GUARD_GRACE:-unset}" > "$FM_HOME/state/arm-received-grace" +printf 'watcher: attached pid=%s (beacon 2s)\n' "$$" +exit 0 SH ;; *) @@ -623,6 +632,20 @@ test_arms_for_x_mode_poll_need_without_inflight() { pass "auto-arm: X-mode poll need arms the cycle even with no tasks in flight" } +test_arms_for_registered_custom_check_without_inflight() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/check-need") + printf '#!/usr/bin/env bash\nexit 0\n' > "$dir/state/issue-comments.check.sh" + chmod 700 "$dir/state/issue-comments.check.sh" + FM_STATE_OVERRIDE="$dir/state" "$ROOT/bin/fm-check-register.sh" issue-comments >/dev/null \ + || fail "fm-check-register.sh could not register the custom check" + write_arm_fixture "$dir" actionable + out=$(run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "a registered custom check must keep the auto-arm active with zero tasks in flight" + [ -e "$dir/state/arm-ran" ] || fail "hook did not arm for the registered custom check" + pass "auto-arm: a registered custom check arms the cycle even with no tasks in flight" +} + test_single_flight_admits_exactly_one_owner() { local dir rc1 rc2 count dir=$(make_primary_dir "$TMP_ROOT/single-flight") @@ -1142,6 +1165,18 @@ test_active_in_marked_secondmate_home() { pass "auto-arm: active in a marked secondmate home" } +test_long_poll_grace_reaches_arm_wrapper() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/long-poll-grace") + : > "$dir/state/task.meta" + write_arm_fixture "$dir" records-grace + out=$(unset FM_GUARD_GRACE; FM_POLL=900 run_autoarm "$dir" 2>/dev/null); status=$? + expect_code 2 "$status" "an unverified close without a healthy watcher must still fail closed" + [ -e "$dir/state/arm-received-grace" ] || fail "arm wrapper never recorded FM_GUARD_GRACE" + [ "$(cat "$dir/state/arm-received-grace")" = 960 ] || fail "arm wrapper must see the poll-derived grace (900+60), got: $(cat "$dir/state/arm-received-grace")" + pass "auto-arm: a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived grace" +} + test_fm_lock_status_still_works_with_shared_lib() { local out out=$(FM_HOME="$TMP_ROOT/lock-status-home" bash "$ROOT/bin/fm-lock.sh" status 2>&1) @@ -1168,6 +1203,7 @@ test_benign_cycle_end_with_live_watcher_is_silent test_positive_recovery_budget_contention_preserves_episode test_owner_mutex_contention_preserves_failure_episode_reset test_arms_for_x_mode_poll_need_without_inflight +test_arms_for_registered_custom_check_without_inflight test_single_flight_admits_exactly_one_owner test_abandoned_owner_claim_is_reclaimed_and_rearms test_arming_claim_with_fresh_beacon_is_never_reclaimed @@ -1187,4 +1223,5 @@ test_superseded_owner_goes_silent_and_never_double_translates test_need_vanished_mid_cycle_closes_quietly test_afk_mid_cycle_suppresses_rewake test_active_in_marked_secondmate_home +test_long_poll_grace_reaches_arm_wrapper test_fm_lock_status_still_works_with_shared_lib diff --git a/tests/fm-cmux-claude-composer-live-e2e.test.sh b/tests/fm-cmux-claude-composer-live-e2e.test.sh index f5c44d26c92..439d9335e97 100755 --- a/tests/fm-cmux-claude-composer-live-e2e.test.sh +++ b/tests/fm-cmux-claude-composer-live-e2e.test.sh @@ -4,6 +4,9 @@ # only one exact fm-test- workspace through the normal scout lifecycle. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" TASK="fm-test-cmux-claude-composer-$$" LAB= @@ -26,16 +29,8 @@ cleanup() { [ -z "$LAB" ] || rm -rf -- "$LAB" } -if [ "${FM_CMUX_CLAUDE_COMPOSER_LIVE:-0}" != 1 ]; then - echo "skip: set FM_CMUX_CLAUDE_COMPOSER_LIVE=1 to run the real cmux Claude composer drift guard" - exit 0 -fi +fm_live_gate opt-in FM_CMUX_CLAUDE_COMPOSER_LIVE claude cmux jq treehouse python3 -command -v claude >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but Claude Code is not installed" -command -v cmux >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but cmux is not installed" -command -v jq >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but jq is not installed" -command -v treehouse >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but treehouse is not installed" -command -v python3 >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but python3 is not installed" cmux ping >/dev/null 2>&1 || fail "FM_CMUX_CLAUDE_COMPOSER_LIVE=1 but the cmux socket is unavailable" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-cmux-claude-composer.XXXXXX") || fail "could not create an isolated cmux Claude lab" diff --git a/tests/fm-codex-continuity-live-e2e.test.sh b/tests/fm-codex-continuity-live-e2e.test.sh index 6d1a6b223ca..ccc7133c832 100755 --- a/tests/fm-codex-continuity-live-e2e.test.sh +++ b/tests/fm-codex-continuity-live-e2e.test.sh @@ -3,10 +3,10 @@ # Codex's bounded foreground-checkpoint supervision path. set -u -if [ "${FM_CODEX_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CODEX_LIVE_E2E=1 to run the Codex continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_CODEX_LIVE_E2E codex ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -15,8 +15,6 @@ fail() { exit 1 } -command -v codex >/dev/null 2>&1 || fail "codex not found" - LAB="$ROOT/.codex-live-e2e.$$" PROJECT="$LAB/project" HOME_DIR="$LAB/fmhome" diff --git a/tests/fm-composer-lib.test.sh b/tests/fm-composer-lib.test.sh index eef745fc820..5ee3e225df9 100755 --- a/tests/fm-composer-lib.test.sh +++ b/tests/fm-composer-lib.test.sh @@ -282,12 +282,75 @@ test_matrix_herdr_halfblock_rule_bounds_bare_wrap() { case "$plain" in *"Run Everything"*) : ;; *) fail "fixture lost its footer content" ;; esac ESC_LOCAL=$(printf '\033') screen=$'transcript\n ▄▄▄▄▄▄▄▄\n'" ${ESC_LOCAL}[2m→ ${ESC_LOCAL}[0;7mA${ESC_LOCAL}[0;2mdd a follow-up${ESC_LOCAL}[0m"$'\n ▀▀▀▀▀▀▀▀\n Cursor Grok 4.5 High · 6.7% Run Everything\n ~/wt · 64cdd3a' - out=$(fm_composer_classify_screen "$CAPS_STYLED" "$(printf '%b' "$screen")") + out=$(fm_composer_classify_screen "$CAPS_STYLED" "$screen") [ "$out" = empty ] \ || fail "an idle cursor composer inside herdr half-block rules must read empty, got '$out'" pass "matrix: herdr half-block rules bound a bare composer's wrap region" } +test_matrix_omp_status_row_bounds_bare_composer() { + # omp (Oh My Pi) draws its status line directly BELOW the borderless `❯` + # composer. Captured live through Herdr on omp 18.1.11 under the captain's + # unicode preset (idle), plus the nerd-preset idle row and the busy spinner + # row from the 18.1.2 investigation. Without the status-row rule the bare + # wrap region swallows that row and an idle omp pane reads `pending`, which + # skipped the doorbell on the first live omp worker. + local idle_unicode idle_nerd busy typed wrapped + idle_unicode=$'transcript line + +❯ + π · ◔ GPT-6-Astra · 🌳 …-workspace · ⑂ detached · ◫ 15.4%/272K ⟲ · (sub)' + idle_nerd=$'transcript line + +❯ + 󰵗 · qwen3:8b · kun-agent-workspace/… · detached ?1 · 36.7%/41K' + busy=$'transcript line + + ⎋ Working… + +❯ + ⠧ 11s · ◔ GPT-6-Astra · ◫ 15.4%/272K' + typed=$'transcript line + +❯ fix the flaky test + π · ◔ GPT-6-Astra · 🌳 …-workspace · ⑂ detached · ◫ 15.4%/272K ⟲ · (sub)' + # Non-vacuousness: each status row is real non-blank content that the wrap + # region would otherwise take as typed input. + _fm_composer_row_is_omp_status ' π · ◔ GPT-6-Astra · 🌳 …-workspace' \ + || fail "the unicode-preset omp status row must be recognized as furniture" + _fm_composer_row_is_omp_status ' 󰵗 · qwen3:8b · kun-agent-workspace/… · detached ?1 · 36.7%/41K' \ + || fail "the nerd-preset omp status row must be recognized as furniture" + _fm_composer_row_is_omp_status ' ⠧ 11s · ◔ GPT-6-Astra' \ + || fail "the busy omp spinner row must be recognized as furniture" + _fm_composer_row_is_omp_status 'fix the flaky test' \ + && fail "ordinary typed text must not be mistaken for omp status furniture" + _fm_composer_row_is_omp_status 'please rerun the suite and report' \ + && fail "ordinary prose must not be mistaken for omp status furniture" + # Only omp's identity cell opens the row: a wrapped typed row that happens + # to begin with a short word and a spaced middle dot is composer input. + _fm_composer_row_is_omp_status 'fix · tests before pushing' \ + && fail "wrapped typed text with a middle dot must not be mistaken for omp status furniture" + # The ascii preset's identity cell is `pi`, but that preset separates its + # cells with ` - `, so a row opening `pi ·` is never omp furniture. + _fm_composer_row_is_omp_status 'pi · e · phi as the three constants' \ + && fail "typed text opening 'pi ·' must not be mistaken for omp status furniture" + _fm_composer_row_is_omp_status ' ⣾ 3s · ◔ GPT-6-Astra' \ + || fail "the status-set omp spinner row must be recognized as furniture" + assert_screen "idle omp (unicode preset)" empty "$CAPS_STYLED" "$idle_unicode" + assert_screen "idle omp (nerd preset)" empty "$CAPS_STYLED" "$idle_nerd" + assert_screen "busy omp keeps an empty composer" empty "$CAPS_STYLED" "$busy" + assert_screen "typed omp text is pending" pending "$CAPS_STYLED" "$typed" + assert_screen "idle omp on a plain capture" empty "$CAPS_PLAIN" "$idle_unicode" + # The boundary must not cut a bare composer's own wrapped input: with the + # cursor on a continuation row that opens `fix · tests`, the composer is a + # proven wrap region and reads pending, exactly as it did before the rule. + wrapped=$'transcript line\n\n❯ please run the suite and then\nfix · tests before pushing' + assert_screen "wrapped typed text with a middle dot stays pending" pending "$CAPS_TMUX" "$wrapped" 3 + wrapped=$'transcript line\n\n❯ document the constants in the order\npi · e · phi with one example each' + assert_screen "wrapped typed text opening 'pi ·' stays pending" pending "$CAPS_TMUX" "$wrapped" 3 + pass "matrix: omp's status row bounds the bare composer's wrap region" +} + test_matrix_pi_separated_needs_identity() { # Real idle pi: a blank row between two solid rules. The blank row alone is # exactly what the strict rule refuses; only structure PLUS a live @@ -617,6 +680,7 @@ test_matrix_codex_dim_hint_row test_matrix_muse_truecolor_glyph_survives_signal_loss test_matrix_cursor_reverse_video_placeholder_remnant test_matrix_herdr_halfblock_rule_bounds_bare_wrap +test_matrix_omp_status_row_bounds_bare_composer test_matrix_pi_separated_needs_identity test_matrix_opencode_leftbar_signals test_matrix_grok_titled_bottom_border diff --git a/tests/fm-composer-matrix-live-e2e.test.sh b/tests/fm-composer-matrix-live-e2e.test.sh index 9bb78ade445..feb94b3be3b 100755 --- a/tests/fm-composer-matrix-live-e2e.test.sh +++ b/tests/fm-composer-matrix-live-e2e.test.sh @@ -27,14 +27,12 @@ # unreadable-composer state and correctly fails that harness's check. set -u -ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -if [ "${FM_COMPOSER_MATRIX_LIVE:-0}" != 1 ]; then - echo "skip: set FM_COMPOSER_MATRIX_LIVE=1 to run the live composer-matrix guard" - exit 0 -fi +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -command -v tmux >/dev/null 2>&1 || { echo "not ok - FM_COMPOSER_MATRIX_LIVE=1 but tmux is not installed" >&2; exit 1; } +fm_live_gate opt-in FM_COMPOSER_MATRIX_LIVE tmux SOCKET="fm-cmx-live-$$" SESSION="cmxlive" diff --git a/tests/fm-control-relaunch.test.sh b/tests/fm-control-relaunch.test.sh index fe2f81bc5c9..b5c58ba1b1d 100755 --- a/tests/fm-control-relaunch.test.sh +++ b/tests/fm-control-relaunch.test.sh @@ -403,7 +403,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { FM_FAKE_TRACE_RELEASE="$launch_release" \ run_control "$dir" rl28 relaunch --note "continue after publication" > "$dir/control.out" & control_pid=$! - while [ ! -e "$prepare" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$prepare" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done @@ -422,7 +422,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { --carry-platform x --carry-max 280 > "$dir/link.out" 2>&1 & link_pid=$! i=0 - while [ ! -e "$waiting" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$waiting" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done @@ -435,7 +435,7 @@ test_relaunch_serializes_concurrent_durable_metadata_publication() { } : > "$launch_release" i=0 - while [ ! -e "$ready" ] && [ "$i" -lt 200 ]; do + while [ ! -e "$ready" ] && [ "$i" -lt 500 ]; do /bin/sleep 0.01 i=$((i + 1)) done diff --git a/tests/fm-control.test.sh b/tests/fm-control.test.sh index 79254f7157e..033cbfd2827 100755 --- a/tests/fm-control.test.sh +++ b/tests/fm-control.test.sh @@ -35,7 +35,7 @@ mkdir -p "$TMP_ROOT" TMP_ROOT=$(cd "$TMP_ROOT" && pwd) trap 'rm -rf "$TMP_ROOT"' EXIT -VERIFIED_HARNESSES="claude codex copilot opencode pi pi-signed grok kimi cursor muse" +VERIFIED_HARNESSES="claude codex copilot opencode pi pi-signed grok kimi cursor muse omp" # The expectation table, written out independently of the implementation so a # silent change to either side shows up here. The fourth field is the composer @@ -49,6 +49,7 @@ verified_adapter_contract() { # <harness> -> exit command, interrupt key, repea opencode) printf '/exit\tEscape\t2\t\n' ;; pi) printf '/quit\tEscape\t1\t\n' ;; pi-signed) printf '/quit\tEscape\t1\t\n' ;; + omp) printf '/quit\tEscape\t1\t\n' ;; grok) printf '/exit\tC-c\t1\t\n' ;; kimi) printf '/exit\tEscape\t1\t\n' ;; cursor) printf '/exit\tEscape\t1\t\n' ;; @@ -269,7 +270,7 @@ test_harness_family_resolution() { copilot:copilot copilot-cli:copilot \ opencode:opencode grok:grok grok-2:grok kimi:kimi cursor:cursor \ cursor-agent:cursor muse:muse muse-bin-0.1.0:muse pi:pi \ - pi-signed:pi-signed; do + pi-signed:pi-signed omp:omp; do recorded=${pair%%:*} want=${pair#*:} got=$(fm_control_harness_family "$recorded") \ @@ -283,6 +284,11 @@ test_harness_family_resolution() { # The signed adapter is a distinct launch profile, not a pi variant. [ "$(fm_control_harness_family pi-signed)" != "$(fm_control_harness_family pi)" ] \ || fail "pi-signed must not collapse into pi" + # omp is exact: an omp* prefix would claim unrelated commands such as ompd. + fm_control_harness_family ompd \ + && fail "ompd must not be guessed into the omp adapter" + fm_control_harness_family comp \ + && fail "comp must not be guessed into the omp adapter" pass "fm-control-lib: a recorded harness resolves to its verified adapter without guessing" } @@ -377,7 +383,7 @@ test_harness_kind_capability() { done fm_control_harness_supports_kind muse secondmate \ && fail "muse has no primary supervision protocol and must not claim a secondmate" - for harness in claude codex copilot opencode pi pi-signed grok kimi; do + for harness in claude codex copilot opencode pi pi-signed grok kimi omp; do fm_control_harness_supports_kind "$harness" secondmate \ || fail "$harness should be able to run a secondmate" done diff --git a/tests/fm-copilot-hooks-live-e2e.test.sh b/tests/fm-copilot-hooks-live-e2e.test.sh index 534ade53e81..648749dc442 100755 --- a/tests/fm-copilot-hooks-live-e2e.test.sh +++ b/tests/fm-copilot-hooks-live-e2e.test.sh @@ -7,14 +7,11 @@ # against the installed CLI. set -u -if [ "${FM_COPILOT_HOOKS_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_COPILOT_HOOKS_LIVE_E2E=1 to run the live Copilot hook discovery guard" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_COPILOT_HOOKS_LIVE_E2E + COPILOT_BIN=${FM_COPILOT_BIN:-$(command -v copilot || true)} [ -n "$COPILOT_BIN" ] && [ -x "$COPILOT_BIN" ] \ || fail "copilot not found; install it or set FM_COPILOT_BIN. This guard refuses to pass without checking the real harness." diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index f5f5f0e2f21..309a7008f8a 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -11,8 +11,15 @@ # source): # (a) active run-step is authoritative -> run-step # (b) needs-decision/blocked log + resumed run = SUPERSEDED -> run-step +# (b2) blocked log claiming the daemon/timeout while the run is fixing with +# fresh activity = superseded BECAUSE THE RUN IS ALIVE; the same claim +# a genuine socket-refusal claim over a stale or terminal run record +# remains blocked, and an ordinary blocked log over a live run keeps the generic +# superseded reading # (c) genuine parked run + needs-decision log = NOT superseded -> run-step # (d) terminal run-step (passed/failed) is authoritative -> run-step +# (d2) terminal failed run whose only failure is an orphaned ci monitor +# after checks read green -> done # (e) cross-branch attribution: this branch's own run found via list lookup # (f) no run + semantic busy -> pane # (g) no run + semantic idle falls to the status-log verb -> status-log @@ -25,6 +32,10 @@ # This is the direct regression pair for the 2026-07-02 herdr incident, # proving the watcher's own absorb-only-when-provably-working predicate # benefits from the fix in both directions. +# (l) coarse runs-ledger fallback: a terminal failed record with the daemon +# provably down (explicit daemon-status probe fails) reads unknown - +# "unverified", never failed; the same record with the daemon up stays +# failed. set -u # shellcheck source=tests/lib.sh @@ -76,6 +87,12 @@ case "${1:-}" in ;; runs) printf '%s\n' "${FM_FAKE_RUNS_LIST:-}" ;; + daemon) + # FM_FAKE_DAEMON_DOWN: the explicit down-probe fails, as the real + # `no-mistakes daemon status` does when the daemon is not running. + [ "${FM_FAKE_DAEMON_DOWN:-0}" = 1 ] && exit 1 + printf '%s\n' 'daemon running (pid 4242)' + exit 0 ;; esac exit 0 SH @@ -198,8 +215,10 @@ reset_fakes() { FM_FAKE_HERDR_HUSK=0 FM_FAKE_HERDR_AGENT_STATUS="" FM_FAKE_CI_LOGS="" + FM_FAKE_DAEMON_DOWN=0 export FM_FAKE_AXI_STATUS FM_FAKE_AXI_STATUS_RUN FM_FAKE_RUNS_LIST FM_FAKE_BUSY FM_FAKE_BUSY_TEXT FM_FAKE_TMUX_MISSING FM_FAKE_TMUX_UNREADABLE export FM_FAKE_HERDR_BUSY FM_FAKE_HERDR_MISSING FM_FAKE_HERDR_READ_FAIL FM_FAKE_HERDR_HUSK FM_FAKE_HERDR_AGENT_STATUS FM_FAKE_CI_LOGS + export FM_FAKE_DAEMON_DOWN } # --- run-object fixtures (TOON, as `no-mistakes axi status` emits) ----------- @@ -231,6 +250,41 @@ run: EOF } +# A fixing run whose active step reports FRESH activity. `axi status` emits the +# active_steps table only while a step is running or fixing, and leaves +# last_activity unprefixed while step-log or agent lifecycle events keep +# arriving - that is the client's own recency verdict. +run_fixing_active_recent() { # <branch> + cat <<EOF +run: + id: "01RUN" + branch: $1 + status: fixing + head: "${FM_FAKE_RUN_HEAD:-abc1234}" + pr: "" + findings: none + active_steps[1]{step,active_for,last_activity,agent_pid,round}: + review,12m3s,8s,44121,"auto-fix 1/3" +EOF +} + +# The same run gone QUIET: the client prefixes last_activity with `quiet` once +# nothing has arrived for longer than its configured quiet warning. This is the +# shape a run record keeps when the daemon really did die under it. +run_fixing_active_quiet() { # <branch> + cat <<EOF +run: + id: "01RUN" + branch: $1 + status: fixing + head: "${FM_FAKE_RUN_HEAD:-abc1234}" + pr: "" + findings: none + active_steps[1]{step,active_for,last_activity,agent_pid,round}: + review,42m8s,"quiet 31m2s",44121,"auto-fix 1/3" +EOF +} + run_top_level_ci() { # <branch> cat <<EOF run: @@ -319,6 +373,78 @@ outcome: failed EOF } +# The 2026-09-05 jr-voice orphaned-CI-monitor shape: every substantive step +# completed, only ci failed (after the shared daemon restarted under its +# merge poll), and GitHub read the PR green and mergeable. +run_failed_ci_orphan() { # <branch> + cat <<EOF +run: + id: "01RUN" + branch: $1 + status: failed + head: "${FM_FAKE_RUN_HEAD:-abc1234}" + pr: "https://github.com/o/r/pull/203" + findings: none +outcome: failed +steps[9]{step,status,findings,duration_ms}: + intent,completed,0,0 + rebase,completed,0,0 + review,completed,0,0 + test,completed,0,0 + document,completed,0,0 + lint,completed,0,0 + push,completed,0,0 + pr,completed,0,0 + ci,failed,0,76127890 +EOF +} + +# Same shape but with no outcome line: only top-level status reads failed. +run_failed_ci_orphan_status_only() { # <branch> + cat <<EOF +run: + id: "01RUN" + branch: $1 + status: failed + head: "${FM_FAKE_RUN_HEAD:-abc1234}" + pr: "https://github.com/o/r/pull/203" + findings: none +steps[9]{step,status,findings,duration_ms}: + intent,completed,0,0 + rebase,completed,0,0 + review,completed,0,0 + test,completed,0,0 + document,completed,0,0 + lint,completed,0,0 + push,completed,0,0 + pr,completed,0,0 + ci,failed,0,76127890 +EOF +} + +# A second failed step (lint) disqualifies the orphaned-monitor reclassification. +run_failed_ci_orphan_second_failure() { # <branch> + cat <<EOF +run: + id: "01RUN" + branch: $1 + status: failed + head: "${FM_FAKE_RUN_HEAD:-abc1234}" + pr: "https://github.com/o/r/pull/203" + findings: none +steps[9]{step,status,findings,duration_ms}: + intent,completed,0,0 + rebase,completed,0,0 + review,completed,0,0 + test,completed,0,0 + document,completed,0,0 + lint,failed,0,0 + push,completed,0,0 + pr,completed,0,0 + ci,failed,0,76127890 +EOF +} + run_ci_monitoring() { # <branch> cat <<EOF run: @@ -417,6 +543,118 @@ test_stale_blocked_superseded() { pass "stale blocked over active run is superseded" } +# A crew whose drive call timed out or was killed by its harness command limit +# routinely blocks claiming the pipeline died. The daemon accepts `respond` +# immediately and runs the fix round in the background, so such a claim over a +# run that is fixing WITH fresh activity is contradicted by the run itself: the +# supervisor answer is to steer a reattach, not to escalate a dead pipeline. +test_daemon_claim_over_live_run_reads_run_alive() { + reset_fakes + local d; d=$(new_case daemon-claim-live) + make_repo_on_branch "$d/wt" fm/feat-dl + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-dl.meta" "window=fm:fm-feat-dl" "worktree=$d/wt" "kind=ship" + printf 'blocked: no-mistakes daemon unreachable, drive run: read response: i/o timeout\n' \ + > "$d/state/feat-dl.status" + FM_FAKE_AXI_STATUS="$(run_fixing_active_recent fm/feat-dl)" + local out; out=$(run_crew_state "$d" feat-dl) + assert_contains "$out" "state: working" "live run beats the crew's death claim" + assert_contains "$out" "source: run-step" "live run -> run-step source" + assert_contains "$out" "run alive" "daemon claim over a live run is named as run alive" + assert_contains "$out" "reattach" "the reading names the reattach steer" + assert_not_contains "$out" "superseded by active run" \ + "the daemon claim gets the sharper reading, not the generic one" + pass "daemon/timeout blocked claim over a live fixing run reads as run alive" +} + +# A genuine refused socket outranks the persisted fixing record, which can +# survive after the daemon exits. +test_socket_refusal_over_stale_fixing_run_reports_blocked() { + reset_fakes + local d; d=$(new_case daemon-socket-refused) + make_repo_on_branch "$d/wt" fm/feat-dq + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-dq.meta" "window=fm:fm-feat-dq" "worktree=$d/wt" "kind=ship" + printf 'blocked: no-mistakes daemon socket refused connections\n' \ + > "$d/state/feat-dq.status" + FM_FAKE_AXI_STATUS="$(run_fixing_active_quiet fm/feat-dq)" + local out; out=$(run_crew_state "$d" feat-dq) + assert_contains "$out" "state: blocked" "socket refusal outranks a stale fixing record" + assert_contains "$out" "source: status-log" "socket refusal remains status-log evidence" + assert_contains "$out" "socket refused connections" "socket failure is preserved" + assert_not_contains "$out" "run alive" "stale fixing record is not reported alive" + + # Exercise the exact alternate wordings emitted by the generated crew rule. + printf 'blocked: no-mistakes daemon socket refuses connections\n' \ + > "$d/state/feat-dq.status" + out=$(run_crew_state "$d" feat-dq) + assert_contains "$out" "state: blocked" "socket-refuses wording outranks a stale fixing record" + assert_not_contains "$out" "state: working" "socket-refuses wording cannot be suppressed by a stale active record" + + printf 'blocked: no-mistakes daemon socket is missing\n' \ + > "$d/state/feat-dq.status" + out=$(run_crew_state "$d" feat-dq) + assert_contains "$out" "state: blocked" "missing socket outranks a stale fixing record" + assert_contains "$out" "source: status-log" "missing socket remains status-log evidence" + assert_not_contains "$out" "state: working" "missing socket cannot be suppressed by a stale active record" + pass "socket refusal or missing socket over a stale fixing run reports blocked" +} + +# A terminal run record can be the final persisted state after the daemon exits. +# Positive socket-failure evidence must not be discarded merely because that +# attributed record no longer has an active status. +test_socket_refusal_over_terminal_run_reports_blocked() { + reset_fakes + local d; d=$(new_case daemon-socket-refused-terminal) + make_repo_on_branch "$d/wt" fm/feat-dqt + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-dqt.meta" "window=fm:fm-feat-dqt" "worktree=$d/wt" "kind=ship" + printf 'blocked: no-mistakes daemon socket refused connections\n' \ + > "$d/state/feat-dqt.status" + FM_FAKE_AXI_STATUS="$(run_failed fm/feat-dqt)" + local out; out=$(run_crew_state "$d" feat-dqt) + assert_contains "$out" "state: blocked" "socket refusal outranks a terminal run record" + assert_contains "$out" "source: status-log" "terminal run cannot suppress socket-failure evidence" + assert_not_contains "$out" "state: failed" "terminal run state is not emitted over socket-failure evidence" + pass "socket refusal over a terminal attributed run reports blocked" +} + +# And the claim half: an ordinary blocked line over the same live run keeps the +# generic reading, so the sharper one cannot fire on every superseded block. +test_ordinary_blocked_over_live_run_keeps_plain_superseded() { + reset_fakes + local d; d=$(new_case ordinary-blocked-live) + make_repo_on_branch "$d/wt" fm/feat-ob + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ob.meta" "window=fm:fm-feat-ob" "worktree=$d/wt" "kind=ship" + printf 'blocked: database upload failed with broken pipe\n' > "$d/state/feat-ob.status" + FM_FAKE_AXI_STATUS="$(run_fixing_active_recent fm/feat-ob)" + local out; out=$(run_crew_state "$d" feat-ob) + assert_contains "$out" "state: working" "ordinary blocked log over an active run -> working" + assert_contains "$out" "superseded by active run" "ordinary blocked keeps the generic reading" + assert_not_contains "$out" "run alive" "broken pipe is not a pipeline-unreachable alias" + pass "broken-pipe blocker over a live run keeps the plain superseded reading" +} + +# The genuine daemon-down case still reaches the supervisor as blocked: the +# socket refused connections and no run is executing anywhere. +test_genuine_daemon_down_reports_blocked() { + reset_fakes + local d; d=$(new_case daemon-down) + make_repo_on_branch "$d/wt" fm/feat-dd + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-dd.meta" "window=fm:fm-feat-dd" "worktree=$d/wt" "kind=ship" "harness=claude" + printf 'blocked: no-mistakes daemon socket refused connections\n' > "$d/state/feat-dd.status" + FM_FAKE_AXI_STATUS="" + FM_FAKE_BUSY=0 + arm_idle_record "$d/state" feat-dd + local out; out=$(run_crew_state "$d" feat-dd) + assert_contains "$out" "state: blocked" "a genuine daemon-down claim with no run stays blocked" + assert_contains "$out" "source: status-log" "no run -> status-log source" + assert_not_contains "$out" "run alive" "nothing is alive to report" + pass "genuine daemon-down blocked line still reports blocked" +} + # (c) genuine parked run + needs-decision log AGREE -> parked, NOT superseded test_genuine_parked_not_superseded() { reset_fakes @@ -712,6 +950,69 @@ test_terminal_failed() { pass "terminal failed run is authoritative" } +test_terminal_failed_ci_orphan_after_green_reads_done() { + reset_fakes + local d; d=$(new_case failed-ci-orphan) + make_repo_on_branch "$d/wt" fm/feat-ci-orphan + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ci-orphan.meta" "window=fm:fm-feat-ci-orphan" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed_ci_orphan fm/feat-ci-orphan)" + FM_FAKE_CI_LOGS="all CI checks passed - still monitoring until merged or closed +daemon shutting down" + local out; out=$(run_crew_state "$d" feat-ci-orphan) + assert_contains "$out" "state: done" "orphaned ci monitor after green must read done, not failed" + assert_contains "$out" "source: run-step" "reclassified held run stays run-step sourced" + assert_contains "$out" "https://github.com/o/r/pull/203" "PR URL surfaced from the run" + assert_not_contains "$out" "state: failed" "monitor death must not read as a failed run" + pass "orphaned ci monitor after green reads as held-for-merge done" +} + +test_terminal_failed_ci_orphan_status_only_reads_done() { + reset_fakes + local d; d=$(new_case failed-ci-orphan-status-only) + make_repo_on_branch "$d/wt" fm/feat-ci-orphan2 + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ci-orphan2.meta" "window=fm:fm-feat-ci-orphan2" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed_ci_orphan_status_only fm/feat-ci-orphan2)" + FM_FAKE_CI_LOGS="all CI checks passed - still monitoring until merged or closed +daemon shutting down" + local out; out=$(run_crew_state "$d" feat-ci-orphan2) + assert_contains "$out" "state: done" "status-only failed orphaned monitor after green reads done" + assert_contains "$out" "https://github.com/o/r/pull/203" "PR URL surfaced from the run" + pass "status-only failed orphaned ci monitor after green reads done" +} + +test_terminal_failed_ci_genuine_red_stays_failed() { + reset_fakes + local d; d=$(new_case failed-ci-genuine-red) + make_repo_on_branch "$d/wt" fm/feat-ci-red + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ci-red.meta" "window=fm:fm-feat-ci-red" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed_ci_orphan fm/feat-ci-red)" + FM_FAKE_CI_LOGS="CI checks running +checks failed: 1 of 2 checks red +daemon shutting down" + local out; out=$(run_crew_state "$d" feat-ci-red) + assert_contains "$out" "state: failed" "a genuinely red check keeps the run failed" + assert_not_contains "$out" "state: done" "genuine CI failure must not reclassify to done" + pass "genuinely failing CI keeps the failed verdict" +} + +test_terminal_failed_ci_orphan_second_failed_step_stays_failed() { + reset_fakes + local d; d=$(new_case failed-ci-second-failure) + make_repo_on_branch "$d/wt" fm/feat-ci-2fail + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-ci-2fail.meta" "window=fm:fm-feat-ci-2fail" "worktree=$d/wt" "kind=ship" + FM_FAKE_AXI_STATUS="$(run_failed_ci_orphan_second_failure fm/feat-ci-2fail)" + FM_FAKE_CI_LOGS="all CI checks passed - still monitoring until merged or closed +daemon shutting down" + local out; out=$(run_crew_state "$d" feat-ci-2fail) + assert_contains "$out" "state: failed" "a second failed step keeps the run failed" + assert_not_contains "$out" "state: done" "a second failed step must not reclassify to done" + pass "a second failed step disqualifies the orphaned-monitor reclassification" +} + # (e) cross-branch attribution: `axi status` returns ANOTHER branch's run (the # routine case once more than one crew validates the same underlying repo # concurrently - they share ONE no-mistakes repo registration), so the helper @@ -746,6 +1047,66 @@ EOF # The runs list is newest-first; a branch with an OLDER completed run must not # shadow its own newer active one - the first (topmost) matching row wins. +test_coarse_socket_refusal_reports_blocked() { + reset_fakes + local d short; d=$(new_case coarse-socket-refused) + make_repo_on_branch "$d/wt" fm/feat-coarse-down + short=$(git -C "$d/wt" rev-parse --short=7 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-coarse-down.meta" "window=fm:fm-feat-coarse-down" "worktree=$d/wt" "kind=ship" + printf 'blocked: no-mistakes daemon connection refused\n' > "$d/state/feat-coarse-down.status" + FM_FAKE_AXI_STATUS="$(run_running fm/other-crew)" + FM_FAKE_RUNS_LIST="$(cat <<EOF + running fm/other-crew aaaaaaa 2026-07-02 22:10 + running fm/feat-coarse-down ${short} 2026-07-02 22:05 +EOF +)" + local out; out=$(run_crew_state "$d" feat-coarse-down) + assert_contains "$out" "state: blocked" "socket refusal outranks a coarse active record" + assert_contains "$out" "source: status-log" "coarse socket refusal remains status-log evidence" + assert_not_contains "$out" "state: working" "coarse active record cannot suppress socket refusal" + pass "socket refusal over a coarse active run reports blocked" +} + +# The coarse fallback has no steps table and no ci log, so the 2026-09-05 +# orphaned-monitor shape (every substantive step completed, only the ci +# monitor failed after the daemon restarted under its merge poll) cannot be +# recognized there. With the daemon provably down, that terminal failed +# record is unverified evidence from a dead instrument and must read unknown, +# never failed - the fleet rule from #3785. The fallback is reached while the +# daemon is answering for another branch, so the probe proves the daemon +# went down after that answer (a flapping daemon under incident load) - the +# two calls are separate socket connections. With the daemon up, the same +# record keeps its failure verdict. +test_coarse_failed_ledger_with_daemon_down_reports_unknown() { + reset_fakes + local d short; d=$(new_case coarse-daemon-down-failed) + make_repo_on_branch "$d/wt" fm/feat-coarsedown + short=$(git -C "$d/wt" rev-parse --short=7 HEAD) + make_fakebin "$d" >/dev/null + fm_write_meta "$d/state/feat-coarsedown.meta" "window=fm:fm-feat-coarsedown" "worktree=$d/wt" "kind=ship" + # The primary `axi status` call answers (another crew's run - the shared + # daemon serves the whole repo), so attribution falls to the coarse runs + # ledger, whose newest row for this branch is terminal failed at this + # worktree's own head. + FM_FAKE_AXI_STATUS="$(run_running fm/other-crew)" + FM_FAKE_RUNS_LIST=" failed fm/feat-coarsedown ${short} 2026-09-05 21:00" + FM_FAKE_DAEMON_DOWN=1 + local out; out=$(run_crew_state "$d" feat-coarsedown) + assert_contains "$out" "state: unknown" "daemon down + failed ledger record -> unknown" + assert_contains "$out" "no-mistakes daemon unreachable; last ledger record failed - unverified" \ + "the unverified detail names the dead instrument" + assert_not_contains "$out" "state: failed" "an instrument failure never reads as work failure" + assert_contains "$out" "source: run-step" "the ledger row is still this branch's attributed run" + + # Daemon provably up again: the same row stays a failure. + FM_FAKE_DAEMON_DOWN=0 + out=$(run_crew_state "$d" feat-coarsedown) + assert_contains "$out" "state: failed" "daemon up keeps the failed verdict over the failed record" + assert_not_contains "$out" "unverified" "no unverified qualifier while the daemon answers" + pass "failed ledger record reads unknown only when the daemon is provably down" +} + test_cross_branch_attribution_picks_most_recent_row() { reset_fakes local d short; d=$(new_case crossbranch-mostrecent) @@ -1873,6 +2234,11 @@ EOF test_active_run_is_authoritative test_stale_needs_decision_superseded test_stale_blocked_superseded +test_daemon_claim_over_live_run_reads_run_alive +test_socket_refusal_over_stale_fixing_run_reports_blocked +test_socket_refusal_over_terminal_run_reports_blocked +test_ordinary_blocked_over_live_run_keeps_plain_superseded +test_genuine_daemon_down_reports_blocked test_genuine_parked_not_superseded test_scalar_gate_parked_not_superseded test_gate_block_parked_not_superseded @@ -1890,7 +2256,13 @@ test_top_level_fixing_ci_running_after_green_stays_working test_top_level_fixing_done_log_stays_working test_terminal_passed test_terminal_failed +test_terminal_failed_ci_orphan_after_green_reads_done +test_terminal_failed_ci_orphan_status_only_reads_done +test_terminal_failed_ci_genuine_red_stays_failed +test_terminal_failed_ci_orphan_second_failed_step_stays_failed test_cross_branch_attribution_via_runs_list +test_coarse_socket_refusal_reports_blocked +test_coarse_failed_ledger_with_daemon_down_reports_unknown test_cross_branch_attribution_picks_most_recent_row test_coarse_run_does_not_probe_other_branch_ci_log_for_ready_status test_other_branch_run_ignored diff --git a/tests/fm-cursor-primary-live-e2e.test.sh b/tests/fm-cursor-primary-live-e2e.test.sh index ad806069986..159dee1674d 100755 --- a/tests/fm-cursor-primary-live-e2e.test.sh +++ b/tests/fm-cursor-primary-live-e2e.test.sh @@ -24,19 +24,15 @@ # path and is the only state left outside the temp dir. set -u -if [ "${FM_CURSOR_PRIMARY_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_CURSOR_PRIMARY_LIVE_E2E=1 to run the live Cursor primary guard" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_CURSOR_PRIMARY_LIVE_E2E tmux jq + +REAL_TMUX=$(command -v tmux) CURSOR_BIN=${FM_CURSOR_BIN:-$(command -v cursor-agent || true)} [ -n "$CURSOR_BIN" ] && [ -x "$CURSOR_BIN" ] \ || fail "cursor-agent not found; install it or set FM_CURSOR_BIN. This guard refuses to pass without checking the real harness." -REAL_TMUX=$(command -v tmux) || fail "tmux not found" -command -v jq >/dev/null 2>&1 || fail "jq not found" CURSOR_VERSION=$("$CURSOR_BIN" --version 2>/dev/null | head -1) [ -n "$CURSOR_VERSION" ] || fail "cursor-agent did not report a version; refusing to claim a verified result" printf 'harness: cursor-agent %s\n' "$CURSOR_VERSION" diff --git a/tests/fm-gate-refuse.test.sh b/tests/fm-gate-refuse.test.sh index 538bd21a799..8ef32c14949 100755 --- a/tests/fm-gate-refuse.test.sh +++ b/tests/fm-gate-refuse.test.sh @@ -206,6 +206,7 @@ case "${1:-}" in for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done printf '%%1\n'; exit 0 ;; capture-pane) printf '╭────╮\n│ │\n╰────╯\n'; exit 0 ;; + list-windows) printf 'fm-lane-ok\n'; exit 0 ;; esac exit 0 SH @@ -256,7 +257,7 @@ test_send_refuses_and_admits() { || fail "send: normal steer was not durably enqueued" assert_not_contains "$(cat "$log")" "literal=1 arg=hello captain" \ "send: normal steer payload must not be typed" - assert_contains "$(cat "$log")" "target=sess:fm-lane-ok literal=1 arg=Firstmate instruction waiting" \ + assert_contains "$(cat "$log")" "target=sess:fm-lane-ok literal=1 arg=: Firstmate instruction waiting" \ "send: normal steer should ring the durable inbox doorbell" pass "fm-send: refuses on marker and gate-worktree backstop; a normal steer uses the inbox" } diff --git a/tests/fm-gotmp.test.sh b/tests/fm-gotmp.test.sh index 43b15c6e818..d29c540e648 100755 --- a/tests/fm-gotmp.test.sh +++ b/tests/fm-gotmp.test.sh @@ -108,6 +108,8 @@ SH # fused backlog close is skipped and the follow-up echo takes the plain-message # path; there is no tasks-axi and no backlog in this fixture. cat > "$fake/bin/fm-tasks-axi-lib.sh" <<'SH' +FM_TASKS_AXI_MIN=0.2.4 +fm_tasks_axi_backend() { printf 'markdown\n'; } fm_tasks_axi_backend_available() { return 1; } fm_tasks_axi_compatible() { return 1; } fm_backlog_backend_manual() { return 1; } @@ -198,6 +200,8 @@ exit 0 SH chmod +x "$fake/bin/fm-fleet-sync.sh" cat > "$fake/bin/fm-tasks-axi-lib.sh" <<'SH' +FM_TASKS_AXI_MIN=0.2.4 +fm_tasks_axi_backend() { printf 'markdown\n'; } fm_tasks_axi_backend_available() { return 1; } fm_tasks_axi_compatible() { return 1; } fm_backlog_backend_manual() { return 1; } diff --git a/tests/fm-grok-continuity-live-e2e.test.sh b/tests/fm-grok-continuity-live-e2e.test.sh index 2caf9ab74f8..a4d8052a4ca 100755 --- a/tests/fm-grok-continuity-live-e2e.test.sh +++ b/tests/fm-grok-continuity-live-e2e.test.sh @@ -3,10 +3,10 @@ # through Grok's tracked background-task notification path. set -u -if [ "${FM_GROK_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_GROK_LIVE_E2E=1 to run the interactive Grok continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_GROK_LIVE_E2E grok tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -15,9 +15,6 @@ fail() { exit 1 } -command -v grok >/dev/null 2>&1 || fail "grok not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" - TMUX=$(command -v tmux) SOCKET="fm-grok-live-e2e-$$" SESSION=grok-live-e2e diff --git a/tests/fm-grok-stop-live-e2e.test.sh b/tests/fm-grok-stop-live-e2e.test.sh index c9d4bcc9a64..b29cee109cb 100755 --- a/tests/fm-grok-stop-live-e2e.test.sh +++ b/tests/fm-grok-stop-live-e2e.test.sh @@ -7,14 +7,11 @@ # window. Cleanup uses only creation-time pane/process identities. set -u -if [ "${FM_GROK_STOP_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_GROK_STOP_LIVE_E2E=1 with FM_GROK_NATIVE_BIN and FM_GROK_LEGACY_BIN" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_GROK_STOP_LIVE_E2E tmux jq + NATIVE_BIN=${FM_GROK_NATIVE_BIN:-} LEGACY_BIN=${FM_GROK_LEGACY_BIN:-} AUTH=${FM_GROK_AUTH_FILE:-$HOME/.grok/auth.json} @@ -25,7 +22,6 @@ ACTIVE_LAB= [ -x "$LEGACY_BIN" ] || fail "FM_GROK_LEGACY_BIN must be an exact executable path" [ -f "$AUTH" ] || fail "FM_GROK_AUTH_FILE must name the already-managed auth artifact" [ -n "$REAL_TMUX" ] || fail "tmux not found" -command -v jq >/dev/null 2>&1 || fail "jq not found" NATIVE_VERSION=$($NATIVE_BIN --version) LEGACY_VERSION=$($LEGACY_BIN --version) diff --git a/tests/fm-harness-adapter-instructions-live-e2e.test.sh b/tests/fm-harness-adapter-instructions-live-e2e.test.sh index 5b693fc0775..fcab6502152 100644 --- a/tests/fm-harness-adapter-instructions-live-e2e.test.sh +++ b/tests/fm-harness-adapter-instructions-live-e2e.test.sh @@ -7,22 +7,17 @@ # unconfigured native harness loaded the references itself. set -u -if [ "${FM_HARNESS_ADAPTER_INSTRUCTION_EVAL:-0}" != 1 ]; then - echo "skip: set FM_HARNESS_ADAPTER_INSTRUCTION_EVAL=1 and FM_HARNESS_ADAPTER_LOCAL_MODEL=<model> to run the local instruction evaluation" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +fm_live_gate opt-in FM_HARNESS_ADAPTER_INSTRUCTION_EVAL curl jq + ROUTER="$ROOT/.agents/skills/harness-adapters/SKILL.md" TMP_ROOT=$(fm_test_tmproot fm-harness-adapter-instructions) EXPECTED_JSON="$TMP_ROOT/expected.json" PROMPT_FILE="$TMP_ROOT/prompt.txt" RESPONSE_JSON="$TMP_ROOT/response.json" -command -v curl >/dev/null 2>&1 || fail "curl is required for the local instruction evaluation" -command -v jq >/dev/null 2>&1 || fail "jq is required for the local instruction evaluation" curl -fsS --max-time 2 http://127.0.0.1:11434/api/tags > "$TMP_ROOT/tags.json" \ || fail "local Ollama is unavailable at 127.0.0.1:11434; no remote provider fallback is allowed" diff --git a/tests/fm-harness-liveness-drift-live-e2e.test.sh b/tests/fm-harness-liveness-drift-live-e2e.test.sh index e0203b85b8a..16d669369c1 100755 --- a/tests/fm-harness-liveness-drift-live-e2e.test.sh +++ b/tests/fm-harness-liveness-drift-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# tests/fm-harness-liveness-drift-live-e2e.test.sh - opt-in drift guard proving +# tests/fm-harness-liveness-drift-live-e2e.test.sh - default-on drift guard proving # every INSTALLED harness is still classified `alive` by the tmux liveness # probe (bin/backends/tmux.sh). # @@ -16,16 +16,17 @@ # unauthenticated harness still starts its process, which is all the liveness # probe reads. # -# Standard CI has no harness binaries or credentials, so this real-harness guard -# is opt-in and on-demand. The portable counterpart in +# Portable serial CI installs the public Pi package but no credentials, so this +# guard checks that available token-free surface there and runs against every installed +# harness on more capable hosts. The portable counterpart in # tests/fm-tmux-agent-liveness.test.sh pins the classifier logic in CI. Run this # guard after any harness upgrade and before trusting refreshed evidence. set -u -if [ "${FM_HARNESS_LIVENESS_DRIFT:-0}" != 1 ]; then - echo "skip: set FM_HARNESS_LIVENESS_DRIFT=1 to run the installed-harness liveness drift guard" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate default-on FM_HARNESS_LIVENESS_DRIFT tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" @@ -33,7 +34,6 @@ fail() { printf 'not ok - %s\n' "$1" >&2; cleanup_all; exit 1; } pass() { printf 'ok - %s\n' "$1"; } note() { printf '# %s\n' "$1"; } -command -v tmux >/dev/null 2>&1 || fail "tmux not found" REAL_TMUX=$(command -v tmux) SOCKET="fm-liveness-drift-$$" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-liveness-drift.XXXXXX") @@ -67,6 +67,18 @@ fm_backend_source tmux || fail "fm_backend_source tmux failed" # order so this guard covers the same binary firstmate would actually launch. resolve_harness_binary() { # <harness> local harness=$1 candidate + # cursor is resolved FIRST, before the generic PATH lookup, and only through + # the verified owner fm-spawn uses. The Cursor agent never installs as + # `cursor`: it installs as `cursor-agent` plus the legacy alias `agent`. A + # machine that also has the Cursor editor does have an executable `cursor` on + # PATH, and launching that one exits immediately, leaving a bare shell in the + # pane that this guard then reports as liveness drift the classifier can do + # nothing about. Asking the owner first also keeps an unrelated executable + # named `agent` rejected here exactly as it would be at launch. + if [ "$harness" = cursor ]; then + fm_cursor_resolve_binary 2>/dev/null && return 0 + return 1 + fi candidate=$(command -v "$harness" 2>/dev/null || true) if [ -n "$candidate" ] && [ -x "$candidate" ]; then printf '%s\n' "$candidate" @@ -76,15 +88,6 @@ resolve_harness_binary() { # <harness> printf '%s\n' "$HOME/.kimi-code/bin/kimi" return 0 fi - # cursor is never on PATH under the name `cursor`: it installs as - # `cursor-agent` plus the legacy alias `agent`, and its user-local install is - # routinely absent from a non-interactive PATH. Resolve it through the same - # verified owner fm-spawn uses, so an unrelated executable named `agent` is - # rejected here exactly as it would be at launch. - if [ "$harness" = cursor ]; then - fm_cursor_resolve_binary 2>/dev/null && return 0 - return 1 - fi return 1 } diff --git a/tests/fm-herdr-submit-confirm-live-e2e.test.sh b/tests/fm-herdr-submit-confirm-live-e2e.test.sh index 9140fec1e6e..c7813b28393 100755 --- a/tests/fm-herdr-submit-confirm-live-e2e.test.sh +++ b/tests/fm-herdr-submit-confirm-live-e2e.test.sh @@ -14,20 +14,17 @@ # Every Herdr call, including adapter calls, is routed through bin/fm-herdr-lab.sh. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } -if [ "${FM_HERDR_SUBMIT_CONFIRM_LIVE:-0}" != 1 ]; then - echo "skip: set FM_HERDR_SUBMIT_CONFIRM_LIVE=1 to run the live Herdr submit-confirmation guard" - exit 0 -fi +fm_live_gate opt-in FM_HERDR_SUBMIT_CONFIRM_LIVE herdr jq claude -command -v herdr >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but herdr is not installed" -command -v jq >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but jq is not installed" -command -v claude >/dev/null 2>&1 || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but Claude Code is not installed" [ -x "$LAB_HELPER" ] || fail "FM_HERDR_SUBMIT_CONFIRM_LIVE=1 but the Herdr lab helper is not executable at $LAB_HELPER" # shellcheck source=tests/herdr-test-safety.sh diff --git a/tests/fm-herdr-version-floor-live-e2e.test.sh b/tests/fm-herdr-version-floor-live-e2e.test.sh index 24463b520dd..5f721214d6e 100755 --- a/tests/fm-herdr-version-floor-live-e2e.test.sh +++ b/tests/fm-herdr-version-floor-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Opt-in live guard for the Herdr presentation version floor. +# Default-on live guard for the Herdr presentation version floor. # # Protocol is the floor's structural signal, and its mapping to real releases # is a vendor-supplied fact that no fixture can prove. The runtime gate checks @@ -10,7 +10,8 @@ # identity the binary actually reports. It fails naming the version and protocol # rather than degrading quietly. # -# It is opt-in because it downloads upstream release binaries over the network. +# A run downloads upstream release binaries over the network but submits no +# prompt, so the shared live gate runs it by default wherever its tools exist. # Run it after every Herdr upgrade and before trusting a refreshed # docs/verification/runtime-backends.md "Presentation version floor" entry. # @@ -20,20 +21,17 @@ # lifecycle operation and no server is involved. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} fail() { printf 'not ok - %s\n' "$1" >&2; exit 1; } pass() { printf 'ok - %s\n' "$1"; } -if [ "${FM_HERDR_VERSION_FLOOR_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_HERDR_VERSION_FLOOR_LIVE_E2E=1 to run the real-release Herdr version-floor guard" - exit 0 -fi +fm_live_gate default-on FM_HERDR_VERSION_FLOOR_LIVE_E2E herdr jq curl shasum -for tool in herdr jq curl shasum; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done [ -x "$LAB_HELPER" ] || { echo "skip: Herdr lab helper not executable at $LAB_HELPER"; exit 0; } case "$(uname -s)/$(uname -m)" in diff --git a/tests/fm-kimi-harness.test.sh b/tests/fm-kimi-harness.test.sh index 3c540856531..672ff1ffda3 100755 --- a/tests/fm-kimi-harness.test.sh +++ b/tests/fm-kimi-harness.test.sh @@ -222,6 +222,8 @@ test_kimi_launch_then_send_is_verified() { assert_present "$task_tmp/gotmp" "kimi spawn did not create its Go temp directory" assert_grep "export GOTMPDIR=$task_tmp/gotmp" "$CASE_DIR/tmux-calls.log" \ "kimi spawn did not export its Go temp directory into the pane" + assert_grep "export FM_TASK_ID=$id" "$CASE_DIR/tmux-calls.log" \ + "kimi spawn did not mark the pane with its task id" assert_grep 'BEGIN FIRSTMATE KIMI TURN-END HOOK' "$HOME_DIR/.kimi-code/config.toml" \ "kimi spawn did not install its guarded global hook region" assert_grep 'token=' "$WT_DIR/.fm-kimi-turnend" "kimi spawn did not write its token pointer" diff --git a/tests/fm-lint.test.sh b/tests/fm-lint.test.sh index 8041c76628e..6de72c97bcf 100755 --- a/tests/fm-lint.test.sh +++ b/tests/fm-lint.test.sh @@ -159,6 +159,10 @@ test_help_reports_the_complete_interface() { assert_contains "$help" "--list-files" "fm-lint.sh --help omitted --list-files" assert_contains "$help" "--help" "fm-lint.sh --help omitted --help" assert_contains "$help" "--fast" "fm-lint.sh --help omitted --fast" + assert_contains "$help" "SC1091" "fm-lint.sh --help omitted the local SC1091 exclusion" + assert_contains "$help" "SC2034" "fm-lint.sh --help omitted the local SC2034 exclusion" + assert_contains "$help" "SC2153" "fm-lint.sh --help omitted the local SC2153 exclusion" + assert_contains "$help" "SC2329" "fm-lint.sh --help omitted the local SC2329 exclusion" pass "fm-lint.sh --help reports the complete executable interface" } @@ -240,7 +244,8 @@ fm_lint_write_diff_file() { # them, so changed-file mode tests can assert exactly which files fm-lint.sh # selected without depending on real ShellCheck findings. When # FM_TEST_MODE_LOG is set, it records the effective analysis mode, treating -# ShellCheck's default as full analysis. +# ShellCheck's default as full analysis. When FM_TEST_FLAG_LOG is set, it +# records whether --external-sources was passed and the --exclude value. fm_lint_stub_shellcheck() { local fakebin=$1 log=$2 : > "$log" @@ -251,13 +256,26 @@ if [ "\${1:-}" = --version ]; then exit 0 fi mode=on +follow=no +exclude=none while [ "\$#" -gt 0 ] && [ "\$1" != -- ]; do - [ "\$1" = --extended-analysis=false ] && mode=off + case "\$1" in + --extended-analysis=false) mode=off ;; + --external-sources) follow=yes ;; + --exclude=*) exclude=\${1#--exclude=} ;; + --exclude) + shift + exclude=\${1:-none} + ;; + esac shift done if [ -n "\${FM_TEST_MODE_LOG:-}" ]; then printf '%s\n' "\$mode" >> "\$FM_TEST_MODE_LOG" fi +if [ -n "\${FM_TEST_FLAG_LOG:-}" ]; then + printf 'external-sources=%s\nexclude=%s\n' "\$follow" "\$exclude" >> "\$FM_TEST_FLAG_LOG" +fi [ "\$#" -eq 0 ] || shift printf '%s\n' "\$@" >> "$log" exit 0 @@ -471,6 +489,282 @@ test_list_files_respects_changed_mode() { pass "fm-lint.sh --list-files reports the would-be changed set in changed mode" } +fm_lint_assert_flag_log() { + local flag_log=$1 expected_follow=$2 expected_exclude=$3 + [ -s "$flag_log" ] || fail "ShellCheck was not invoked; flag log is empty" + awk -v follow="$expected_follow" -v exclude="$expected_exclude" ' + BEGIN { bad=0; saw=0 } + /^external-sources=/ { saw=1; if ($0 != "external-sources=" follow) bad=1 } + /^exclude=/ { if ($0 != "exclude=" exclude) bad=1 } + END { exit (saw && !bad) ? 0 : 1 } + ' "$flag_log" \ + || fail "ShellCheck flags were not external-sources=$expected_follow exclude=$expected_exclude"$'\n'"$(cat "$flag_log")" +} + +test_changed_mode_drops_external_sources_and_excludes_cross_file_codes() { + local tmp fakebin log flag_log mode_log diff_file telemetry out target + tmp=$(fm_test_tmproot fm-lint-local-nox) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + mode_log="$tmp/mode.log" + telemetry="$tmp/telemetry.tsv" + fm_lint_stub_shellcheck "$fakebin" "$log" + diff_file="$tmp/diff.nul" + target="bin/fm-afk-launch.sh" + fm_lint_write_diff_file "$diff_file" "$target" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature \ + FM_TEST_GIT_DIFF_FILE="$diff_file" \ + FM_TEST_FLAG_LOG="$flag_log" FM_TEST_MODE_LOG="$mode_log" \ + "$LINT" --telemetry "$telemetry" 2>&1) \ + || fail "changed-mode local lint failed"$'\n'"$out" + [ "$(cat "$log")" = "$target" ] \ + || fail "changed-mode lint did not run ShellCheck on exactly the changed file"$'\n'"logged: $(cat "$log")" + [ "$(cat "$mode_log")" = on ] \ + || fail "changed-mode local lint disabled dataflow analysis" + fm_lint_assert_flag_log "$flag_log" no "SC1091,SC2034,SC2153,SC2329" + assert_contains "$out" "source following disabled" \ + "changed-mode local lint did not disclose dropped source following" + assert_grep $'analysis_mode\tlocal' "$telemetry" \ + "telemetry did not record local analysis mode" + assert_grep $'source_directives\t3' "$telemetry" \ + "telemetry did not count the changed root's source directives" + assert_grep $'source_followed_directives\t0' "$telemetry" \ + "telemetry reported followed sources in no-external-sources mode" + pass "fm-lint.sh changed mode drops source following and excludes cross-file codes" +} + +test_changed_mode_invokes_shellcheck_once_per_root() { + local tmp fakebin log flag_log diff_file out first second invocation_count + tmp=$(fm_test_tmproot fm-lint-local-per-root) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + fm_lint_stub_shellcheck "$fakebin" "$log" + diff_file="$tmp/diff.nul" + first="bin/fm-install-shellcheck.sh" + second="bin/fm-lint-workflows.sh" + fm_lint_write_diff_file "$diff_file" "$first" "$second" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature FM_TEST_GIT_DIFF_FILE="$diff_file" \ + FM_TEST_FLAG_LOG="$flag_log" "$LINT" 2>&1) \ + || fail "changed-mode per-root lint failed"$'\n'"$out" + [ "$(LC_ALL=C sort "$log")" = "$first"$'\n'"$second" ] \ + || fail "changed-mode lint did not analyze both changed roots"$'\n'"logged: $(cat "$log")" + invocation_count=$(grep -c '^external-sources=' "$flag_log" || true) + [ "$invocation_count" -eq 2 ] \ + || fail "changed-mode lint used $invocation_count ShellCheck calls for two roots" + fm_lint_assert_flag_log "$flag_log" no "SC1091,SC2034,SC2153,SC2329" + pass "fm-lint.sh changed mode invokes ShellCheck once per root" +} + +test_ci_keeps_external_sources_without_local_exclusions() { + local tmp fakebin log flag_log mode_log fixture out + tmp=$(fm_test_tmproot fm-lint-ci-follow) + fakebin=$(fm_fakebin "$tmp") + fixture="$tmp/fixture.sh" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + mode_log="$tmp/mode.log" + cat > "$fixture" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${1:-ok}" +SH + fm_lint_stub_shellcheck "$fakebin" "$log" + + out=$(PATH="$fakebin:$PATH" CI=true GITHUB_ACTIONS=true FM_LINT_JOBS=1 \ + FM_TEST_FLAG_LOG="$flag_log" FM_TEST_MODE_LOG="$mode_log" \ + "$LINT" "$fixture" 2>&1) \ + || fail "CI lint with explicit path failed"$'\n'"$out" + [ "$(cat "$mode_log")" = on ] \ + || fail "CI lint disabled dataflow analysis" + fm_lint_assert_flag_log "$flag_log" yes none + pass "fm-lint.sh CI keeps source following without the local exclusion list" +} + +test_main_branch_keeps_external_sources() { + local tmp fakebin log flag_log out + tmp=$(fm_test_tmproot fm-lint-main-follow) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + fm_lint_stub_shellcheck "$fakebin" "$log" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=main \ + FM_TEST_FLAG_LOG="$flag_log" "$LINT" 2>&1) \ + || fail "main-branch lint failed"$'\n'"$out" + fm_lint_assert_flag_log "$flag_log" yes none + pass "fm-lint.sh on main keeps source following without the local exclusion list" +} + +test_merge_base_less_keeps_external_sources() { + local tmp fakebin log flag_log out + tmp=$(fm_test_tmproot fm-lint-nomergebase-follow) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + fm_lint_stub_shellcheck "$fakebin" "$log" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature FM_TEST_GIT_MERGE_BASE_OK=0 \ + FM_TEST_FLAG_LOG="$flag_log" "$LINT" 2>&1) \ + || fail "merge-base-less lint failed"$'\n'"$out" + fm_lint_assert_flag_log "$flag_log" yes none + pass "fm-lint.sh without a merge-base keeps source following without the local exclusion list" +} + +test_explicit_path_keeps_external_sources() { + local tmp fakebin log flag_log out target + tmp=$(fm_test_tmproot fm-lint-explicit-follow) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + fm_lint_stub_shellcheck "$fakebin" "$log" + target="bin/fm-install-shellcheck.sh" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature \ + FM_TEST_FLAG_LOG="$flag_log" "$LINT" "$target" 2>&1) \ + || fail "explicit-path lint failed"$'\n'"$out" + fm_lint_assert_flag_log "$flag_log" yes none + pass "fm-lint.sh explicit paths keep source following" +} + +test_fast_mode_on_a_local_branch_keeps_source_following() { + local tmp fakebin log flag_log mode_log diff_file out target + tmp=$(fm_test_tmproot fm-lint-fast-follow) + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + log="$tmp/shellcheck.log" + flag_log="$tmp/flags.log" + mode_log="$tmp/mode.log" + fm_lint_stub_shellcheck "$fakebin" "$log" + diff_file="$tmp/diff.nul" + target="bin/fm-install-shellcheck.sh" + fm_lint_write_diff_file "$diff_file" "$target" + + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature \ + FM_TEST_GIT_DIFF_FILE="$diff_file" \ + FM_TEST_FLAG_LOG="$flag_log" FM_TEST_MODE_LOG="$mode_log" \ + "$LINT" --fast 2>&1) \ + || fail "fast local-branch lint failed"$'\n'"$out" + [ "$(cat "$mode_log")" = off ] \ + || fail "fast local-branch lint did not disable extended analysis" + fm_lint_assert_flag_log "$flag_log" yes none + pass "fm-lint.sh --fast on a local branch keeps source following" +} + +test_changed_mode_hides_cross_file_codes_that_ci_still_sees() { + if ! pinned_ready; then + pass "SKIP (ShellCheck $REQUIRED not resolved): changed-mode exclusion behavior" + return + fi + local tmp fakebin diff_file fixture out rc + tmp=$(fm_test_tmproot fm-lint-local-exclude-behavior) + fixture="$ROOT/tests/fm-lint-local-exclude-fixture.test.sh" + printf '%s\n' "$fixture" >> "$FM_TEST_CLEANUP_REGISTRY" + cat > "$fixture" <<'SH' +#!/usr/bin/env bash +# Assigned here and only consumed by a library the local gate does not follow. +cross_file_only=1 +outer() { + ( + # Defined here and only invoked by a library the local gate does not follow. + cross_file_helper() { + printf 'ok\n' + } + printf 'hi\n' + ) +} +outer +SH + fakebin=$(fm_fakebin "$tmp") + fm_lint_stub_git "$fakebin" + diff_file="$tmp/diff.nul" + fm_lint_write_diff_file "$diff_file" "tests/fm-lint-local-exclude-fixture.test.sh" + + rc=0 + out=$(PATH="$fakebin:$PATH" GITHUB_ACTIONS='' CI='' FM_LINT_JOBS=1 \ + FM_TEST_GIT_BRANCH=feature \ + FM_TEST_GIT_DIFF_FILE="$diff_file" "$LINT" 2>&1) || rc=$? + [ "$rc" -eq 0 ] \ + || fail "changed-mode local lint failed a cross-file-only fixture"$'\n'"$out" + assert_not_contains "$out" "SC2034" "changed-mode local lint still reported SC2034" + assert_not_contains "$out" "SC2329" "changed-mode local lint still reported SC2329" + + rc=0 + out=$("$LINT" "$fixture" 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "explicit-path lint passed a cross-file-only fixture"$'\n'"$out" + assert_contains "$out" "SC2034" "explicit-path lint did not keep SC2034" + assert_contains "$out" "SC2329" "explicit-path lint did not keep SC2329" + rm -f "$fixture" + pass "fm-lint.sh changed mode excludes cross-file codes that explicit paths still report" +} + +# One ShellCheck process per root. Passing the whole canonical set in a +# single invocation still follows in-set sources and is not the no-x posture. +fm_lint_nox_one_root() { + local index=$1 path=$2 outdir=$3 + shellcheck --norc --format gcc -- "$path" > "$outdir/$index" || true +} + +test_local_exclusion_list_covers_every_no_external_sources_code() { + if ! pinned_ready; then + pass "SKIP (ShellCheck $REQUIRED not resolved): local exclusion completeness" + return + fi + local tmp files_file out unexpected code path found i batch + local -a files + tmp=$(fm_test_tmproot fm-lint-nox-complete) + files_file="$tmp/files" + CI=true "$LINT" --list-files > "$files_file" + [ -s "$files_file" ] || fail "CI --list-files returned no canonical lint roots" + files=() + while IFS= read -r path; do + [ -n "$path" ] || continue + files+=("$path") + done < "$files_file" + [ "${#files[@]}" -gt 0 ] || fail "CI --list-files returned no readable lint roots" + mkdir -p "$tmp/gcc" + i=0 + batch=0 + for path in "${files[@]}"; do + i=$((i + 1)) + fm_lint_nox_one_root "$i" "$path" "$tmp/gcc" & + batch=$((batch + 1)) + if [ "$batch" -eq 4 ]; then + wait + batch=0 + fi + done + wait + found=$(find "$tmp/gcc" -type f | wc -l | tr -d '[:space:]') + [ "$found" = "${#files[@]}" ] \ + || fail "completeness sweep linted $found roots, expected ${#files[@]}" + out=$(cat "$tmp/gcc"/* 2>/dev/null || true) + unexpected= + while IFS= read -r code; do + [ -n "$code" ] || continue + case "$code" in + SC1091|SC2034|SC2153|SC2329) ;; + *) unexpected="${unexpected}${unexpected:+ }$code" ;; + esac + done < <(printf '%s\n' "$out" | sed -n 's/.*\[\(SC[0-9][0-9]*\)\].*/\1/p' | LC_ALL=C sort -u) + [ -z "$unexpected" ] \ + || fail "no-external-sources pass emitted codes outside the local exclusion list: $unexpected" + pass "local exclusion list covers every no-external-sources ShellCheck code" +} + test_pins_an_explicit_version() { [ -n "$REQUIRED" ] || fail "fm-lint.sh --required-version printed nothing" # The captain-agreed pin: adopt ShellCheck 0.11.0's rule set consistently, @@ -1021,3 +1315,12 @@ test_main_branch_forces_full_lint test_explicit_path_bypasses_changed_logic test_zero_changed_files_exits_clean test_list_files_respects_changed_mode +test_changed_mode_drops_external_sources_and_excludes_cross_file_codes +test_changed_mode_invokes_shellcheck_once_per_root +test_ci_keeps_external_sources_without_local_exclusions +test_main_branch_keeps_external_sources +test_merge_base_less_keeps_external_sources +test_explicit_path_keeps_external_sources +test_fast_mode_on_a_local_branch_keeps_source_following +test_changed_mode_hides_cross_file_codes_that_ci_still_sees +test_local_exclusion_list_covers_every_no_external_sources_code diff --git a/tests/fm-live-gate.test.sh b/tests/fm-live-gate.test.sh new file mode 100755 index 00000000000..c2e3b4e1ca1 --- /dev/null +++ b/tests/fm-live-gate.test.sh @@ -0,0 +1,220 @@ +#!/usr/bin/env bash +# Behavior tests for tests/lib.sh's fm_live_gate, the single decision every +# live-harness guard opens with, and for the wiring that makes that decision +# reach the whole family. +# +# The gate is what turns "a live guard exists" into "a live guard actually ran +# on the machine that has the harness", so the cases below drive it the way a +# guard does: real scripts, executed as separate processes, with a fakebin PATH +# standing in for a host that does or does not have the tool. Nothing here reads +# tests/lib.sh's source text. +# +# The family sweep at the end runs every real live guard with FM_LIVE=0 and +# requires the shared refusal line, which is the only way to prove each guard is +# wired to the shared gate rather than to a private env check of its own. It is +# cheap because a disabled gate exits before a guard touches a harness. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-live-gate) +BIN="$TMP_ROOT/bin" +mkdir -p "$BIN" + +# A stand-in for a harness this host has: present on the fakebin PATH, and +# nothing the gate can confuse with a real one. +cat > "$BIN/fmfakeharness" <<'SH' +#!/usr/bin/env bash +exit 0 +SH +chmod +x "$BIN/fmfakeharness" + +clean_env() { + env -i \ + HOME="${HOME:-$TMP_ROOT}" \ + PATH="${PATH:-/usr/bin:/bin}" \ + TMPDIR="${TMPDIR:-/tmp}" \ + LANG="${LANG:-C}" \ + TERM="${TERM:-dumb}" \ + FM_TEST_SKIP_ORPHAN_REAP=1 \ + "$@" +} + +# guard <name> <gate args...>: write a guard script that opens with the shared +# gate and, if the gate lets it through, reports that it ran. +guard() { + local name=$1 + shift + local path="$TMP_ROOT/$name.test.sh" + { + printf '#!/usr/bin/env bash\nset -u\n' + printf '. "%s/tests/lib.sh"\n' "$ROOT" + printf 'fm_live_gate' + printf ' %s' "$@" + printf '\n' + printf 'printf "ran\\n"\n' + } > "$path" + chmod +x "$path" + printf '%s\n' "$path" +} + +# run_guard <path> [env assignment ...]: execute a guard on a PATH that carries +# only the fakebin plus the system essentials, capturing stdout and stderr. +run_guard() { + local path=$1 + shift + local out rc + set +e + out=$(clean_env "$@" PATH="$BIN:/usr/bin:/bin" "$path" 2>&1) + rc=$? + set -e + printf '%s\n' "$rc" + printf '%s\n' "$out" +} + +test_default_on_runs_when_the_tool_is_installed() { + local path result + path=$(guard default-present default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path") + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "a default-on guard with its tool installed must succeed: $result" + assert_contains "$result" ran "a default-on guard must run with no variable set when its tool is installed" +} + +test_default_on_skips_and_names_the_absent_tool() { + local path result + path=$(guard default-absent default-on FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path") + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an absent tool must be a skip, not a failure: $result" + assert_contains "$result" "skip: live: fmmissingharness absent" \ + "a capability skip must name the tool this host does not have" + assert_not_contains "$result" ran "an absent tool must stop the guard before it runs" +} + +test_opt_in_stays_off_until_asked() { + local path result + path=$(guard optin-idle opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path") + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an unrequested opt-in guard must skip cleanly: $result" + assert_contains "$result" "skip: live: opt-in; set FM_FAKE_LIVE=1 to run" \ + "an opt-in skip must name the variable that turns the guard on" + assert_not_contains "$result" ran "a token-spending guard must not run unasked" +} + +test_own_variable_turns_an_opt_in_guard_on() { + local path result + path=$(guard optin-on opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_FAKE_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "FM_FAKE_LIVE=1 must run the guard: $result" + assert_contains "$result" ran "an explicitly requested opt-in guard must run" +} + +test_requested_run_fails_rather_than_skipping_on_an_absent_tool() { + local path result + path=$(guard optin-strict opt-in FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path" FM_FAKE_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 1 ] || fail "a demanded run with no tool must fail, not skip: $result" + assert_contains "$result" "FM_FAKE_LIVE was requested but fmmissingharness is not installed" \ + "the hard failure must name the request and the missing tool" + assert_not_contains "$result" "skip:" "a demanded run must never report itself as a skip" +} + +test_fm_live_turns_the_whole_family_off_and_on() { + local path result + path=$(guard fmlive-off default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=0) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "FM_LIVE=0 must skip cleanly: $result" + assert_contains "$result" "skip: live: disabled by FM_LIVE=0" "FM_LIVE=0 must say why it skipped" + + path=$(guard fmlive-on opt-in FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=1) + assert_contains "$result" ran "FM_LIVE=1 must turn an opt-in guard on" + + path=$(guard fmlive-on-strict opt-in FM_FAKE_LIVE fmmissingharness) + result=$(run_guard "$path" FM_LIVE=1) + [ "$(printf '%s' "$result" | sed -n 1p)" = 1 ] || fail "FM_LIVE=1 must make an absent tool a failure: $result" +} + +test_a_guards_own_variable_wins_over_fm_live() { + local path result + path=$(guard own-off default-on FM_FAKE_LIVE fmfakeharness) + result=$(run_guard "$path" FM_LIVE=1 FM_FAKE_LIVE=0) + [ "$(printf '%s' "$result" | sed -n 1p)" = 0 ] || fail "an explicit per-guard opt-out must skip cleanly: $result" + assert_contains "$result" "skip: live: disabled by FM_FAKE_LIVE=0" \ + "a guard's own 0 must win over FM_LIVE=1 and say so" + assert_not_contains "$result" ran "a guard switched off by name must not run" +} + +test_any_of_several_entry_points_turns_a_guard_on() { + local path result + path=$(guard multi opt-in FM_FAKE_LIVE,FM_FAKE_ALT_LIVE fmfakeharness) + result=$(run_guard "$path") + assert_contains "$result" "set FM_FAKE_LIVE=1 to run" \ + "a multi-entry guard must point at its primary variable when idle" + result=$(run_guard "$path" FM_FAKE_ALT_LIVE=1) + assert_contains "$result" ran "a secondary entry point must also turn the guard on" +} + +test_gate_lets_a_guard_drive_the_real_fleet_scripts_under_a_gate_marker() { + # The nine live guards that never sourced the shared helpers used to be + # refused by bin/fm-gate-refuse-lib.sh whenever the pipeline ran them, because + # the gate marker is set for every no-mistakes gate agent. Opening with the + # shared gate is what carries the test-suite bypass into them. + local path out rc + path="$TMP_ROOT/bypass.test.sh" + { + printf '#!/usr/bin/env bash\nset -u\n' + printf '. "%s/tests/lib.sh"\n' "$ROOT" + printf 'fm_live_gate default-on FM_FAKE_LIVE fmfakeharness\n' + printf '. "%s/bin/fm-gate-refuse-lib.sh"\n' "$ROOT" + printf 'if fm_is_gate_agent; then printf "refused\\n"; else printf "allowed\\n"; fi\n' + } > "$path" + chmod +x "$path" + set +e + out=$(clean_env NO_MISTAKES_GATE=1 PATH="$BIN:/usr/bin:/bin" "$path" 2>&1) + rc=$? + set -e + expect_code 0 "$rc" "a guard opened with the shared gate must not be refused" + assert_contains "$out" allowed \ + "the shared gate must carry the test-suite bypass so a live guard can drive the real fleet scripts" +} + +test_every_live_guard_is_wired_to_the_shared_gate() { + local script out listing checked=0 + listing=$("$ROOT/bin/fm-test-run.sh" --family live-harness-optin --list) \ + || fail "could not list the live-harness family" + while IFS= read -r script; do + [ -n "$script" ] || continue + script="$ROOT/$script" + # bin/fm-test-run.sh runs every script through bash, so a guard that is not + # marked executable is still a real suite member here. + out=$(clean_env FM_LIVE=0 bash "$script" 2>&1) || fail "$(basename "$script") must exit 0 when live guards are disabled" + assert_contains "$out" "skip: live: disabled by FM_LIVE=0" \ + "$(basename "$script") must open with the shared live gate" + checked=$((checked + 1)) + done <<EOF +$listing +EOF + [ "$checked" -ge 20 ] || fail "expected the whole live-guard family to be swept, saw only $checked" + pass "all $checked live guards refuse together on FM_LIVE=0" +} + +test_default_on_runs_when_the_tool_is_installed +pass "a default-on guard runs wherever its tools are installed" +test_default_on_skips_and_names_the_absent_tool +pass "an absent tool is a named capability skip, not a silent pass" +test_opt_in_stays_off_until_asked +pass "a prompt-submitting guard stays off until it is asked for" +test_own_variable_turns_an_opt_in_guard_on +pass "a guard's own variable turns it on" +test_requested_run_fails_rather_than_skipping_on_an_absent_tool +pass "a demanded run refuses to pass as a skip" +test_fm_live_turns_the_whole_family_off_and_on +pass "FM_LIVE switches the whole family" +test_a_guards_own_variable_wins_over_fm_live +pass "a guard's own setting wins over FM_LIVE" +test_any_of_several_entry_points_turns_a_guard_on +pass "any entry point of a multi-mode guard turns it on" +test_gate_lets_a_guard_drive_the_real_fleet_scripts_under_a_gate_marker +pass "the shared gate carries the gate-refusal bypass into every live guard" +test_every_live_guard_is_wired_to_the_shared_gate diff --git a/tests/fm-muse-harness.test.sh b/tests/fm-muse-harness.test.sh index a5436ac8b3e..a655b4d7ec7 100755 --- a/tests/fm-muse-harness.test.sh +++ b/tests/fm-muse-harness.test.sh @@ -912,6 +912,35 @@ test_muse_trusts_no_record_sources() { pass "muse trusts no busy record source" } +test_spawn_environment_allowlist_credential_preflight() { + local setting rec case_dir home proj wt fakebin id out status + for setting in withheld allowed stored; do + rec=$(make_spawn_case "allowlist-$setting") + IFS='|' read -r case_dir home proj wt fakebin id <<EOF +$rec +EOF + : > "$home/config/launch-env-allowlist" + case "$setting" in + allowed) printf 'META_API_KEY\n' > "$home/config/launch-env-allowlist" ;; + stored) + mkdir -p "$home/xdgconfig/muse" + printf '{"schema_version":1}\n' > "$home/xdgconfig/muse/auth.json" + ;; + esac + out=$(run_muse_spawn "$home" "$proj" "$wt" "$fakebin" "$id" --mode no-mistakes --yolo off) + status=$? + if [ "$setting" = withheld ]; then + expect_code 1 "$status" "withheld Muse key must not satisfy preflight" + assert_contains "$out" "no worker-reachable credential" "missing credential explanation" + assert_absent "$home/state/$id.meta" "withheld Muse key still launched a worker" + else + expect_code 0 "$status" "Muse $setting credential must remain usable: $out" + fi + done + pass "Muse preflight respects the allowlist while retaining stored authentication" +} + +test_spawn_environment_allowlist_credential_preflight test_detects_versioned_process_ancestor test_detection_is_anchored test_spawn_clears_inherited_foreign_harness_markers diff --git a/tests/fm-muse-signals-live-e2e.test.sh b/tests/fm-muse-signals-live-e2e.test.sh index 85874ac9da5..a406b9af601 100755 --- a/tests/fm-muse-signals-live-e2e.test.sh +++ b/tests/fm-muse-signals-live-e2e.test.sh @@ -1,6 +1,9 @@ #!/usr/bin/env bash set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" MUSE_BIN=$(command -v muse 2>/dev/null || true) REAL_TMUX=$(command -v tmux 2>/dev/null || true) @@ -116,14 +119,7 @@ if [ "${1:-}" = --ansi-self-test ]; then exit 0 fi -if [ "${FM_MUSE_SIGNALS_LIVE:-0}" != 1 ]; then - echo "skip: set FM_MUSE_SIGNALS_LIVE=1 to run the real Muse signal drift guard" - exit 0 -fi - -[ -x "$MUSE_BIN" ] || fail "FM_MUSE_SIGNALS_LIVE=1 but no real muse executable is installed on PATH" -[ -x "$REAL_TMUX" ] || fail "FM_MUSE_SIGNALS_LIVE=1 but tmux is not installed" -command -v node >/dev/null 2>&1 || fail "node is required to inspect Muse's serialized session protocol" +fm_live_gate opt-in FM_MUSE_SIGNALS_LIVE muse tmux node LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-muse-signals.XXXXXX") || fail "could not create the isolated Muse lab" trap cleanup EXIT diff --git a/tests/fm-nm-test-contract.test.sh b/tests/fm-nm-test-contract.test.sh new file mode 100755 index 00000000000..a7f17011fde --- /dev/null +++ b/tests/fm-nm-test-contract.test.sh @@ -0,0 +1,26 @@ +#!/usr/bin/env bash +# Contract: parsed .no-mistakes.yaml must leave commands.test absent or empty. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +NM="$ROOT/.no-mistakes.yaml" + +test_nm_has_no_deterministic_test_command() { + command -v ruby >/dev/null 2>&1 \ + || fail "ruby is required to parse .no-mistakes.yaml for this contract" + local val + val=$(ruby -ryaml -e ' +doc = YAML.load_file(ARGV[0]) || {} +cmds = doc["commands"] || {} +val = cmds.is_a?(Hash) ? cmds["test"] : nil +puts (val.nil? || val == false || val == "") ? "" : val.inspect +' "$NM") || fail "failed to parse .no-mistakes.yaml as YAML" + if [ -n "$val" ]; then + fail "commands.test must be absent or empty so Test stays intent-targeted; got: $val" + fi + pass "no-mistakes does not configure commands.test" +} + +test_nm_has_no_deterministic_test_command diff --git a/tests/fm-omp-harness.test.sh b/tests/fm-omp-harness.test.sh new file mode 100755 index 00000000000..0b3f597fedd --- /dev/null +++ b/tests/fm-omp-harness.test.sh @@ -0,0 +1,587 @@ +#!/usr/bin/env bash +# tests/fm-omp-harness.test.sh - the portable regression for the omp (Oh My Pi) +# adapter: detection, session-lock identity, tmux liveness classification, the +# spawn launch line and worker posture overlay, pre-launch model validation, the +# per-task busy-state extension, the extension supervision model and ownership +# proof, and the two tracked primary extensions driven over a fake omp API. +# +# omp's identity, launch, and lifecycle checks are HARNESS-DEPENDENT: their +# verdicts come from what the vendor emits (a process name, a settings schema, +# an extension event). This suite pins the LOGIC with real processes, a fake +# omp binary, and a plain Node host, so CI enforces it with no omp installed; +# FM_OMP_LIVE_E2E=1 tests/fm-omp-primary-live-e2e.test.sh is the live guard that +# catches vendor drift against a real omp. Neither replaces the other. +# +# The load-bearing contracts: +# 1. omp publishes no marker; the anchored process name `omp` is the ancestry +# evidence, and ompd/comp never identify. +# 2. FM_OMP_HARNESS=omp is a precedence override that needs a real omp +# ancestor: it beats an inherited CLAUDECODE under omp and is inert when it +# leaks into a worker whose ancestry holds no omp. +# 3. Every omp launch clears foreign markers, carries the tracked posture +# overlay, --auto-approve, --cwd, and (for a crewmate) one -e pointing at +# state/<id>.omp-ext.ts; a secondmate launch names no -e at all. +# 4. A <provider>/<id> model is validated only when `omp models --json` lists +# that provider; an unlisted provider passes through with a notice. +# 5. Busy state: agent_start is busy, agent_end with willContinue stays busy, +# a plain agent_end is idle, turn_end is a notification only. +# 6. The turn-end guard extension compels one continuation on exit 2 and +# stands down when the payload already carries stop_hook_active. +# 7. The watch extension arms through fm_watch_arm_omp and delivers an +# actionable close as one follow-up. +set -u + +# shellcheck source=tests/fixtures.sh +. "$(dirname "${BASH_SOURCE[0]}")/fixtures.sh" + +# shellcheck source=bin/fm-busy-lib.sh +. "$ROOT/bin/fm-busy-lib.sh" +# shellcheck source=bin/fm-composer-lib.sh +. "$ROOT/bin/fm-composer-lib.sh" +# shellcheck source=bin/fm-control-lib.sh +. "$ROOT/bin/fm-control-lib.sh" +# shellcheck source=bin/fm-session-lock-lib.sh +. "$ROOT/bin/fm-session-lock-lib.sh" + +HARNESS="$ROOT/bin/fm-harness.sh" +TMP_ROOT=$(fm_test_tmproot fm-omp-harness) +export NODE_NO_WARNINGS=1 + +# A process whose kernel-recorded identity is the bare name `omp`: a SYMLINK to +# the system shell, never a copy (a copied platform binary fails macOS code +# signing). macOS reports the symlink name through `ps -o comm=`, which is the +# exact signal under test. Every `-c` body below ends in a no-op so bash does +# not exec-optimize the single command away and replace the named process. +make_named_shells() { # <dir> -> echoes <bindir> + local dir=$1 name + mkdir -p "$dir" + for name in omp ompd comp; do + ln -sf /bin/bash "$dir/$name" + done + printf '%s' "$dir" +} + +# --- 1. Detection -------------------------------------------------------------- + +test_detection_anchored_name_and_marker_precedence() { + local bin out + bin=$(make_named_shells "$TMP_ROOT/named") + # shellcheck disable=SC2016 # the quoted body expands inside the named shell + out=$(env -u CLAUDECODE -u FM_OMP_HARNESS -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + "$bin/omp" -c '"$1"; :' _ "$HARNESS") + [ "$out" = omp ] || fail "a process named omp must detect as omp, got '$out'" + for decoy in ompd comp; do + # shellcheck disable=SC2016 # the quoted body expands inside the named shell + out=$(env -u CLAUDECODE -u FM_OMP_HARNESS -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + "$bin/$decoy" -c '"$1"; :' _ "$HARNESS") + [ "$out" != omp ] || fail "'$decoy' merely contains omp and must not detect as omp" + done + # The marker beats an inherited CLAUDECODE only under a real omp ancestor. + # shellcheck disable=SC2016 # the quoted body expands inside the named shell + out=$(env -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_OMP_HARNESS=omp \ + "$bin/omp" -c '"$1"; :' _ "$HARNESS") + [ "$out" = omp ] || fail "FM_OMP_HARNESS under an omp ancestor must outrank an inherited CLAUDECODE, got '$out'" + # ...and is inert when it leaks into a worker with no omp ancestor. + # shellcheck disable=SC2016 # the quoted body expands inside the named shell + out=$(env -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS CLAUDECODE=1 FM_OMP_HARNESS=omp \ + bash -c '"$1"; :' _ "$HARNESS") + [ "$out" = claude ] || fail "a leaked FM_OMP_HARNESS without an omp ancestor must not relabel a claude worker, got '$out'" + pass "fm-harness: omp detects by its anchored name; the marker is a precedence override that needs real omp ancestry" +} + +test_lock_identity_and_liveness_classification() { + fm_harness_process_matches omp '' || fail "session-lock identity must accept the exact omp name" + fm_harness_process_matches /usr/local/bin/omp 'omp --cwd /x' || fail "session-lock identity must accept an omp path" + ! fm_harness_process_matches ompd '' || fail "session-lock identity must not accept ompd" + ! fm_harness_process_matches comp '' || fail "session-lock identity must not accept comp" + # shellcheck source=bin/fm-backend.sh + . "$ROOT/bin/fm-backend.sh" + fm_backend_source tmux || fail "fm_backend_source tmux failed" + [ "$(fm_backend_tmux_classify_process_name omp)" = agent ] || fail "tmux liveness must classify omp as an agent" + [ "$(fm_backend_tmux_classify_process_name /opt/omp/bin/omp)" = agent ] || fail "tmux liveness must classify an omp path as an agent" + [ "$(fm_backend_tmux_classify_process_name ompd)" != agent ] || fail "tmux liveness must not classify ompd as an agent" + [ "$(fm_backend_tmux_classify_process_name comp)" != agent ] || fail "tmux liveness must not classify comp as an agent" + pass "session lock and tmux liveness: omp is anchored, decoys stay out" +} + +# --- 2. Launch --------------------------------------------------------------- + +# A fake omp that answers `models --json` with a two-provider catalog and exits +# 0 for everything else (the launch itself is only recorded by the fake tmux). +make_fake_omp() { # <fakebin> + cat > "$1/omp" <<'SH' +#!/usr/bin/env bash +case "$1" in + models) + printf '%s\n' '{"models":[{"provider":"openai-codex","id":"gpt-6-astra","selector":"openai-codex/gpt-6-astra"},{"provider":"ollama","id":"qwen3:8b","selector":"ollama/qwen3:8b"}]}' + ;; +esac +exit 0 +SH + chmod +x "$1/omp" +} + +make_spawn_case() { # <name> <harness> <id> + local name=$1 harness=$2 id=$3 case_dir home proj wt fakebin + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + wt="$case_dir/wt" + fakebin=$(make_spawn_fakebin "$case_dir/fake" claude) + make_fake_omp "$fakebin" + fm_test_spawn_home "$home" "$harness" + fm_git_worktree "$proj" "$wt" "wt-$name" + fm_test_spawn_brief "$home" "$id" + : > "$case_dir/launch.log" + printf '%s\n' "$case_dir|$home|$proj|$wt|$fakebin|$case_dir/launch.log" +} + +read_case_record() { + # shellcheck disable=SC2034 # CASE_DIR is part of the shared record shape + IFS='|' read -r CASE_DIR HOME_DIR PROJ_DIR WT_DIR FAKEBIN_DIR LAUNCH_LOG <<EOF +$1 +EOF +} + +run_scout_spawn() { # <home> <wt> <fakebin> <launch-log> <spawn-args...> + local home=$1 wt=$2 fakebin=$3 launchlog=$4 + shift 4 + FM_FAKE_LAUNCH_LOG="$launchlog" fm_test_run_spawn "$home" "$wt" "$fakebin" "$@" --scout +} + +test_spawn_launch_line_and_worker_wiring() { + local rec id=omp-launch-q1 out status launch state + rec=$(make_spawn_case launch omp "$id") + read_case_record "$rec" + out=$(run_scout_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness omp --model openai-codex/gpt-6-astra --effort medium) + status=$? + expect_code 0 "$status" "omp scout spawn should succeed: $out" + assert_contains "$out" "spawned $id harness=omp" "spawn did not report the omp harness" + state="$HOME_DIR/state" + assert_grep "harness=omp" "$state/$id.meta" "meta missing harness=omp" + assert_grep "model=openai-codex/gpt-6-astra" "$state/$id.meta" "meta missing the pinned model" + assert_grep "effort=medium" "$state/$id.meta" "meta missing the pinned effort" + assert_present "$state/$id.omp-ext.ts" "omp spawn did not write the per-task extension" + launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS -u GEMINI_CLI -u CURSOR_AGENT -u CURSOR_INVOKED_AS FM_OMP_HARNESS=omp OMP_SKIP_SETUP=1 '$FAKEBIN_DIR/omp'" \ + "omp launch did not clear foreign markers and establish its own at the launch boundary" + assert_contains "$launch" "--config '$ROOT/.omp/fm-worker-overlay.yml' --auto-approve --cwd '$WT_DIR'" \ + "omp launch did not carry the tracked posture overlay, --auto-approve, and the pinned working directory" + assert_contains "$launch" "--model 'openai-codex/gpt-6-astra' --thinking 'medium' -e '$state/$id.omp-ext.ts'" \ + "omp launch did not pass the model, thinking level, and the state-resident worker extension" + assert_contains "$launch" "encode launch-brief < '$HOME_DIR/data/$id/launch-brief.md'" "omp launch lost the canonical typed launch-brief envelope" + case "$launch" in + *"-e '$state/$id.omp-ext.ts' \"\$("*) ;; + *) fail "omp launch must keep exactly one positional brief after the extension flag: $launch" ;; + esac + [ "$(fm_busy_classify tmux fake:w omp "$id" "$state")" = "busy fm-spawn" ] \ + || fail "omp spawn must seed the busy-state contract" + pass "fm-spawn: the omp launch line clears markers, pins posture, and wires the state-resident extension" +} + +test_spawn_model_validation_scoped_to_listed_providers() { + local rec id out status + rec=$(make_spawn_case model-refused omp omp-model-refused-q2) + read_case_record "$rec" + id=omp-model-refused-q2 + out=$(run_scout_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness omp --model openai-codex/gpt-nope) + status=$? + expect_code 1 "$status" "a model absent from a listed provider must refuse" + assert_contains "$out" "is not listed by 'omp models --json' although provider 'openai-codex' is" "refusal did not name the listing evidence" + assert_absent "$HOME_DIR/state/$id.meta" "a refused spawn must publish no record" + + rec=$(make_spawn_case model-bridge omp omp-model-bridge-q3) + read_case_record "$rec" + id=omp-model-bridge-q3 + out=$(run_scout_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness omp --model claude-bridge/claude-opus-4-8) + status=$? + expect_code 0 "$status" "an extension-registered provider must pass through: $out" + assert_contains "$out" "notice: omp provider 'claude-bridge' is not in 'omp models --json'" "pass-through did not state its reason" + assert_contains "$(cat "$LAUNCH_LOG")" "--model 'claude-bridge/claude-opus-4-8'" "pass-through model did not reach the launch line" + + rec=$(make_spawn_case model-fuzzy omp omp-model-fuzzy-q4) + read_case_record "$rec" + id=omp-model-fuzzy-q4 + out=$(run_scout_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness omp --model astra) + status=$? + expect_code 0 "$status" "a bare fuzzy pattern is omp's own matcher's job: $out" + pass "fm-spawn: omp model validation is scoped to providers the listing can prove" +} + +test_secondmate_launch_relies_on_discovery() { + # A seeded secondmate home, launched for real through fm-spawn on omp: the + # launch must carry the posture overlay and pin --cwd to the home, and must + # name NO -e, because omp auto-discovers the home's tracked .omp/extensions + # and a file named both ways loads twice. + local world home fakebin launchlog out status launch + world="$TMP_ROOT/secondmate" + home="$world/sm" + mkdir -p "$world/home/state" "$world/home/data" "$world/home/config" "$home/bin" "$home/data" + printf '# Firstmate\n' > "$home/AGENTS.md" + printf 'sm\n' > "$home/.fm-secondmate-home" + printf 'charter\n' > "$home/data/charter.md" + fakebin=$(make_spawn_fakebin "$world/fake" claude) + make_fake_omp "$fakebin" + launchlog="$world/launch.log" + : > "$launchlog" + # FM_BACKEND=tmux pins the fake tmux even where the developer shell carries a + # live Herdr environment; without it auto-detection would spawn a real pane. + out=$(PATH="$fakebin:$PATH" TMUX='fake,1,0' FM_BACKEND=tmux CLAUDECODE=1 \ + FM_ROOT_OVERRIDE='' FM_HOME="$world/home" \ + FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ + FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ + FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ + "$ROOT/bin/fm-spawn.sh" sm "$home" omp --secondmate 2>&1) + status=$? + expect_code 0 "$status" "omp secondmate spawn should succeed: $out" + assert_grep "harness=omp" "$world/home/state/sm.meta" "secondmate meta missing harness=omp" + launch=$(cat "$launchlog") + case "$launch" in + *" -e "*) fail "an omp secondmate launch must name no -e: omp auto-discovers .omp/extensions and a file named both ways loads twice: $launch" ;; + esac + assert_contains "$launch" "--config '$ROOT/.omp/fm-worker-overlay.yml' --auto-approve --cwd '$home'" "secondmate launch lost the posture overlay or the pinned home directory: $launch" + assert_contains "$launch" "FM_OMP_HARNESS=omp OMP_SKIP_SETUP=1 '$fakebin/omp'" "secondmate launch lost the omp marker or executable" + assert_contains "$launch" "FM_SUPERVISION_MODEL=extension" "an omp secondmate must run the extension supervision model" + assert_absent "$world/home/state/sm.omp-ext.ts" "a secondmate must not receive a per-task worker extension" + pass "fm-spawn: a real omp secondmate launch relies on auto-discovery while crewmates load one -e" +} + +test_secondmate_config_pinned_model_is_validated() { + # The same seeded secondmate home, but the harness and model come from the + # primary's config/secondmate-harness rather than the command line: the + # durable pin lands on MODEL after the harness case arm, so an unlisted id + # under a listed provider must still be refused before endpoint creation. + local world home fakebin launchlog out status + world="$TMP_ROOT/secondmate-config-model" + home="$world/sm" + mkdir -p "$world/home/state" "$world/home/data" "$world/home/config" "$home/bin" "$home/data" + printf '# Firstmate\n' > "$home/AGENTS.md" + printf 'sm\n' > "$home/.fm-secondmate-home" + printf 'charter\n' > "$home/data/charter.md" + printf 'omp openai-codex/gpt-nope\n' > "$world/home/config/secondmate-harness" + fakebin=$(make_spawn_fakebin "$world/fake" claude) + make_fake_omp "$fakebin" + launchlog="$world/launch.log" + : > "$launchlog" + out=$(PATH="$fakebin:$PATH" TMUX='fake,1,0' FM_BACKEND=tmux CLAUDECODE=1 \ + FM_ROOT_OVERRIDE='' FM_HOME="$world/home" \ + FM_STATE_OVERRIDE="$world/home/state" FM_DATA_OVERRIDE="$world/home/data" \ + FM_PROJECTS_OVERRIDE="$world/home/projects" FM_CONFIG_OVERRIDE="$world/home/config" \ + FM_SPAWN_NO_GUARD=1 FM_FAKE_LAUNCH_LOG="$launchlog" \ + "$ROOT/bin/fm-spawn.sh" sm "$home" --secondmate 2>&1) + status=$? + expect_code 1 "$status" "a config-pinned unlisted omp model must refuse the secondmate spawn: $out" + assert_contains "$out" "omp model 'openai-codex/gpt-nope' is not listed by 'omp models --json' although provider 'openai-codex' is" \ + "the refusal did not name the config-pinned model under its listed provider: $out" + assert_absent "$world/home/state/sm.meta" "a refused secondmate spawn must publish no sm.meta" + [ ! -s "$launchlog" ] || fail "a refused secondmate spawn must record no launch: $(cat "$launchlog")" + pass "fm-spawn: the config/secondmate-harness model pin is validated against the omp catalog before launch" +} + +# --- 3. Busy state ------------------------------------------------------------- + +drive_omp_ext() { # <ext-path> <mode> + EXT_PATH="$1" MODE="$2" node --input-type=module 2>&1 <<'EOF' +import { pathToFileURL } from "node:url"; +const mod = await import(pathToFileURL(process.env.EXT_PATH).href); +const handlers = {}; +mod.default({ on: (name, fn) => { handlers[name] = fn; } }); +// ctx.isIdle() reads false at a natural TUI agent_end on omp; the extension +// must go idle on a plain agent_end regardless of it. +const ctx = { isIdle: () => false }; +switch (process.env.MODE) { + case "handlers": console.log(Object.keys(handlers).sort().join(" ")); break; + case "agent-start": await handlers["agent_start"]({ type: "agent_start" }, ctx); break; + case "end-continuing": await handlers["agent_end"]({ type: "agent_end", willContinue: true }, ctx); break; + case "end-final": await handlers["agent_end"]({ type: "agent_end" }, ctx); break; + case "turn-end": await handlers["turn_end"]({ type: "turn_end", turnIndex: 0 }, ctx); break; + default: throw new Error("unknown mode " + process.env.MODE); +} +if (process.env.MODE === "turn-end") { + await new Promise((resolve) => setTimeout(resolve, 200)); +} +EOF +} + +test_busy_extension_lifecycle() { + local rec id=omp-busy-q5 out state ext + rec=$(make_spawn_case busy omp "$id") + read_case_record "$rec" + out=$(run_scout_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --harness omp) + expect_code 0 $? "omp spawn should succeed: $out" + state="$HOME_DIR/state" + ext="$state/$id.omp-ext.ts" + assert_present "$ext" "omp spawn did not write the per-task extension" + out=$(drive_omp_ext "$ext" handlers) || fail "handler listing failed: $out" + case " $out " in + *" agent_settled "*) fail "the omp extension must not listen for agent_settled (omp has no such event)" ;; + esac + for handler in agent_start agent_end turn_end; do + case " $out " in + *" $handler "*) ;; + *) fail "the omp extension must register $handler, got '$out'" ;; + esac + done + + rm -f "$state/$id.turn-ended" + out=$(drive_omp_ext "$ext" turn-end) || fail "turn_end drive failed: $out" + [ -f "$state/$id.turn-ended" ] || fail "turn_end no longer touches the notification marker" + [ "$(fm_busy_classify tmux fake:w omp "$id" "$state")" = "busy fm-spawn" ] || fail "turn_end must stay a notification, not a state edge" + + out=$(drive_omp_ext "$ext" agent-start) || fail "agent_start drive failed: $out" + [ "$(fm_busy_classify tmux fake:w omp "$id" "$state")" = "busy omp-ext" ] || fail "agent_start must classify 'busy omp-ext'" + + out=$(drive_omp_ext "$ext" end-continuing) || fail "continuing agent_end drive failed: $out" + [ "$(fm_busy_classify tmux fake:w omp "$id" "$state")" = "busy omp-ext" ] || fail "agent_end with willContinue must stay busy (a session_stop continuation is coming)" + + out=$(drive_omp_ext "$ext" end-final) || fail "final agent_end drive failed: $out" + [ "$(fm_busy_classify tmux fake:w omp "$id" "$state")" = "idle omp-ext" ] || fail "a plain agent_end must classify 'idle omp-ext'" + + # A record from another harness's writer is never trusted for omp. + fm_busy_source_trusted omp pi-ext && fail "omp must not trust the Pi extension's records" + fm_busy_source_trusted omp omp-ext || fail "omp must trust its own extension's records" + pass "omp extension: agent_start busy, willContinue stays busy, plain agent_end idle, turn_end a notification" +} + +# --- 4. Control, composer, supervision model ----------------------------------- + +test_control_composer_and_model_tables() { + [ "$(fm_control_exit_command omp)" = /quit ] || fail "omp exit command must be /quit" + [ "$(fm_control_interrupt_key omp)" = Escape ] || fail "omp interrupt key must be Escape" + [ "$(fm_control_interrupt_repeat omp)" = 1 ] || fail "omp interrupts on a single press" + [ -z "$(fm_control_interrupt_clear_key omp)" ] || fail "omp leaves its composer empty and needs no clear key" + [ "$(fm_control_harness_wiring_paths omp /wt /st id1)" = "/st/id1.omp-ext.ts" ] || fail "omp wiring path must be the state-resident extension" + printf 'Working…\n' | fm_busy_lines_match omp || fail "omp busy regex must match the TUI ellipsis form" + printf 'Working...\n' | fm_busy_lines_match omp && fail "omp busy regex must not match the three-dot form no supervised pane renders" + printf ' ⠧ 11s · gpt-6-astra\n' | fm_busy_lines_match omp || fail "omp busy regex must match the braille spinner plus elapsed cell" + printf ' ⣾ 3s · gpt-6-astra\n' | fm_busy_lines_match omp || fail "omp busy regex must match the status-set spinner frames, not only the activity set" + printf ' 󰵗 · gpt-6-astra · 36.7%%/41K\n' | fm_busy_lines_match omp && fail "an idle omp status row must not read busy" + printf 'esc to interrupt\n' | fm_busy_lines_match omp && fail "omp must not borrow Claude's footer" + local bin out + bin=$(make_named_shells "$TMP_ROOT/named-model") + # shellcheck disable=SC2016 # the quoted body expands inside the named shell + out=$(env -u CLAUDECODE -u FM_OMP_HARNESS -u PI_CODING_AGENT -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u FM_SUPERVISION_MODEL \ + "$bin/omp" -c '. "$1"; fm_supervision_model' _ "$ROOT/bin/fm-wake-lib.sh") + [ "$out" = extension ] || fail "an omp primary must run the extension supervision model, got '$out'" + pass "control, composer, and supervision-model tables carry omp's verified values" +} + +# --- 5. Ownership proof -------------------------------------------------------- + +# Stand up the durable evidence a live omp session leaves behind: both tracked +# extensions under the case root and one marker per extension recording that +# build plus the session pid in state/.lock. +record_omp_session() { # <root> <home> <session-pid> [omit] [drift] + local root=$1 home=$2 session_pid=$3 omit=${4:-} drift=${5:-} pair source marker version + mkdir -p "$root/.omp/extensions" "$home/state" + for pair in \ + "fm-primary-omp-watch.ts:.omp-watch-extension-loaded:watch" \ + "fm-primary-turnend-guard.ts:.omp-turnend-extension-loaded:turnend"; do + source=${pair%%:*} + marker=${pair#*:}; marker=${marker%%:*} + printf '// %s\n' "${pair##*:}" > "$root/.omp/extensions/$source" + [ "$omit" = "${pair##*:}" ] && continue + if [ "$drift" = "${pair##*:}" ]; then + version="sha256:0000000000000000000000000000000000000000000000000000000000000000" + else + version=$(bash -c '. "$1"; fm_pi_extension_version "$2"' _ "$ROOT/bin/fm-wake-lib.sh" "$root/.omp/extensions/$source") || return 1 + fi + printf '%s\n%s\n' "$version" "$session_pid" > "$home/state/$marker" + done + printf '%s\n' "$session_pid" > "$home/state/.lock" +} + +owns() { # <root> <home> + bash -c '. "$1"; fm_omp_extension_owns_supervision "$2" "$3"' _ "$ROOT/bin/fm-wake-lib.sh" "$2/state" "$1" +} + +test_ownership_proof_is_omp_keyed() { + local root home pid + sleep 60 & + pid=$! + root="$TMP_ROOT/own/root"; home="$TMP_ROOT/own/home" + record_omp_session "$root" "$home" "$pid" || fail "could not record the omp session" + owns "$root" "$home" || fail "a live session that loaded both omp extensions must own supervision" + bash -c '. "$1"; fm_pi_extension_owns_supervision "$2" "$3"' _ "$ROOT/bin/fm-wake-lib.sh" "$home/state" "$root" \ + && fail "omp markers must never satisfy the Pi proof" + bash -c '. "$1"; fm_extension_owns_supervision "$2" "$3"' _ "$ROOT/bin/fm-wake-lib.sh" "$home/state" "$root" \ + || fail "the shared extension proof must accept the omp pair" + + root="$TMP_ROOT/own-drift/root"; home="$TMP_ROOT/own-drift/home" + record_omp_session "$root" "$home" "$pid" "" watch || fail "could not record the drifted session" + owns "$root" "$home" && fail "a session that loaded an older watch build must not own supervision" + root="$TMP_ROOT/own-omit/root"; home="$TMP_ROOT/own-omit/home" + record_omp_session "$root" "$home" "$pid" turnend || fail "could not record the partial session" + owns "$root" "$home" && fail "a session missing the turn-end guard extension must not own supervision" + root="$TMP_ROOT/own-dead/root"; home="$TMP_ROOT/own-dead/home" + record_omp_session "$root" "$home" "$pid" || fail "could not record the dead session" + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + owns "$root" "$home" && fail "a dead session must not own supervision" + + # The pull-guard verdict tolerates the extension's own hand-off only with the proof. + sleep 60 & + pid=$! + root="$TMP_ROOT/own-verdict/root"; home="$TMP_ROOT/own-verdict/home" + record_omp_session "$root" "$home" "$pid" || fail "could not record the verdict session" + touch "$home/state/.last-watcher-beat" + local verdict + verdict=$(FM_SUPERVISION_MODEL=extension FM_HOME="$home" bash -c ' + . "$1"; fm_watcher_supervision_verdict "$2" "$3" 999 "$4" "$5"; printf "%s %s" "$FM_WATCHER_VERDICT_OK" "$FM_WATCHER_VERDICT_REASON"' \ + _ "$ROOT/bin/fm-wake-lib.sh" "$home/state" "$root/bin/fm-watch.sh" "$home" "$root") + [ "${verdict%% *}" = true ] || fail "an unheld lock with a fresh beacon and the omp proof must be healthy, got '$verdict'" + rm -f "$home/state/.omp-turnend-extension-loaded" + verdict=$(FM_SUPERVISION_MODEL=extension FM_HOME="$home" bash -c ' + . "$1"; fm_watcher_supervision_verdict "$2" "$3" 999 "$4" "$5"; printf "%s %s" "$FM_WATCHER_VERDICT_OK" "$FM_WATCHER_VERDICT_REASON"' \ + _ "$ROOT/bin/fm-wake-lib.sh" "$home/state" "$root/bin/fm-watch.sh" "$home" "$root") + [ "$verdict" = "false no-watcher" ] || fail "without the proof the same hand-off must alarm as no-watcher, got '$verdict'" + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + pass "fm-wake-lib: the omp ownership proof is keyed on its own extensions and gates the hand-off tolerance" +} + +# --- 6. The tracked primary extensions over a fake omp API ---------------------- + +install_omp_extension_fixture() { # <repo> + local repo=$1 + mkdir -p "$repo/.omp/extensions" "$repo/.pi/extensions/lib" "$repo/bin" "$repo/node_modules/typebox" + cp "$ROOT/.omp/extensions/fm-primary-turnend-guard.ts" "$ROOT/.omp/extensions/fm-primary-omp-watch.ts" "$repo/.omp/extensions/" + cp "$ROOT/.pi/extensions/lib/fm-operational-input.ts" "$ROOT/.pi/extensions/lib/fm-sessionstart-supervisor.mjs" "$repo/.pi/extensions/lib/" + cp "$ROOT/bin/fm-operational-input.sh" "$repo/bin/" + chmod +x "$repo/bin/fm-operational-input.sh" + printf '{"name":"typebox","type":"module","exports":"./index.js"}\n' > "$repo/node_modules/typebox/package.json" + printf 'export const Type = { Object(p) { return { type: "object", properties: p }; } };\n' > "$repo/node_modules/typebox/index.js" +} + +test_turnend_guard_extension_compels_one_continuation() { + local repo home out status + repo="$TMP_ROOT/guard/repo"; home="$TMP_ROOT/guard/home" + install_omp_extension_fixture "$repo" + mkdir -p "$home/state" + cat > "$repo/bin/fm-turnend-guard.sh" <<'SH' +#!/usr/bin/env bash +payload=$(cat); printf '%s\n' "$payload" >> "${FM_GUARD_LOG:?}" +case "$payload" in *'"stop_hook_active":true'*) exit 0 ;; esac +printf 'guard says: repair with fm_watch_arm_omp\n' >&2; exit 2 +SH + cat > "$repo/bin/fm-arm-pretool-check.sh" <<'SH' +#!/usr/bin/env bash +case "$*" in *fm-watch-arm.sh*'&'*) printf 'fm watcher-arm seatbelt: blocked\n' >&2; exit 2 ;; esac; exit 0 +SH + printf '#!/usr/bin/env bash\nexit 0\n' > "$repo/bin/fm-cd-pretool-check.sh" + # shellcheck disable=SC2016 # $2 expands in the generated script + printf '#!/usr/bin/env bash\nprintf "OMP DIGEST source=%%s\\n" "$2"\n' > "$repo/bin/fm-sessionstart-run.sh" + chmod +x "$repo/bin/"*.sh + out=$(FM_GUARD_LOG="$TMP_ROOT/guard/guard.log" FM_HOME="$home" EXT="$repo/.omp/extensions/fm-primary-turnend-guard.ts" node --input-type=module 2>&1 <<'EOF' +import { pathToFileURL } from "node:url"; +import { readFileSync, existsSync } from "node:fs"; +const handlers = new Map(); +const pi = { on(e, h) { handlers.set(e, h); }, sendMessage() {} }; +const mod = await import(pathToFileURL(process.env.EXT).href); +mod.default(pi); +for (const name of ["session_start", "before_agent_start", "session_compact", "session_shutdown", "tool_call", "session_stop"]) { + if (!handlers.has(name)) throw new Error(`${name} handler was not registered`); +} +if (handlers.has("agent_settled")) throw new Error("omp guard must not listen for agent_settled"); +const ctx = { sessionManager: { getSessionId: () => "s1" } }; +handlers.get("session_start")({ type: "session_start" }, ctx); +const first = await handlers.get("before_agent_start")({ type: "before_agent_start", prompt: "hi" }, ctx); +if (!first?.message?.content?.includes("FIRSTMATE_OP: v1 session-start: OMP DIGEST source=startup")) throw new Error(`first start did not deliver a startup digest: ${JSON.stringify(first)}`); +if (first.message.display !== false || first.message.customType !== "firstmate-sessionstart-nudge") throw new Error("digest message lost its persistent shape"); +// A later in-process session_start is a replacement and maps to clear. +handlers.get("session_start")({ type: "session_start" }, ctx); +const second = await handlers.get("before_agent_start")({ type: "before_agent_start", prompt: "hi" }, ctx); +if (!second?.message?.content?.includes("source=clear")) throw new Error(`in-process replacement did not map to clear: ${JSON.stringify(second)}`); +const allowed = await handlers.get("tool_call")({ type: "tool_call", toolName: "bash", input: { command: "ls" } }, {}); +if (allowed.block) throw new Error("an ordinary command was blocked"); +const blocked = await handlers.get("tool_call")({ type: "tool_call", toolName: "bash", input: { command: "bin/fm-watch-arm.sh &" } }, {}); +if (blocked.block !== true || !blocked.reason.includes("seatbelt")) throw new Error(`backgrounded arm was not blocked: ${JSON.stringify(blocked)}`); +const r1 = await handlers.get("session_stop")({ type: "session_stop", stop_hook_active: false }, {}); +if (r1?.continue !== true) throw new Error(`guard exit 2 did not compel a continuation: ${JSON.stringify(r1)}`); +if (!r1.additionalContext.startsWith("⁣FIRSTMATE_OP: v1 turn-end-guard: ")) throw new Error(`continuation context is not typed operational input: ${r1.additionalContext}`); +if (!r1.additionalContext.includes("TURN WOULD END BLIND") || !r1.additionalContext.includes("repair with fm_watch_arm_omp")) throw new Error("continuation dropped the guard text"); +const r2 = await handlers.get("session_stop")({ type: "session_stop", stop_hook_active: true }, {}); +if (r2 !== undefined) throw new Error(`the flagged second stop must stand down, got ${JSON.stringify(r2)}`); +const payloads = readFileSync(process.env.FM_GUARD_LOG, "utf8").trim().split("\n"); +if (payloads.join("|") !== '{"stop_hook_active":false}|{"stop_hook_active":true}') throw new Error(`guard payloads were ${payloads.join("|")}`); +if (!existsSync(`${process.env.FM_HOME}/state/.omp-turnend-extension-loaded`)) throw new Error("loaded marker was not written"); +await handlers.get("session_shutdown")({}, {}); +EOF +) + status=$? + expect_code 0 "$status" "omp turn-end guard extension contract: $out" + [ -z "$out" ] || fail "omp guard extension test printed output: $out" + pass ".omp turn-end guard: digest delivery, seatbelt block, one compelled continuation, flagged stop stands down" +} + +test_watch_extension_arms_and_delivers() { + local repo home out status + repo="$TMP_ROOT/watch/repo"; home="$TMP_ROOT/watch/home" + install_omp_extension_fixture "$repo" + mkdir -p "$home/state" + # The first arm child closes with one actionable reason; every successor + # stays up, so exactly one wake exists to consume. + cat > "$repo/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +printf 'watcher: started pid=%s (beacon 0s) recovery-generation=gen-1\n' "$$" +if [ ! -e "${FM_HOME:?}/state/.e2e-fired" ]; then + : > "$FM_HOME/state/.e2e-fired" + sleep 1 + printf 'signal: omp-e2e done\n' + exit 0 +fi +sleep 30 +SH + chmod +x "$repo/bin/fm-watch-arm.sh" + out=$(FM_HOME="$home" FM_ROOT_OVERRIDE="$repo" FM_OMP_ARM_READY_TIMEOUT_MS=3000 FM_WATCH_REARM_RETRY_LIMIT=1 FM_WATCH_REARM_RETRY_BASE_MS=5 FM_WATCH_REARM_RETRY_MAX_MS=10 \ + EXT="$repo/.omp/extensions/fm-primary-omp-watch.ts" node --input-type=module 2>&1 <<'EOF' +import { pathToFileURL } from "node:url"; +import { writeFileSync, existsSync, readFileSync } from "node:fs"; +writeFileSync(`${process.env.FM_HOME}/state/.lock`, `${process.pid}\n`); +const handlers = new Map(); let tool = null; let command = null; const sent = []; +const pi = { + on(e, h) { handlers.set(e, h); }, + registerCommand(n, o) { if (n === "fm-watch-arm-omp") command = o.handler; }, + registerTool(t) { tool = t; }, + // omp sendUserMessage returns synchronously, not a promise. + sendUserMessage(m, o) { sent.push({ m, o }); return undefined; }, +}; +const mod = await import(pathToFileURL(process.env.EXT).href); +mod.default(pi); +if (!tool || tool.name !== "fm_watch_arm_omp") throw new Error("fm_watch_arm_omp was not registered"); +if (!command) throw new Error("/fm-watch-arm-omp was not registered"); +if (tool.parameters?.type !== "object") throw new Error("tool parameters must be an empty object schema"); +const result = await tool.execute(); +if (!/^watcher: started omp extension arm child 1;/.test(result.content[0].text)) throw new Error(`unexpected arm result: ${result.content[0].text}`); +const marker = readFileSync(`${process.env.FM_HOME}/state/.omp-watch-extension-loaded`, "utf8").split("\n"); +if (marker[1] !== String(process.pid)) throw new Error("loaded marker must record the session pid"); +const again = await tool.execute(); +if (!/^watcher: unchanged - omp extension already owns an arm child/.test(again.content[0].text)) throw new Error(`redundant arm was not an ownership no-op: ${again.content[0].text}`); +await new Promise((r) => setTimeout(r, 2500)); +if (sent.length !== 1) throw new Error(`expected one follow-up wake, saw ${sent.length}: ${JSON.stringify(sent)}`); +if (!sent[0].m.startsWith("⁣FIRSTMATE_OP: v1 watcher: FIRSTMATE WATCHER WAKE: signal: omp-e2e done")) throw new Error(`unexpected wake text: ${sent[0].m}`); +if (sent[0].o?.deliverAs !== "followUp") throw new Error("wake must be delivered as a follow-up"); +// The wake is consumed when omp starts the next run with that exact prompt. +await handlers.get("before_agent_start")({ type: "before_agent_start", prompt: sent[0].m }, {}); +await handlers.get("session_shutdown")({}, {}); +if (existsSync(`${process.env.FM_HOME}/state/extensions/omp-primary-watch/session-replacement-actionable.json`)) throw new Error("a consumed wake must not ride the replacement handoff"); +process.exit(0); +EOF +) + status=$? + expect_code 0 "$status" "omp watch extension contract: $out" + [ -z "$out" ] || fail "omp watch extension test printed output: $out" + pass ".omp watch extension: fm_watch_arm_omp arms once, repeats as a no-op, and delivers an actionable close as one follow-up" +} + +test_detection_anchored_name_and_marker_precedence +test_lock_identity_and_liveness_classification +test_spawn_launch_line_and_worker_wiring +test_spawn_model_validation_scoped_to_listed_providers +test_secondmate_launch_relies_on_discovery +test_secondmate_config_pinned_model_is_validated +test_busy_extension_lifecycle +test_control_composer_and_model_tables +test_ownership_proof_is_omp_keyed +test_turnend_guard_extension_compels_one_continuation +test_watch_extension_arms_and_delivers diff --git a/tests/fm-omp-primary-live-e2e.test.sh b/tests/fm-omp-primary-live-e2e.test.sh new file mode 100755 index 00000000000..333c9b27c37 --- /dev/null +++ b/tests/fm-omp-primary-live-e2e.test.sh @@ -0,0 +1,313 @@ +#!/usr/bin/env bash +# Opt-in credentialed omp (Oh My Pi) primary regression in an isolated lab +# checkout. It drives a real omp through its JSON-RPC stdio mode so no terminal +# multiplexer is needed, uses the captain's existing omp login without copying +# any credential, and defaults to the captain-approved openai-codex model. +# +# It proves, against the installed omp, everything the portable suite +# (tests/fm-omp-harness.test.sh) can only pin over a fake API: +# 1. both tracked .omp/extensions load by auto-discovery alone; +# 2. the session-start digest reaches model context before the first turn +# and the session lock names the omp process (ancestry detection); +# 3. fm_watch_arm_omp starts a real watcher, an actionable close spawns a +# ledger-linked successor, and the wake arrives as one follow-up turn; +# 4. with the successor watcher frozen until its beacon passes the lab grace, +# the next turn end is genuinely unsupervised, so session_stop must compel +# the turn-end guard continuation and the model reaches for the tool. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_OMP_LIVE_E2E omp node jq + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +unset NO_MISTAKES_GATE + +fail() { + printf 'not ok - %s\n' "$1" >&2 + if [ -f "${RPC_LOG:-}" ]; then + printf '# rpc frame types seen:\n' >&2 + grep -o '"type":"[a-z_]*"' "$RPC_LOG" 2>/dev/null | sort | uniq -c | sort -rn | head -30 >&2 + printf '# guard spy log:\n' >&2 + tail -12 "${GUARD_SPY_LOG:-/dev/null}" >&2 2>/dev/null + printf '# last stderr lines:\n' >&2 + tail -5 "${RPC_ERR:-/dev/null}" >&2 + if [ "${FM_OMP_LIVE_KEEP:-0}" = 1 ]; then + printf '# lab kept at %s\n' "$LAB" >&2 + trap - EXIT + exec 3>&- 2>/dev/null || true + [ -z "$OMP_PID" ] || kill -TERM "$OMP_PID" 2>/dev/null || true + fi + fi + exit 1 +} +pass() { printf 'ok - %s\n' "$1"; } +note() { printf '# %s\n' "$1"; } + +OMP_VERSION=$(omp --version 2>/dev/null | head -1) +MODEL=${FM_OMP_LIVE_MODEL:-openai-codex/gpt-6-astra} +# The guard's beacon grace for this lab. The watcher beats every FM_POLL=1s, so +# a 20s grace is comfortably healthy in normal operation and lets stage 3 make +# the beacon stale by freezing the watcher for a bounded time instead of killing +# it: a killed watcher closes its arm child and the extension re-arms within +# milliseconds, which would keep the guard from ever firing. +GUARD_GRACE=20 +LAB="$ROOT/.omp-live-e2e.$$" +PROJECT="$LAB/project" +RPC_IN="$LAB/rpc.in" +RPC_LOG="$LAB/rpc.log" +RPC_ERR="$LAB/rpc.err" +OMP_PID= + +lab_pid_is_safe() { + local pid=$1 command + command=$(ps -p "$pid" -o command= 2>/dev/null || true) + case "$command" in + *"$LAB"*) return 0 ;; + *) return 1 ;; + esac +} + +# Every process the lab started names the lab path on its command line (omp +# itself, the session-start supervisor and its runner, the watcher and its arm +# child), so cleanup reaps by that path rather than by remembered pids: an omp +# rpc process that outlives its closed stdin, or a detached session-start +# worker, would otherwise survive the lab that created it. +lab_pids() { + ps -axo pid=,command= | awk -v lab="$LAB" 'index($0, lab) { print $1 }' +} + +reap_lab() { + local pid + for pid in $(lab_pids); do kill -TERM "$pid" 2>/dev/null || true; done + sleep 1 + for pid in $(lab_pids); do kill -KILL "$pid" 2>/dev/null || true; done +} + +cleanup() { + exec 3>&- 2>/dev/null || true + if [ -n "$OMP_PID" ]; then + kill -TERM "$OMP_PID" 2>/dev/null || true + fi + reap_lab + rm -rf "$LAB" +} +trap cleanup EXIT + +# --- lab checkout: the tracked tree plus this working tree's pending edits ---- +mkdir -p "$LAB" +git clone -q "$ROOT" "$PROJECT" || fail "could not clone the repository into the lab" +while IFS= read -r path; do + [ -n "$path" ] || continue + [ -f "$ROOT/$path" ] || continue + mkdir -p "$PROJECT/$(dirname "$path")" + cp "$ROOT/$path" "$PROJECT/$path" +done <<EOF +$(git -C "$ROOT" ls-files --modified --others --exclude-standard) +EOF +mkdir -p "$PROJECT/state" "$PROJECT/config" "$PROJECT/data" +# A spy in front of the real turn-end guard: every invocation records the +# payload the extension sent and the exit code the real guard returned, which +# proves the compelled continuation (a payload with stop_hook_active true can +# only come from a stop omp raised for the continuation itself) independently of +# whether the rpc stream echoes additionalContext. +GUARD_SPY_LOG="$LAB/guard-spy.log" +mv "$PROJECT/bin/fm-turnend-guard.sh" "$PROJECT/bin/fm-turnend-guard.real.sh" +cat > "$PROJECT/bin/fm-turnend-guard.sh" <<SH +#!/usr/bin/env bash +payload=\$(cat) +printf '%s' "\$payload" | "\$(dirname "\$0")/fm-turnend-guard.real.sh" "\$@" +rc=\$? +printf 'rc=%s payload=%s\n' "\$rc" "\$payload" >> '$GUARD_SPY_LOG' +exit "\$rc" +SH +chmod +x "$PROJECT/bin/fm-turnend-guard.sh" +[ -f "$PROJECT/.omp/extensions/fm-primary-omp-watch.ts" ] || fail "lab checkout is missing the omp watch extension" +[ -f "$PROJECT/.omp/extensions/fm-primary-turnend-guard.ts" ] || fail "lab checkout is missing the omp turn-end extension" + +# --- rpc plumbing -------------------------------------------------------------- +rpc_send() { # <json-line> + printf '%s\n' "$1" >&3 +} + +wait_for_log() { # <fixed-string> <attempts> + local expected=$1 attempts=${2:-240} i=0 + while [ "$i" -lt "$attempts" ]; do + grep -Fq -- "$expected" "$RPC_LOG" 2>/dev/null && return 0 + sleep 0.5 + i=$((i + 1)) + done + return 1 +} + +wait_for_file() { # <path> <attempts> + local path=$1 attempts=${2:-240} i=0 + while [ "$i" -lt "$attempts" ]; do + [ -f "$path" ] && return 0 + sleep 0.5 + i=$((i + 1)) + done + return 1 +} + +# Model-issued invocations of one extension tool, counted from the +# tool-execution frames rather than from text, because a tool result is echoed +# by several frame kinds. omp exposes extension tools to some models (verified: +# the openai-codex family on 18.1.11) through its virtual-file bridge, where the +# model invokes the tool by WRITING xd://<tool-name>; a direct call and a bridge +# write are the same invocation and are counted together. +tool_call_count() { # <tool-name> + local n + n=$(jq -r --arg t "$1" 'select(.type == "tool_execution_start" and (.toolName == $t or (.toolName == "write" and (.args.path // "") == ("xd://" + $t)))) | .type' "$RPC_LOG" 2>/dev/null | grep -c . 2>/dev/null) || true + printf '%s' "${n:-0}" +} + +# Every assistant text delta from the rpc event stream, joined, since <line>. +assistant_text_since() { # <line-number> + tail -n +"$1" "$RPC_LOG" | jq -r 'select(.type == "message_update") | .assistantMessageEvent | select(.type == "text_delta") | .delta' 2>/dev/null | tr -d '\n' +} + +agent_end_count() { + local n + n=$(jq -r 'select(.type == "agent_end") | .type' "$RPC_LOG" 2>/dev/null | grep -c . 2>/dev/null) || true + printf '%s' "${n:-0}" +} + +wait_for_agent_ends() { # <count> <attempts> + local want=$1 attempts=${2:-360} i=0 + while [ "$i" -lt "$attempts" ]; do + [ "$(agent_end_count)" -ge "$want" ] && return 0 + sleep 0.5 + i=$((i + 1)) + done + return 1 +} + +mkfifo "$RPC_IN" || fail "could not create the rpc fifo" +: > "$RPC_LOG" +( + cd "$PROJECT" && + env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + -u FM_HOME -u FM_ROOT_OVERRIDE -u FM_STATE_OVERRIDE -u FM_CONFIG_OVERRIDE -u FM_DATA_OVERRIDE \ + FM_OMP_HARNESS=omp OMP_SKIP_SETUP=1 FM_POLL=1 FM_SIGNAL_GRACE=0 FM_HEARTBEAT=600 \ + FM_GUARD_GRACE="$GUARD_GRACE" \ + omp --mode rpc --no-session --cwd "$PROJECT" --config "$PROJECT/.omp/fm-worker-overlay.yml" --auto-approve \ + --model "$MODEL" --thinking low < "$RPC_IN" > "$RPC_LOG" 2> "$RPC_ERR" +) & +OMP_PID=$! +exec 3> "$RPC_IN" + +wait_for_log '"type":"ready"' 240 || fail "omp $OMP_VERSION did not print its rpc ready frame: $(tail -5 "$RPC_ERR")" +wait_for_file "$PROJECT/state/.omp-turnend-extension-loaded" 60 || fail "omp $OMP_VERSION did not auto-discover the turn-end guard extension" +wait_for_file "$PROJECT/state/.omp-watch-extension-loaded" 60 || fail "omp $OMP_VERSION did not auto-discover the watch extension" +pass "omp $OMP_VERSION: both tracked .omp/extensions loaded by auto-discovery with no -e and no trust dialog" + +# --- 1. session-start digest and lock identity --------------------------------- +rpc_send '{"id":"p1","type":"prompt","message":"From the Firstmate session-start digest already in your context, reply with the single line that begins with SESSION START - and nothing else. Do not run any tool."}' +wait_for_agent_ends 1 360 || fail "omp did not finish the first turn: $(tail -3 "$RPC_ERR")" +first=$(assistant_text_since 1) +case "$first" in + *"SESSION START - $PROJECT"*) ;; + *) fail "the session-start digest did not reach model context before the first turn; reply was: $first" ;; +esac +lock_pid=$(sed -n '1p' "$PROJECT/state/.lock" 2>/dev/null || true) +omp_real_pid=$(pgrep -P "$OMP_PID" -x omp 2>/dev/null | head -1 || true) +[ -n "$omp_real_pid" ] || omp_real_pid=$OMP_PID +[ "$lock_pid" = "$omp_real_pid" ] || fail "the session lock names pid '$lock_pid', not the omp process $omp_real_pid; ancestry detection failed" +[ -f "$PROJECT/state/.session-start-complete" ] || fail "session start did not record completion" +pass "omp $OMP_VERSION: before_agent_start delivered the digest into model context and the lock names the omp process" + +# --- 2. watcher arm, successor, and wake delivery ------------------------------ +: > "$PROJECT/state/omp-e2e.meta" +rpc_send '{"id":"p2","type":"prompt","message":"Call the fm_watch_arm_omp tool exactly once now, then reply with its result text verbatim and nothing else. Never run bin/fm-watch-arm.sh through bash."}' +wait_for_log "watcher: started omp extension arm child 1" 360 || fail "omp did not render the initial watcher tool result: $(tail -3 "$RPC_ERR")" +wait_for_agent_ends 2 360 || fail "omp did not finish the arm turn" +watcher_pid=$(cat "$PROJECT/state/.watch.lock/pid" 2>/dev/null || true) +if [ -z "$watcher_pid" ] || ! kill -0 "$watcher_pid" 2>/dev/null; then + fail "no live watcher holds the lab home lock after fm_watch_arm_omp" +fi +pass "omp $OMP_VERSION: fm_watch_arm_omp started a live watcher through the extension" + +printf 'done: omp live e2e watcher fire\n' > "$PROJECT/state/omp-e2e.status" +i=0 +while [ "$i" -lt 240 ]; do + grep -Eq 'reason=actionable-signal.*successor=started:[0-9]+' "$PROJECT/state/.watch-cycle-exits.log" 2>/dev/null && break + sleep 0.5 + i=$((i + 1)) +done +grep -Eq 'reason=actionable-signal.*successor=started:[0-9]+' "$PROJECT/state/.watch-cycle-exits.log" 2>/dev/null \ + || fail "omp extension did not start and ledger-link a successor after the actionable close" +wait_for_log "FIRSTMATE WATCHER WAKE: signal:" 240 || fail "the actionable close was not delivered to main as a watcher follow-up" +wait_for_agent_ends 3 360 || fail "omp did not finish the wake turn" +arm_calls=$(tool_call_count fm_watch_arm_omp) +[ "$arm_calls" -eq 1 ] || fail "the model re-armed from memory instead of the extension (fm_watch_arm_omp call count $arm_calls)" +pass "omp $OMP_VERSION: an actionable close spawned a ledger-linked successor and woke main exactly once" + +# --- 3. the compelled turn-end guard continuation ------------------------------- +# Freeze the successor watcher (SIGSTOP) so its beacon goes stale past the lab +# grace while its arm child stays attached: the extension sees no close and +# schedules no retry, so the next turn end is genuinely unsupervised and +# session_stop must compel the guard continuation. The model's repair call then +# returns the extension's ownership no-op (it still owns the frozen arm), and +# thawing the watcher restores the same cycle. +successor_pid=$(cat "$PROJECT/state/.watch.lock/pid" 2>/dev/null || true) +[ -n "$successor_pid" ] || fail "no successor watcher recorded before the guard probe" +lab_pid_is_safe "$successor_pid" || fail "refusing to freeze a watcher outside the lab ($successor_pid)" +kill -STOP "$successor_pid" 2>/dev/null || fail "could not freeze the successor watcher" +thaw() { kill -CONT "$successor_pid" 2>/dev/null || true; } +i=0 +while [ "$i" -lt 60 ]; do + age=$(( $(date +%s) - $(stat -f %m "$PROJECT/state/.last-watcher-beat" 2>/dev/null || stat -c %Y "$PROJECT/state/.last-watcher-beat" 2>/dev/null || date +%s) )) + [ "$age" -gt "$GUARD_GRACE" ] && break + sleep 1 + i=$((i + 1)) +done +[ "$age" -gt "$GUARD_GRACE" ] || { thaw; fail "the frozen watcher's beacon never went stale (age ${age}s)"; } +rpc_send '{"id":"p3","type":"prompt","message":"Reply with exactly GUARD_PROBE and nothing else. Do not call any tool unless a later instruction in this turn tells you supervision is off."}' +# The guard must have refused a stop (rc=2) and omp must then have raised the +# continuation's own stop with stop_hook_active true. +i=0 +while [ "$i" -lt 360 ]; do + grep -q '^rc=2 ' "$GUARD_SPY_LOG" 2>/dev/null && grep -q 'stop_hook_active":true' "$GUARD_SPY_LOG" 2>/dev/null && break + sleep 0.5 + i=$((i + 1)) +done +grep -q '^rc=2 ' "$GUARD_SPY_LOG" 2>/dev/null || { thaw; fail "the turn-end guard never refused a stop while the watcher was frozen (spy log: $(cat "$GUARD_SPY_LOG" 2>/dev/null))"; } +grep -q 'stop_hook_active":true' "$GUARD_SPY_LOG" 2>/dev/null || { thaw; fail "omp did not raise the compelled continuation's own stop (spy log: $(cat "$GUARD_SPY_LOG" 2>/dev/null))"; } +i=0 +while [ "$i" -lt 360 ]; do + [ "$(tool_call_count fm_watch_arm_omp)" -ge 2 ] && break + sleep 0.5 + i=$((i + 1)) +done +[ "$(tool_call_count fm_watch_arm_omp)" -ge 2 ] \ + || { thaw; fail "the model did not reach for fm_watch_arm_omp after the compelled continuation"; } +wait_for_agent_ends 4 360 || { thaw; fail "omp did not settle after the compelled continuation"; } +thaw +sleep 3 +repaired_pid=$(cat "$PROJECT/state/.watch.lock/pid" 2>/dev/null || true) +if [ -z "$repaired_pid" ] || ! kill -0 "$repaired_pid" 2>/dev/null; then + fail "no live watcher after the guard stage" +fi +pass "omp $OMP_VERSION: session_stop compelled the guard continuation (guard rc=2, then a stop_hook_active stop) and the model reached for fm_watch_arm_omp" + +# --- shutdown ------------------------------------------------------------------- +# omp documents that closing rpc stdin disposes the session and exits 0. On +# 18.1.11 the process outlived its closed stdin for longer than 30s in this lab +# while its session-start supervisor child was still attached, so the exit is +# recorded as a note rather than asserted: it is omp's shutdown behavior, not +# Firstmate's supervision contract, and cleanup reaps the lab either way. +exec 3>&- +i=0 +while [ "$i" -lt 60 ]; do + kill -0 "$OMP_PID" 2>/dev/null || break + sleep 0.5 + i=$((i + 1)) +done +if kill -0 "$OMP_PID" 2>/dev/null; then + note "omp $OMP_VERSION did not exit within 30s of its rpc stdin closing; terminating the lab session" +else + note "omp $OMP_VERSION exited on its own after its rpc stdin closed" +fi +note "omp $OMP_VERSION model=$MODEL: every live omp primary assertion passed" diff --git a/tests/fm-opencode-primary-live-e2e.test.sh b/tests/fm-opencode-primary-live-e2e.test.sh index edbade79eeb..8b96cae6ff6 100755 --- a/tests/fm-opencode-primary-live-e2e.test.sh +++ b/tests/fm-opencode-primary-live-e2e.test.sh @@ -3,10 +3,10 @@ # FM_HOME. Existing OpenCode credentials stay in their managed store. set -u -if [ "${FM_OPENCODE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_OPENCODE_LIVE_E2E=1 to run the interactive OpenCode continuity regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_OPENCODE_LIVE_E2E opencode tmux sqlite3 ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -16,10 +16,6 @@ fail() { exit 1 } -command -v opencode >/dev/null 2>&1 || fail "opencode not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" -command -v sqlite3 >/dev/null 2>&1 || fail "sqlite3 not found" - TMUX=$(command -v tmux) SOCKET="fm-opencode-live-e2e-$$" SESSION=opencode-live-e2e diff --git a/tests/fm-pi-branch-extension.test.sh b/tests/fm-pi-branch-extension.test.sh index b38834679be..c39f2604463 100755 --- a/tests/fm-pi-branch-extension.test.sh +++ b/tests/fm-pi-branch-extension.test.sh @@ -107,6 +107,10 @@ export class ModelRuntime { constructor() { this.models = (globalThis.__fmBranchStaticModels?.() ?? []).map((model) => ({ ...model })); this.authenticated = new Set(this.models.filter((model) => model.storedAuth !== false).map((model) => model.provider)); + this.registeredProviderConfigs = new Map(); + // Like the real runtime, a registered provider's credentials are only + // known once refresh() has run for it; registration alone is provisional. + this.pendingAuth = new Set(); } static async create() { const queuedError = globalThis.__fmModelRuntimeErrors?.shift(); @@ -116,6 +120,18 @@ export class ModelRuntime { (globalThis.__fmModelRuntimes ??= []).push(runtime); return runtime; } + registerProvider(providerId, config) { + this.registeredProviderConfigs.set(providerId, config); + for (const model of config.models ?? []) { + this.models.push({ ...model, provider: providerId }); + } + if (config.oauth || config.apiKey) this.pendingAuth.add(providerId); + } + async refresh(options) { + for (const providerId of options?.providers ?? this.pendingAuth) { + if (this.pendingAuth.delete(providerId)) this.authenticated.add(providerId); + } + } getModel(provider, id) { return this.models.find((model) => model.provider === provider && model.id === id); } @@ -447,6 +463,8 @@ const modelRegistry = { getAvailable: () => registryModels.filter((model) => model.mainAvailable !== false).slice(), find: (provider, id) => registryModels.find((model) => model.provider === provider && model.id === id), hasConfiguredAuth: (model) => model.mainAvailable !== false, + getRegisteredProviderConfig: (providerId) => globalThis.__fmExtensionProviderConfigs?.get(providerId), + getRegisteredProviderIds: () => [...(globalThis.__fmExtensionProviderConfigs?.keys() ?? [])], }; function makeCtx(extra) { return { @@ -1821,6 +1839,71 @@ EOF pass "pre-drain eligibility re-check excludes a newly main-owned row without deferring eligible work" } +# A needs-decision signal wakes main independently, but it must not veto an +# already accepted routine delivery at the branch's pre-drain recheck. The +# grant serializes the actors: branch owns only the routine row, while the +# decision row remains main-owned. If the prompted branch then fails, rejecting +# the settlement releases the grant so watcher fallback can replay both rows. +test_branch_predrain_needs_decision_keeps_routine_row_branch_eligible() { + local repo home out status + repo="$TMP_ROOT/predrain-needs-decision-root" + home="$TMP_ROOT/predrain-needs-decision-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { bus, fire, home, makeOffer, realRoot }; })()`); +const { bus, fire, home, makeOffer, realRoot } = globalThis.__t; +import { spawnSync } from "node:child_process"; +import { existsSync, readFileSync, writeFileSync } from "node:fs"; + +fire("session_start", {}); +writeFileSync( + `${home}/state/.wake-queue`, + "1\t1\tsignal\tbranch-driver.status\tsignal: routine progress\n" + + "2\t2\tsignal\tdecision-task.status\tneeds-decision: decision-task.status\n", +); +let releasePrompt; +globalThis.__fmPromptGate = new Promise((resolve) => { releasePrompt = resolve; }); +const offer = makeOffer("signal: branch-driver.status"); +bus.emit("fm-branch-supervision:dispatch", offer); +if (!offer.accepted) throw new Error("branch refused the routine offer before its mixed-queue recheck"); +for (let i = 0; i < 250 && !globalThis.__fmPromptStarted; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +if (!globalThis.__fmPromptStarted) { + throw new Error("a co-present needs-decision row vetoed the accepted routine branch prompt"); +} +const snapshot = readFileSync(`${home}/state/.branch-eligible-rows`, "utf8").trim().split("\n"); +if (!snapshot.includes("1") || snapshot.includes("2")) { + throw new Error(`mixed queue granted the wrong rows to branch: ${snapshot}`); +} +releasePrompt(); +const failure = await offer.settlement.then(() => null, (error) => error); +if (!(failure instanceof Error) || !failure.message.includes("produced no durable outcome")) { + throw new Error(`accepted wake settled without delivery instead of rejecting to fallback: ${String(failure)}`); +} +if (existsSync(`${home}/state/.branch-eligible-rows`)) { + throw new Error("failed branch prompt retained its routine-row grant"); +} +const drain = spawnSync("bash", [`${realRoot}/bin/fm-wake-drain.sh`], { + encoding: "utf8", + env: { ...process.env, FM_HOME: home, FM_STATE_OVERRIDE: `${home}/state`, FM_ROOT_OVERRIDE: realRoot }, +}); +if (drain.status !== 0) throw new Error(`main fallback drain failed: ${drain.stderr}`); +if (!drain.stdout.includes("\t1\tsignal\tbranch-driver.status\t") || + !drain.stdout.includes("\t2\tsignal\tdecision-task.status\t")) { + throw new Error(`fallback did not receive the released mixed queue: ${drain.stdout}`); +} +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "a mixed needs-decision recheck must keep routine branch delivery live: $out" + pass "a co-present needs-decision row neither vetoes nor falsely settles routine branch delivery" +} + test_settled_branch_prompt_releases_unacknowledged_grant() { local repo home out status repo="$TMP_ROOT/settled-grant-root" @@ -3678,7 +3761,17 @@ test_branch_dispatch_classifies_main_only_rows_and_writes_the_eligible_snapshot( LIB="$repo/.pi/extensions/lib/fm-branch-dispatch.ts" FM_HOME="$home" GRANT="$ROOT/bin/fm-wake-grant.sh" \ node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' import { pathToFileURL } from "node:url"; -import { readFileSync, writeFileSync } from "node:fs"; +import { syncBuiltinESMExports } from "node:module"; +import fs, { readFileSync, symlinkSync, unlinkSync, writeFileSync } from "node:fs"; + +const originalReadFileSync = fs.readFileSync; +let countedStatusPath = ""; +let countedStatusReads = 0; +fs.readFileSync = function(path, ...args) { + if (String(path) === countedStatusPath) countedStatusReads += 1; + return originalReadFileSync.call(this, path, ...args); +}; +syncBuiltinESMExports(); const { activateEligibleRowsOwner, scopeForUnreadWake, writeEligibleRowsSnapshot, releaseEligibleRowsSnapshot, BRANCH_ELIGIBLE_ROWS_FILE } = await import(pathToFileURL(process.env.LIB).href); @@ -3705,6 +3798,157 @@ for (const row of mainOnlyRows) { if (scope.corrupted) throw new Error(`an ordinary main-only row must not read as corrupted: ${row}`); } +// A needs-decision signal row is a main-only class too, marked by payload +// rather than kind (docs/pi-supervision-branch.md "Autonomy"): it is excluded +// from eligibleSeqs and named in needsDecisionKeys. A later stale row under the +// task's window alias remains individually claimable, while task-identity +// precedence keeps its complete wake on main until the decision row is read. +writeFileSync( + `${state}/.wake-queue`, + [ + "1\t1\tsignal\ttask-a.status\tneeds-decision: task-a.status", + "1\t2\tstale\tfm-window\tstale: later routine reminder", + ].join("\n"), +); +const needsDecisionMixed = scopeForUnreadWake(state, false); +if (!needsDecisionMixed.eligible) { + throw new Error(`an unread needs-decision row must not erase a later stale row: ${JSON.stringify(needsDecisionMixed)}`); +} +if (needsDecisionMixed.eligibleSeqs.join(",") !== "2") { + throw new Error(`a needs-decision row must be excluded from eligibleSeqs: ${JSON.stringify(needsDecisionMixed)}`); +} +if (needsDecisionMixed.needsDecisionKeys.join(",") !== "task-a.status") { + throw new Error(`needsDecisionKeys must name the excluded row: ${JSON.stringify(needsDecisionMixed)}`); +} +if (needsDecisionMixed.taskByWakeKey["task-a.status"] !== "task-a" || + needsDecisionMixed.taskByWakeKey["fm-window"] !== "task-a") { + throw new Error(`status and stale aliases did not resolve to one task: ${JSON.stringify(needsDecisionMixed)}`); +} +if (needsDecisionMixed.corrupted) { + throw new Error(`a needs-decision row must not read as corrupted: ${JSON.stringify(needsDecisionMixed)}`); +} + +// A queue holding only a needs-decision row is ordinary main-only absence, +// exactly like a queue holding only a check row. +writeFileSync(`${state}/.wake-queue`, "1\t1\tsignal\ttask-a.status\tneeds-decision: task-a.status"); +const needsDecisionOnly = scopeForUnreadWake(state, false); +if (needsDecisionOnly.eligible || needsDecisionOnly.eligibleSeqs.length !== 0 || needsDecisionOnly.corrupted) { + throw new Error(`a needs-decision-only queue must be ordinary main-only absence: ${JSON.stringify(needsDecisionOnly)}`); +} + +// A captain-held task's bounded stale recheck is itself a decision wake. It is +// excluded while an unrelated routine row remains independently branch-owned. +writeFileSync(`${state}/task-a.status`, "captain-held [key=route]: awaiting the captain\n \t \n"); +writeFileSync( + `${state}/.wake-queue`, + [ + "1\t1\tstale\tfm-window\tstale: fm-window (awaiting the captain)", + "1\t2\tsignal\ttask-a.status\tsignal: routine follow-up", + ].join("\n"), +); +const captainHeldMixed = scopeForUnreadWake(state, false); +if (!captainHeldMixed.eligible || captainHeldMixed.eligibleSeqs.join(",") !== "2") { + throw new Error(`a captain-held stale row was offered to the branch: ${JSON.stringify(captainHeldMixed)}`); +} +if (captainHeldMixed.needsDecisionKeys.join(",") !== "fm-window") { + throw new Error(`the captain-held stale key was not marked main-owned: ${JSON.stringify(captainHeldMixed)}`); +} + +writeFileSync(`${state}/task-a.status`, "captain-held [key=route]: awaiting a second captain reminder\n \n"); +writeFileSync( + `${state}/.wake-queue`, + [ + "1\t1\tstale\tfm-window\tstale: fm-window (first reminder)", + "1\t2\tstale\tfm-window\tstale: fm-window (second reminder)", + "1\t3\tsignal\ttask-a.status\tsignal: routine follow-up", + ].join("\n"), +); +countedStatusPath = `${state}/task-a.status`; +countedStatusReads = 0; +const repeatedCaptainHeld = scopeForUnreadWake(state, false); +if (countedStatusReads !== 1) { + throw new Error(`one status was read ${countedStatusReads} times for repeated stale rows`); +} +if (!repeatedCaptainHeld.eligible || repeatedCaptainHeld.eligibleSeqs.join(",") !== "3" || + repeatedCaptainHeld.needsDecisionKeys.join(",") !== "fm-window,fm-window") { + throw new Error(`repeated stale reminders changed classification: ${JSON.stringify(repeatedCaptainHeld)}`); +} +const repeatedCaptainHeldNextScan = scopeForUnreadWake(state, false); +if (countedStatusReads !== 1 || repeatedCaptainHeldNextScan.needsDecisionKeys.join(",") !== "fm-window,fm-window") { + throw new Error(`an unchanged status was not reused across scans: reads=${countedStatusReads} scope=${JSON.stringify(repeatedCaptainHeldNextScan)}`); +} +writeFileSync(`${state}/task-a.status`, "captain-held [key=route]: awaiting the captain\nworking: resumed after answer\n"); +const changedCaptainHeld = scopeForUnreadWake(state, false); +if (countedStatusReads !== 2 || changedCaptainHeld.eligibleSeqs.join(",") !== "1,2,3" || + changedCaptainHeld.needsDecisionKeys.length !== 0) { + throw new Error(`a changed status did not invalidate its cached decision: reads=${countedStatusReads} scope=${JSON.stringify(changedCaptainHeld)}`); +} +countedStatusPath = ""; +writeFileSync( + `${state}/.wake-queue`, + [ + "1\t1\tstale\tfm-window\tstale: fm-window (awaiting the captain)", + "1\t2\tsignal\ttask-a.status\tsignal: routine follow-up", + ].join("\n"), +); + +// A later unrelated status does not mask a still-open durable decision. The +// stale row remains main-owned while the routine signal stays branch-owned. +writeFileSync( + `${state}/task-a.status`, + "needs-decision [key=cleanup]: choose destructive cleanup\nworking: routine follow-up\n", +); +const openDecisionMixed = scopeForUnreadWake(state, false); +if (!openDecisionMixed.eligible || openDecisionMixed.eligibleSeqs.join(",") !== "2") { + throw new Error(`an open-decision stale row was offered to the branch: ${JSON.stringify(openDecisionMixed)}`); +} +if (openDecisionMixed.needsDecisionKeys.join(",") !== "fm-window") { + throw new Error(`the open-decision stale key was not marked main-owned: ${JSON.stringify(openDecisionMixed)}`); +} + +process.env.FM_CLASSIFY_RESOLVE_VERB = "answered"; +writeFileSync( + `${state}/task-a.status`, + "needs-decision [key=cleanup]: choose destructive cleanup\nanswered [key=cleanup]: remove generated files\n", +); +const customResolved = scopeForUnreadWake(state, false); +if (!customResolved.eligible || customResolved.eligibleSeqs.slice().sort().join(",") !== "1,2" || + customResolved.needsDecisionKeys.length !== 0) { + throw new Error(`a custom resolution verb left the stale decision open: ${JSON.stringify(customResolved)}`); +} + +process.env.FM_CLASSIFY_CAPTAIN_HELD_VERB = "awaiting-captain"; +writeFileSync(`${state}/task-a.status`, "awaiting-captain [key=cleanup]: awaiting the captain\n"); +const customHeld = scopeForUnreadWake(state, false); +if (!customHeld.eligible || customHeld.eligibleSeqs.join(",") !== "2" || + customHeld.needsDecisionKeys.join(",") !== "fm-window") { + throw new Error(`a custom captain-held verb was offered to the branch: ${JSON.stringify(customHeld)}`); +} +delete process.env.FM_CLASSIFY_RESOLVE_VERB; +delete process.env.FM_CLASSIFY_CAPTAIN_HELD_VERB; + +process.env.FM_CLASSIFY_RESERVED_KEY_PREFIXES = "secret-"; +writeFileSync( + `${state}/task-a.status`, + "needs-decision [key=pending-reply-x]: choose destructive cleanup\nworking: routine follow-up\n", +); +const customReservedPrefixes = scopeForUnreadWake(state, false); +if (!customReservedPrefixes.eligible || customReservedPrefixes.eligibleSeqs.join(",") !== "2" || + customReservedPrefixes.needsDecisionKeys.join(",") !== "fm-window") { + throw new Error(`configured reserved prefixes lost an open stale decision: ${JSON.stringify(customReservedPrefixes)}`); +} +delete process.env.FM_CLASSIFY_RESERVED_KEY_PREFIXES; + +writeFileSync(`${state}/symlink-target.status`, "needs-decision: external choice\n"); +unlinkSync(`${state}/task-a.status`); +symlinkSync(`${state}/symlink-target.status`, `${state}/task-a.status`); +const symlinkedStatus = scopeForUnreadWake(state, false); +if (!symlinkedStatus.corrupted || symlinkedStatus.eligible || symlinkedStatus.needsDecisionKeys.length !== 0) { + throw new Error(`a symlinked status file influenced stale routing: ${JSON.stringify(symlinkedStatus)}`); +} +unlinkSync(`${state}/task-a.status`); +writeFileSync(`${state}/task-a.status`, "working: routine work\n"); + writeFileSync( `${state}/.wake-queue`, [ @@ -4542,6 +4786,98 @@ EOF pass "a failed cursor write re-delivers a routine note exactly once more while a captain outcome stays deduplicated" } +test_extension_registered_provider_resolves_in_the_branch() { + local repo home out status + repo="$TMP_ROOT/extprov-root" + home="$TMP_ROOT/extprov-home" + mkdir -p "$home/state" "$home/config" + install_pi_branch_extension_fixture "$repo" + PLUGIN="$repo/.pi/extensions/fm-branch-supervision.ts" FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" \ + DRIVER_PRELUDE="$DRIVER_PRELUDE" node --input-type=module > "$TMP_ROOT/node-output" 2>&1 <<'EOF' +const prelude = process.env.DRIVER_PRELUDE; +await eval(`(async () => { ${prelude}; globalThis.__t = { fire, dispatch, settle, makeCtx, registryModels, uiSelections, uiPrompts, notices, commands, home }; })()`); +const { fire, dispatch, settle, makeCtx, registryModels, uiSelections, uiPrompts, notices, commands, home } = globalThis.__t; +import { readFileSync, writeFileSync } from "node:fs"; + +// An extension-registered provider exists only in main's registry, never in +// the isolated branch runtime's static catalog. Registering its config on +// main's registry is what makes it resolvable for the branch. +registryModels.push( + { provider: "anthropic", id: "main-model" }, + // Available in main's registry but absent from the branch runtime's static + // catalog, exactly like a provider an extension registered at runtime. + { provider: "devin", id: "swe-1-7", branchAvailable: false }, +); +globalThis.__fmExtensionProviderConfigs = new Map([ + [ + "devin", + { + name: "Devin (Cognition)", + api: "devin-cloud", + baseUrl: "https://server.codeium.com", + models: [{ id: "swe-1-7", name: "SWE 1.7", reasoning: false, input: ["text"], cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, contextWindow: 200000, maxTokens: 8192 }], + oauth: { name: "Devin (Cognition / Windsurf)", login: async () => ({}), refreshToken: async (c) => c, getApiKey: (c) => c.access }, + streamSimple: () => {}, + }, + ], +]); + +await fire("session_start", {}, makeCtx()); + +// The picker must offer the extension-registered model: it is available in +// main's registry and resolvable in the branch runtime once its registration +// is copied across. +const command = commands.get("supervision-model"); +if (!command) throw new Error("the supervision-model command was not registered"); +uiSelections.push("devin/swe-1-7"); +await command.handler("", makeCtx()); +const offered = uiPrompts[0]; +if (!offered.options.includes("devin/swe-1-7")) { + throw new Error(`the picker must offer an extension-registered provider the branch can run: ${JSON.stringify(offered.options)}`); +} +if (readFileSync(`${home}/config/supervision-branch-model`, "utf8") !== "devin/swe-1-7\n") { + throw new Error("the extension-registered pick was not persisted"); +} +dispatch("signal: extension provider pin"); +await settle(() => (globalThis.__fmSessions ?? []).length === 1, "pinned extension-provider branch build"); +const pinned = globalThis.__fmSessions[0].options.model; +if (!pinned || pinned.provider !== "devin" || pinned.id !== "swe-1-7") { + throw new Error(`the extension-registered pin did not bind the branch: ${JSON.stringify(pinned)}`); +} +// Copying the provider registration must not loosen the branch's isolation: +// the devin-pinned session still loads no extensions, skills, or context files. +const pinnedLoader = globalThis.__fmLoaders.at(-1); +for (const key of ["noExtensions", "noSkills", "noContextFiles"]) { + if (pinnedLoader.options[key] !== true) throw new Error(`devin-pinned branch loader must keep ${key}`); +} + +// Without the registration, the same pin is unavailable and the branch +// refuses to build rather than silently downgrading. +globalThis.__fmExtensionProviderConfigs = new Map(); +await fire("session_shutdown", {}); +await fire("session_start", {}, makeCtx()); +const unregisteredOffer = dispatch("signal: unregistered provider pin"); +if (!unregisteredOffer.accepted) throw new Error("unregistered-pin wake was not initially accepted"); +const unregisteredFailure = await unregisteredOffer.settlement.then( + () => null, + (error) => error, +); +if ( + !(unregisteredFailure instanceof Error) || + !unregisteredFailure.message.includes("devin/swe-1-7") || + !unregisteredFailure.message.includes("supervision model pin") +) { + throw new Error(`the unregistered pin did not reject with its own name: ${String(unregisteredFailure)}`); +} +if ((globalThis.__fmSessions ?? []).length !== 1) throw new Error("an unregistered pin must not build a second branch session"); +process.exit(0); +EOF + status=$? + out=$(cat "$TMP_ROOT/node-output") + expect_code 0 "$status" "an extension-registered provider must resolve in the isolated branch runtime: $out" + pass "an extension-registered provider resolves in the isolated branch runtime" +} + test_outcomes_tool_uses_stock_execution_and_export_consumers test_real_pi_picker_primitives_stay_bounded_and_searchable test_branch_dispatch_two_stage_filter_and_prefix_contract @@ -4554,6 +4890,7 @@ test_branch_default_on_heartbeat_afk_and_fallback test_branch_predrain_recheck_keeps_a_heartbeat_a_co_present_check_arrives_under test_branch_report_refuses_a_task_the_wake_did_not_name test_branch_predrain_recheck_excludes_new_main_owned_row_without_deferring_eligible_work +test_branch_predrain_needs_decision_keeps_routine_row_branch_eligible test_settled_branch_prompt_releases_unacknowledged_grant test_post_construction_provider_error_falls_back_latches_and_recovers_on_cooldown test_selection_change_does_not_corrupt_inflight_provider_state @@ -4569,6 +4906,7 @@ test_supervision_model_picker_is_bounded_searchable_and_branch_only test_branch_model_picker_keeps_follow_main_first_under_ranking test_branch_effort_pin_applies_and_absent_pin_follows_main test_unpinned_branch_follows_main_effort_changes_live +test_extension_registered_provider_resolves_in_the_branch test_supervision_model_command_picks_effort_after_the_model test_unusable_model_pin_falls_back_to_main test_replacement_activation_cleans_leases_and_retries_failure diff --git a/tests/fm-pi-branch-live-e2e.test.sh b/tests/fm-pi-branch-live-e2e.test.sh index 06e2bad1458..5ad17da8fde 100644 --- a/tests/fm-pi-branch-live-e2e.test.sh +++ b/tests/fm-pi-branch-live-e2e.test.sh @@ -30,13 +30,10 @@ # (docs/verification/runtime-backends.md). set -u -if [ "${FM_PI_BRANCH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_BRANCH_LIVE_E2E=1 to run the real-SDK Pi branch regression" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_PI_BRANCH_LIVE_E2E npm jq node export NODE_NO_WARNINGS=1 PI_PACKAGE_DIR=${FM_PI_PACKAGE_DIR:-"$(npm root -g)/@earendil-works/pi-coding-agent"} diff --git a/tests/fm-pi-branch-responsiveness-live-e2e.test.sh b/tests/fm-pi-branch-responsiveness-live-e2e.test.sh index 00c8d2f8e2f..65dd75cd824 100755 --- a/tests/fm-pi-branch-responsiveness-live-e2e.test.sh +++ b/tests/fm-pi-branch-responsiveness-live-e2e.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Opt-in live guard for the ONE thing only a real Pi TUI can answer: whether +# Default-on live guard for the ONE thing only a real Pi TUI can answer: whether # the captain can still type and see the screen repaint while a supervision # outcome is being delivered into his session. # @@ -28,17 +28,10 @@ # which runs the same reconcile-and-deliver chain an arriving outcome runs. set -u -if [ "${FM_PI_BRANCH_RESPONSIVENESS_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_BRANCH_RESPONSIVENESS_E2E=1 to run the real-Pi delivery responsiveness guard" - exit 0 -fi - # shellcheck source=tests/lib.sh . "$(dirname "${BASH_SOURCE[0]}")/lib.sh" -command -v pi >/dev/null 2>&1 || fail "pi not installed: the real-Pi responsiveness guard cannot report a verdict without it" -command -v tmux >/dev/null 2>&1 || fail "tmux not installed: the real-Pi responsiveness guard cannot drive a pane without it" -command -v node >/dev/null 2>&1 || fail "node not installed: the real-Pi responsiveness guard cannot time keystroke echo without it" +fm_live_gate default-on FM_PI_BRANCH_RESPONSIVENESS_E2E pi tmux node PI_VERSION=$(pi --version 2>/dev/null || printf 'unknown') TMUX=$(command -v tmux) @@ -57,7 +50,7 @@ cleanup() { trap cleanup EXIT cp "$ROOT/.pi/extensions/fm-branch-supervision.ts" "$PROJECT/.pi/extensions/fm-branch-supervision.ts" -for lib in fm-async-exec fm-branch-dispatch fm-branch-model-picker fm-calm-visibility fm-operational-input; do +for lib in fm-async-exec fm-branch-dispatch fm-branch-model-picker fm-calm-visibility fm-operational-input fm-process-ancestry; do cp "$ROOT/.pi/extensions/lib/$lib.ts" "$PROJECT/.pi/extensions/lib/$lib.ts" done diff --git a/tests/fm-pi-primary-live-e2e.test.sh b/tests/fm-pi-primary-live-e2e.test.sh index a6680f4b7df..f76a5f9c8ed 100755 --- a/tests/fm-pi-primary-live-e2e.test.sh +++ b/tests/fm-pi-primary-live-e2e.test.sh @@ -4,10 +4,10 @@ # copying credentials and pins the captain-approved openai-codex model. set -u -if [ "${FM_PI_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_PI_LIVE_E2E=1 to run the isolated interactive Pi regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_PI_LIVE_E2E pi tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -17,9 +17,6 @@ fail() { exit 1 } -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" - TMUX=$(command -v tmux) SOCKET="fm-pi-live-e2e-$$" SESSION=pi-live-e2e diff --git a/tests/fm-pi-primary-types.test.sh b/tests/fm-pi-primary-types.test.sh index 650ce7a4626..56030968660 100755 --- a/tests/fm-pi-primary-types.test.sh +++ b/tests/fm-pi-primary-types.test.sh @@ -4,12 +4,12 @@ set -u ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -command -v npm >/dev/null 2>&1 || { echo "skip: npm not found for Pi extension typecheck"; exit 0; } -command -v tsc >/dev/null 2>&1 || { echo "skip: tsc not found for Pi extension typecheck"; exit 0; } +command -v npm >/dev/null 2>&1 || { echo "skip: Pi extension typecheck prerequisite not found: npm"; exit 0; } +command -v tsc >/dev/null 2>&1 || { echo "skip: Pi extension typecheck prerequisite not found: tsc"; exit 0; } PI_PACKAGE_DIR=${FM_PI_PACKAGE_DIR:-"$(npm root -g)/@earendil-works/pi-coding-agent"} if [ ! -f "$PI_PACKAGE_DIR/package.json" ]; then - echo "skip: installed @earendil-works/pi-coding-agent package not found" + echo "skip: Pi extension typecheck prerequisite not found: installed @earendil-works/pi-coding-agent package" exit 0 fi if [ ! -d "$PI_PACKAGE_DIR/node_modules/typebox" ] || \ diff --git a/tests/fm-pi-watch-extension.test.sh b/tests/fm-pi-watch-extension.test.sh index 751334a5d05..7a8fab58822 100755 --- a/tests/fm-pi-watch-extension.test.sh +++ b/tests/fm-pi-watch-extension.test.sh @@ -811,6 +811,376 @@ CLASSES pass "every main-only check class still reaches main, never the supervision branch" } +# A surfaced captain-held signal uses the existing decision-owned payload, so a +# co-present routine row cannot take the signal close away from main. +test_pi_captain_held_signal_stays_on_main() { + local repo home plugin log stop out status + repo="$TMP_ROOT/pi-captain-held-whitespace-root" + home="$TMP_ROOT/pi-captain-held-whitespace-home" + log="$TMP_ROOT/pi-captain-held-whitespace.log" + stop="$TMP_ROOT/pi-captain-held-whitespace.stop" + mkdir -p "$repo/bin" "$home/state" "$home/config" "$home/projects/approved" + install_pi_watch_extension_fixture "$repo" + plugin="$repo/.pi/extensions/fm-primary-pi-watch.ts" + printf 'project=%s/projects/approved\nwindow=fm-window\n' "$home" > "$home/state/task-a.meta" + printf 'captain-held [key=route]: awaiting the captain\n \t \n' > "$home/state/task-a.status" + cat > "$repo/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = --handling-delivered ]; then exit 0; fi +printf 'arm=%s\n' "$$" >> "${FM_ARM_LOG:?}" +count=$(grep -c '^arm=' "$FM_ARM_LOG") +if [ "$count" -eq 1 ]; then + printf 'watcher: started pid=%s (beacon fresh)\n' "$$" + printf 'signal: task-a.status\n' + exit 0 +fi +printf 'watcher: started pid=%s (beacon fresh) recovery-generation=fixture-generation\n' "$$" +trap 'exit 0' TERM INT +while [ ! -e "$FM_STOP_FILE" ]; do sleep 0.02; done +SH + chmod +x "$repo/bin/fm-watch-arm.sh" + out=$(PLUGIN="$plugin" FM_HOME="$home" FM_ROOT_OVERRIDE="$repo" FM_ARM_LOG="$log" FM_STOP_FILE="$stop" \ + node --input-type=module 2>&1 <<'EOF' +import { writeFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const offers = []; +let prompt = ""; +let tool = null; +const handlers = new Map(); +const bus = { + on(channel, handler) { + handlers.set(channel, [...(handlers.get(channel) ?? []), handler]); + return () => {}; + }, + emit(channel, data) { + for (const handler of handlers.get(channel) ?? []) handler(data); + }, +}; +bus.on("fm-branch-supervision:dispatch", (offer) => { + offers.push({ message: offer.message, eligible: offer.eligible }); + if (offer.eligible) offer.accept(); +}); +const pi = { + on() {}, + events: bus, + registerCommand() {}, + registerTool(candidate) { + if (candidate.name === "fm_watch_arm_pi") tool = candidate; + }, + sendUserMessage: async (message) => { + prompt = message; + }, +}; +writeFileSync(`${process.env.FM_HOME}/state/.lock`, `${process.pid}\n`); +writeFileSync( + `${process.env.FM_HOME}/state/.wake-queue`, + "1\t1\tsignal\ttask-a.status\tneeds-decision: task-a.status\n" + + "1\t2\tsignal\ttask-a.status\tsignal: routine follow-up\n", +); +const mod = await import(pathToFileURL(process.env.PLUGIN).href); +mod.default(pi); +await tool.execute("tool-call-captain-held-signal", {}, undefined, undefined, {}); +for (let i = 0; i < 250 && !prompt; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +if (offers.length !== 1 || offers[0].eligible !== false) { + throw new Error(`a captain-held signal trigger was offered to the branch: ${JSON.stringify(offers)}`); +} +if (!prompt.includes("FIRSTMATE WATCHER WAKE: signal: task-a.status")) { + throw new Error(`a captain-held signal trigger did not reach main: ${prompt}`); +} +writeFileSync(process.env.FM_STOP_FILE, "stop\n"); +process.exit(0); +EOF + ) + status=$? + expect_code 0 "$status" "a captain-held signal trigger must stay on main: $out" + [ -z "$out" ] || fail "Pi captain-held signal test printed output: $out" + pass "a captain-held signal trigger reaches main with routine rows present" +} + +# An unread second-mate pending-reply escalation keeps a later stale reminder +# under the same task's window alias on main until the decision row is read. +test_pi_unread_pending_reply_forces_later_stale_alias_to_main() { + local repo home plugin log stop out status + repo="$TMP_ROOT/pi-mixed-signal-root" + home="$TMP_ROOT/pi-mixed-signal-home" + log="$TMP_ROOT/pi-mixed-signal.log" + stop="$TMP_ROOT/pi-mixed-signal.stop" + mkdir -p "$repo/bin" "$home/state" "$home/config" "$home/projects/approved" + install_pi_watch_extension_fixture "$repo" + plugin="$repo/.pi/extensions/fm-primary-pi-watch.ts" + printf 'project=%s/projects/approved\nwindow=fm-a\n' "$home" > "$home/state/task-a.meta" + printf 'blocked [key=pending-reply-0123456789abcdef]: pending-reply-missed: task=task-a pending-reply-id=0123456789abcdef request=finish report\n' \ + > "$home/state/task-a.status" + cat > "$repo/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = --handling-delivered ]; then exit 0; fi +printf 'arm=%s\n' "$$" >> "${FM_ARM_LOG:?}" +count=$(grep -c '^arm=' "$FM_ARM_LOG") +if [ "$count" -eq 1 ]; then + printf 'watcher: started pid=%s (beacon fresh)\n' "$$" + printf 'stale: fm-a (routine reminder)\n' + exit 0 +fi +printf 'watcher: started pid=%s (beacon fresh) recovery-generation=fixture-generation\n' "$$" +trap 'exit 0' TERM INT +while [ ! -e "$FM_STOP_FILE" ]; do sleep 0.02; done +SH + chmod +x "$repo/bin/fm-watch-arm.sh" + out=$(PLUGIN="$plugin" FM_HOME="$home" FM_ROOT_OVERRIDE="$repo" FM_ARM_LOG="$log" FM_STOP_FILE="$stop" \ + node --input-type=module 2>&1 <<'EOF' +import { writeFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const offers = []; +let prompt = ""; +let tool = null; +const handlers = new Map(); +const bus = { + on(channel, handler) { + handlers.set(channel, [...(handlers.get(channel) ?? []), handler]); + return () => {}; + }, + emit(channel, data) { + for (const handler of handlers.get(channel) ?? []) handler(data); + }, +}; +bus.on("fm-branch-supervision:dispatch", (offer) => { + offers.push({ message: offer.message, eligible: offer.eligible, projects: [...offer.projects] }); + if (offer.eligible) offer.accept(); +}); +const pi = { + on() {}, + events: bus, + registerCommand() {}, + registerTool(candidate) { + if (candidate.name === "fm_watch_arm_pi") tool = candidate; + }, + sendUserMessage: async (message) => { + prompt = message; + }, +}; +writeFileSync(`${process.env.FM_HOME}/state/.lock`, `${process.pid}\n`); +writeFileSync( + `${process.env.FM_HOME}/state/.wake-queue`, + "1\t1\tsignal\ttask-a.status\tneeds-decision: task-a.status\n" + + "2\t2\tstale\tfm-a\tstale: fm-a (routine reminder)\n", +); +const mod = await import(pathToFileURL(process.env.PLUGIN).href); +mod.default(pi); +await tool.execute("tool-call-later-stale-alias", {}, undefined, undefined, {}); +for (let i = 0; i < 250 && offers.length === 0; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +for (let i = 0; i < 250 && !prompt; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +if (offers.length !== 1 || offers[0].eligible !== false) { + throw new Error(`a later stale alias bypassed its unread decision: ${JSON.stringify(offers)}`); +} +if (!prompt.includes("FIRSTMATE WATCHER WAKE: stale: fm-a")) { + throw new Error(`a later stale alias with an unread decision did not wake main: ${prompt}`); +} +writeFileSync(process.env.FM_STOP_FILE, "stop\n"); +process.exit(0); +EOF + ) + status=$? + expect_code 0 "$status" "an unread pending-reply escalation must keep a later stale alias on main: $out" + [ -z "$out" ] || fail "Pi unread pending-reply alias test printed output: $out" + pass "an unread pending-reply escalation keeps later stale aliases on main" +} + +# The captain's accepted rule names ONE coalesced trigger batch, not only a +# single status file with two rows: when a watcher poll bundles a routine +# signal for one task and a needs-decision signal for a DIFFERENT task into +# one "signal: <files>" close, the whole batch - both files - must reach main +# together, never split so the routine file goes to the branch while only the +# decision file goes to main. A `some`-vs-`every` regression on the file-list +# cross-reference would pass the same-key test above (both entries share one +# key) but must fail here, where the two files are genuinely distinct. +test_pi_distinct_files_mixed_batch_routes_whole_batch_to_main() { + local repo home plugin log stop out status + repo="$TMP_ROOT/pi-distinct-mixed-batch-root" + home="$TMP_ROOT/pi-distinct-mixed-batch-home" + log="$TMP_ROOT/pi-distinct-mixed-batch.log" + stop="$TMP_ROOT/pi-distinct-mixed-batch.stop" + mkdir -p "$repo/bin" "$home/state" "$home/config" "$home/projects/approved" + install_pi_watch_extension_fixture "$repo" + plugin="$repo/.pi/extensions/fm-primary-pi-watch.ts" + printf 'project=%s/projects/approved\nwindow=fm-a\n' "$home" > "$home/state/task-a.meta" + printf 'project=%s/projects/approved\nwindow=fm-b\n' "$home" > "$home/state/task-b.meta" + cat > "$repo/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = --handling-delivered ]; then exit 0; fi +printf 'arm=%s\n' "$$" >> "${FM_ARM_LOG:?}" +count=$(grep -c '^arm=' "$FM_ARM_LOG") +if [ "$count" -eq 1 ]; then + printf 'watcher: started pid=%s (beacon fresh)\n' "$$" + printf 'signal: task-a.status task-b.status\n' + exit 0 +fi +printf 'watcher: started pid=%s (beacon fresh) recovery-generation=fixture-generation\n' "$$" +trap 'exit 0' TERM INT +while [ ! -e "$FM_STOP_FILE" ]; do sleep 0.02; done +SH + chmod +x "$repo/bin/fm-watch-arm.sh" + out=$(PLUGIN="$plugin" FM_HOME="$home" FM_ROOT_OVERRIDE="$repo" FM_ARM_LOG="$log" FM_STOP_FILE="$stop" \ + node --input-type=module 2>&1 <<'EOF' +import { writeFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const offers = []; +let prompt = ""; +let tool = null; +const handlers = new Map(); +const bus = { + on(channel, handler) { + handlers.set(channel, [...(handlers.get(channel) ?? []), handler]); + return () => {}; + }, + emit(channel, data) { + for (const handler of handlers.get(channel) ?? []) handler(data); + }, +}; +bus.on("fm-branch-supervision:dispatch", (offer) => { + offers.push({ message: offer.message, eligible: offer.eligible, projects: [...offer.projects] }); + if (offer.eligible) offer.accept(); +}); +const pi = { + on() {}, + events: bus, + registerCommand() {}, + registerTool(candidate) { + if (candidate.name === "fm_watch_arm_pi") tool = candidate; + }, + sendUserMessage: async (message) => { + prompt = message; + }, +}; +writeFileSync(`${process.env.FM_HOME}/state/.lock`, `${process.pid}\n`); +writeFileSync( + `${process.env.FM_HOME}/state/.wake-queue`, + "1\t1\tsignal\ttask-a.status\tneeds-decision: task-a.status\n" + + "2\t2\tsignal\ttask-b.status\tsignal: task-b.status\n", +); +const mod = await import(pathToFileURL(process.env.PLUGIN).href); +mod.default(pi); +await tool.execute("tool-call-distinct-mixed-batch", {}, undefined, undefined, {}); +for (let i = 0; i < 250 && offers.length === 0; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +for (let i = 0; i < 250 && !prompt; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +if (offers.length !== 1 || offers[0].eligible !== false) { + throw new Error(`a distinct-file mixed batch was split, offering the routine file to the branch: ${JSON.stringify(offers)}`); +} +if (!prompt.includes("FIRSTMATE WATCHER WAKE: signal: task-a.status task-b.status")) { + throw new Error(`the distinct-file mixed batch did not reach main as one whole close: ${prompt}`); +} +writeFileSync(process.env.FM_STOP_FILE, "stop\n"); +process.exit(0); +EOF + ) + status=$? + expect_code 0 "$status" "a distinct-file mixed batch must route wholly to main, not split: $out" + [ -z "$out" ] || fail "Pi distinct-file mixed-batch test printed output: $out" + pass "a mixed batch of two distinct files - one routine, one needs-decision - routes wholly to main" +} + +# Independent heartbeat handling: a needs-decision row sitting elsewhere in the +# unread queue is excluded and non-vetoing exactly like a check-kind row, so it +# must neither block nor ride along with a co-present, otherwise-eligible +# heartbeat scan (docs/pi-supervision-branch.md "Heartbeat routing"). +test_pi_heartbeat_is_not_ridden_into_main_by_a_co_present_needs_decision() { + local repo home plugin log stop out status + repo="$TMP_ROOT/pi-heartbeat-needs-decision-root" + home="$TMP_ROOT/pi-heartbeat-needs-decision-home" + log="$TMP_ROOT/pi-heartbeat-needs-decision.log" + stop="$TMP_ROOT/pi-heartbeat-needs-decision.stop" + mkdir -p "$repo/bin" "$home/state" "$home/config" + install_pi_watch_extension_fixture "$repo" + plugin="$repo/.pi/extensions/fm-primary-pi-watch.ts" + cat > "$repo/bin/fm-watch-arm.sh" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = --handling-delivered ]; then exit 0; fi +printf 'arm=%s\n' "$$" >> "${FM_ARM_LOG:?}" +count=$(grep -c '^arm=' "$FM_ARM_LOG") +if [ "$count" -eq 1 ]; then + printf 'watcher: started pid=%s (beacon fresh)\n' "$$" + printf 'heartbeat\n' + exit 0 +fi +printf 'watcher: started pid=%s (beacon fresh) recovery-generation=fixture-generation\n' "$$" +trap 'exit 0' TERM INT +while [ ! -e "$FM_STOP_FILE" ]; do sleep 0.02; done +SH + chmod +x "$repo/bin/fm-watch-arm.sh" + out=$(PLUGIN="$plugin" FM_HOME="$home" FM_ROOT_OVERRIDE="$repo" FM_ARM_LOG="$log" FM_STOP_FILE="$stop" node --input-type=module 2>&1 <<'EOF' +import { writeFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const offers = []; +let prompt = ""; +let tool = null; +const handlers = new Map(); +const bus = { + on(channel, handler) { + handlers.set(channel, [...(handlers.get(channel) ?? []), handler]); + return () => {}; + }, + emit(channel, data) { + for (const handler of handlers.get(channel) ?? []) handler(data); + }, +}; +bus.on("fm-branch-supervision:dispatch", (offer) => { + offers.push({ message: offer.message, heartbeat: offer.heartbeat, eligible: offer.eligible }); + if (offer.eligible) offer.accept(); +}); +const pi = { + on() {}, + events: bus, + registerCommand() {}, + registerTool(candidate) { + if (candidate.name === "fm_watch_arm_pi") tool = candidate; + }, + sendUserMessage: async (message) => { + prompt = message; + }, +}; +writeFileSync(`${process.env.FM_HOME}/state/.lock`, `${process.pid}\n`); +writeFileSync( + `${process.env.FM_HOME}/state/.wake-queue`, + "1\t1\theartbeat\theartbeat\theartbeat\n2\t2\tsignal\tdecision-task.status\tneeds-decision: decision-task.status\n", +); +const mod = await import(pathToFileURL(process.env.PLUGIN).href); +mod.default(pi); +await tool.execute("tool-call-heartbeat-needs-decision", {}, undefined, undefined, {}); +for (let i = 0; i < 250 && offers.length === 0; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +for (let i = 0; i < 25 && !prompt; i += 1) { + await new Promise((resolve) => setTimeout(resolve, 10)); +} +if (offers.length !== 1 || offers[0].heartbeat !== true || offers[0].eligible !== true) { + throw new Error(`a co-present needs-decision row made the heartbeat offer ineligible: ${JSON.stringify(offers)}`); +} +if (prompt) { + throw new Error(`a co-present needs-decision row rode the heartbeat into main: ${prompt}`); +} +writeFileSync(process.env.FM_STOP_FILE, "stop\n"); +process.exit(0); +EOF + ) + status=$? + expect_code 0 "$status" "a heartbeat must not ride a co-present needs-decision row into main: $out" + [ -z "$out" ] || fail "Pi heartbeat/needs-decision test printed output: $out" + pass "a co-present needs-decision row neither vetoes nor rides a heartbeat into main" +} + test_pi_heartbeat_restoration_failure_stays_on_main() { local repo home plugin log out status repo="$TMP_ROOT/pi-heartbeat-restoration-failure-root" @@ -3627,6 +3997,10 @@ test_pi_branch_offer_owns_actionable_wake test_pi_branch_offer_flags_heartbeat test_pi_heartbeat_is_not_ridden_into_main_by_a_co_present_check test_pi_main_only_check_classes_stay_on_main +test_pi_captain_held_signal_stays_on_main +test_pi_unread_pending_reply_forces_later_stale_alias_to_main +test_pi_distinct_files_mixed_batch_routes_whole_batch_to_main +test_pi_heartbeat_is_not_ridden_into_main_by_a_co_present_needs_decision test_pi_heartbeat_restoration_failure_stays_on_main test_pi_watcher_failure_never_offered_to_branch test_pi_handling_delivery_failure_is_typed_once diff --git a/tests/fm-pi-windows-shell-invocation.test.sh b/tests/fm-pi-windows-shell-invocation.test.sh new file mode 100755 index 00000000000..b85493c6c13 --- /dev/null +++ b/tests/fm-pi-windows-shell-invocation.test.sh @@ -0,0 +1,120 @@ +#!/usr/bin/env bash +# Native-Windows Pi extension regression for invoking tracked Bash owners through bash. +set -u + +ROOT=$(cd "$(dirname "$0")/.." && pwd) +. "$ROOT/tests/lib.sh" + +TMP_ROOT=$(fm_test_tmproot fm-pi-windows-shell-invocation) + +if [ "$(node -p 'process.platform')" != win32 ]; then + echo "skip: native Windows Node required" + exit 0 +fi + +project="$TMP_ROOT/project" +mkdir -p "$project/.pi/extensions/lib" "$project/bin" "$project/state" +cp "$ROOT/.pi/extensions/fm-primary-turnend-guard.ts" "$project/.pi/extensions/" +cp "$ROOT/.pi/extensions/lib/fm-operational-input.ts" \ + "$ROOT/.pi/extensions/lib/fm-process-ancestry.ts" \ + "$ROOT/.pi/extensions/lib/fm-sessionstart-supervisor.mjs" "$project/.pi/extensions/lib/" + +cat >"$project/bin/fm-sessionstart-run.sh" <<'SH' +#!/usr/bin/env bash +printf 'sessionstart:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-cd-pretool-check.sh" <<'SH' +#!/usr/bin/env bash +printf 'cd:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-arm-pretool-check.sh" <<'SH' +#!/usr/bin/env bash +printf 'arm:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-turnend-guard.sh" <<'SH' +#!/usr/bin/env bash +cat >/dev/null +printf 'turnend:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +SH +cat >"$project/bin/fm-operational-input.sh" <<'SH' +#!/usr/bin/env bash +printf 'operational:%s\n' "$*" >> "$FM_WINDOWS_SHELL_LOG" +input=$(cat) +if [ "$1" = encode ]; then + printf 'encoded:%s:%s\n' "$2" "$input" +else + printf 'not-operational\n' +fi +SH +chmod +x "$project/bin/"*.sh + +log="$project/state/calls" +out=$(EXT="$project/.pi/extensions/fm-primary-turnend-guard.ts" \ + FM_HOME="$project" FM_ROOT_OVERRIDE="$project" FM_WINDOWS_SHELL_LOG="$log" \ + FM_OPERATIONAL_INPUT_SCRIPT="$project/bin/fm-operational-input.sh" \ + node --input-type=module 2>&1 <<'JS' +import { spawn } from "node:child_process"; +import { readFileSync } from "node:fs"; +import { pathToFileURL } from "node:url"; + +const handlers = new Map(); +const pi = { + on(event, handler) { handlers.set(event, handler); }, + sendMessage() {}, +}; +const extension = await import(`${pathToFileURL(process.env.EXT).href}?windows=${Date.now()}`); +extension.default(pi); +const ctx = { sessionManager: { getSessionId: () => "windows-test" } }; +handlers.get("session_start")({ reason: "startup" }, ctx); +await handlers.get("before_agent_start")({}, ctx); +await handlers.get("tool_call")({ type: "tool_call", toolName: "bash", input: { command: "printf test" } }); +await handlers.get("agent_settled")({}, ctx); +const operational = await import(`${new URL("./lib/fm-operational-input.ts", pathToFileURL(process.env.EXT)).href}?windows=${Date.now()}`); +operational.classifyFirstmateOperationalText("probe"); +let asyncInvocation; +const encoded = await operational.encodeFirstmateOperationalInputWith( + (command, args, { input }) => { + asyncInvocation = { command, args: [...args], input }; + return new Promise((resolve, reject) => { + const child = spawn(command, args, { stdio: ["pipe", "pipe", "ignore"] }); + let stdout = ""; + child.stdout.on("data", (chunk) => { stdout += chunk; }); + child.on("error", reject); + child.on("close", (status) => resolve({ status, stdout })); + child.stdin.end(input); + }); + }, + "branch-outcome", + "branch result", +); +if (encoded !== "encoded:branch-outcome:branch result\n") { + throw new Error(`unexpected encoded branch outcome: ${encoded}`); +} +if ( + asyncInvocation.command !== "bash" || + asyncInvocation.args.join("\0") !== [ + process.env.FM_OPERATIONAL_INPUT_SCRIPT, + "encode", + "branch-outcome", + ].join("\0") || + asyncInvocation.input !== "branch result" +) { + throw new Error(`unexpected async invocation: ${JSON.stringify(asyncInvocation)}`); +} +const calls = readFileSync(process.env.FM_WINDOWS_SHELL_LOG, "utf8"); +for (const expected of [ + "sessionstart:--source startup --pi-prerequisite", + "cd:--command printf test", + "arm:--command printf test", + "turnend:", + "operational:classify", + "operational:encode branch-outcome", +]) { + if (!calls.includes(expected)) throw new Error(`missing ${expected} in:\n${calls}`); +} +JS +) +status=$? +expect_code 0 "$status" "native-Windows Pi shell seams${out:+: $out}" +[ -z "$out" ] || fail "native-Windows Pi shell seam test printed output: $out" +pass "Pi session-start, pre-tool, turn-end, and operational-input seams invoke Bash owners on native Windows" diff --git a/tests/fm-pr-check-security.test.sh b/tests/fm-pr-check-security.test.sh index 5bce7184d86..fddee11d938 100755 --- a/tests/fm-pr-check-security.test.sh +++ b/tests/fm-pr-check-security.test.sh @@ -1115,6 +1115,12 @@ test_postrename_poll_validation_revokes_and_retries() { local gate link_target_mode expected_link_target_mode for artifact in data registration check; do for action in type mode device content; do + # The device fault is injected by a fake stat on PATH; on Darwin the + # device helper now calls /usr/bin/stat directly, so the fake can never + # fire there. Skip the device action on Darwin. + if [ "$action" = device ] && [ "$(uname)" = Darwin ]; then + continue + fi dir=$(make_case "poll-final-$artifact-$action") state="$dir/home/state" write_poll_meta "$state" task-a https://github.com/o/r/pull/1 diff --git a/tests/fm-procevent-when.test.sh b/tests/fm-procevent-when.test.sh index 259286beb85..396c48df8d9 100755 --- a/tests/fm-procevent-when.test.sh +++ b/tests/fm-procevent-when.test.sh @@ -21,23 +21,10 @@ export FM_PROCEVENT_CLAIM_ROOT="$TMP_ROOT/claims" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } when() { FM_HOME="$1" "$ROOT/bin/fm-procevent-when.sh" "${@:2}"; } -# Every home this suite arms is tracked so teardown can stop any runner still -# blocked on a condition that never fires. -WHEN_HOMES=() -when_teardown() { - local home seen=$'\n' - for home in ${WHEN_HOMES[@]+"${WHEN_HOMES[@]}"}; do - case "$seen" in - *$'\n'"$home"$'\n'*) continue ;; - esac - seen+="$home"$'\n' - FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true - done - fm_test_cleanup -} -trap when_teardown EXIT - -new_home() { mkdir -p "$1/state"; WHEN_HOMES+=("$1"); } +# Every home this suite arms is registered with tests/lib.sh, which sweeps it +# from every cleanup path so a runner still blocked on a condition that never +# fires cannot survive the run. +new_home() { mkdir -p "$1/state"; fm_test_track_procevent_home "$1"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } diff --git a/tests/fm-procevent.test.sh b/tests/fm-procevent.test.sh index fd04351a304..5c5a95112f3 100755 --- a/tests/fm-procevent.test.sh +++ b/tests/fm-procevent.test.sh @@ -24,9 +24,14 @@ BLOCKER="$TMP_ROOT/blocker.sh" cat > "$BLOCKER" <<'SH' #!/usr/bin/env bash # Blocks until the trigger exists, then emits its payload. Completion is the -# event; nothing here polls on a schedule. +# event; nothing here polls on a schedule. The wait is bounded so a stub that +# escapes its test cannot keep spawning processes indefinitely. trigger=$1; shift -while [ ! -e "$trigger" ]; do sleep 0.05; done +[ -z "${BLOCKER_READY:-}" ] || printf 'ready\n' > "$BLOCKER_READY" +while [ ! -e "$trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.05 +done [ -n "${BLOCKER_STDERR:-}" ] && printf 'noise on stderr\n' >&2 [ -n "${BLOCKER_EXIT:-}" ] && exit "$BLOCKER_EXIT" printf '%s\n' "$@" @@ -35,31 +40,17 @@ chmod +x "$BLOCKER" pe() { FM_HOME="$1" "$ROOT/bin/fm-procevent.sh" "${@:2}"; } -# Every source this suite registers is tracked so teardown can stop its runner. -# A runner started by reconcile is detached and reparented, so a source that -# never completes outlives the suite unless it is retired explicitly - removing -# the fixture directory does not stop an already-running child. -PE_TRACKED=() +# Every home this suite registers a source in is tracked so teardown can stop +# its runners. A runner started by reconcile is detached and reparented, so a +# source that never completes outlives the suite unless its home is swept - +# removing the fixture directory does not stop an already-running child. +# tests/lib.sh owns that sweep and runs it from every cleanup path. pe_register() { # <home> <adapter> <source-id> -- <argv>... local home=$1 adapter=$2 id=$3 shift 3 - PE_TRACKED+=("$home|$id") + fm_test_track_procevent_home "$home" pe "$home" register "$adapter" "$id" "$@" } - -procevent_teardown() { - local entry home seen=$'\n' - for entry in ${PE_TRACKED[@]+"${PE_TRACKED[@]}"}; do - home=${entry%%|*} - case "$seen" in - *$'\n'"$home"$'\n'*) continue ;; - esac - seen+="$home"$'\n' - FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true - done - fm_test_cleanup -} -trap procevent_teardown EXIT new_home() { mkdir -p "$1/state"; } wake_payloads() { awk -F '\t' '{print $5}' "$1/state/.wake-queue" 2>/dev/null; } @@ -432,7 +423,7 @@ pe_adapter() { # <home> <command>...: run the runner against the fixture adapte } HPUBLISH="$TMP_ROOT/hpublish"; new_home "$HPUBLISH" -PE_TRACKED+=("$HPUBLISH|publish-src") +fm_test_track_procevent_home "$HPUBLISH" pe_adapter "$HPUBLISH" register applying publish-src -- /bin/echo "apply after publish" >/dev/null mkdir "$HPUBLISH/state/.wake-queue" out=$(pe_adapter "$HPUBLISH" start publish-src 2>&1) @@ -464,7 +455,7 @@ pass "automatic application waits for durable publication and failed publication # channel is the announcement. The declaration never silences a capture the # adapter could NOT apply - that one still publishes for the handler. HSELF="$TMP_ROOT/hself"; new_home "$HSELF" -PE_TRACKED+=("$HSELF|self-src") +fm_test_track_procevent_home "$HSELF" pe_adapter "$HSELF" register selfann self-src -- /bin/echo "self announced" >/dev/null out=$(pe_adapter "$HSELF" start self-src 2>&1) assert_contains "$out" "autohandled: self-src" "the self-announcing adapter did not apply its own capture" @@ -495,7 +486,7 @@ rm -f "$HSELF/state/selfann-fail" pass "a self-announcing adapter applies quietly and still publishes what it could not apply" HTERM="$TMP_ROOT/hterm"; new_home "$HTERM" -PE_TRACKED+=("$HTERM|ends-src") +fm_test_track_procevent_home "$HTERM" pe_adapter "$HTERM" register endnow ends-src -- /bin/echo "terminal payload" >/dev/null out=$(pe_adapter "$HTERM" start ends-src) assert_contains "$out" "captured:" "a terminal result is still captured durably" @@ -519,7 +510,7 @@ assert_contains "$out" "published=0" "an acknowledged terminal result stops bein pass "an adapter-classified terminal result is captured once, announced, and retires its source automatically" HOPEN="$TMP_ROOT/hopen"; new_home "$HOPEN" -PE_TRACKED+=("$HOPEN|open-src") +fm_test_track_procevent_home "$HOPEN" pe_adapter "$HOPEN" register openended open-src -- /bin/echo "open payload" >/dev/null out=$(pe_adapter "$HOPEN" start open-src) assert_contains "$out" "captured:" "a result from an adapter with no terminal verdict is captured" @@ -529,12 +520,22 @@ pe_adapter "$HOPEN" retire open-src >/dev/null pass "a source stays armed unless its own adapter classifies the result terminal" HREPLACE="$TMP_ROOT/hreplace"; new_home "$HREPLACE" -PE_TRACKED+=("$HREPLACE|replace-src") +fm_test_track_procevent_home "$HREPLACE" OLD_TRIGGER="$TMP_ROOT/replace-old-trigger" -pe_adapter "$HREPLACE" register endnow replace-src -- "$BLOCKER" "$OLD_TRIGGER" "old terminal payload" >/dev/null +OLD_STARTED="$TMP_ROOT/replace-old-started" +REPLACE_BLOCKER="$TMP_ROOT/replace-blocker.sh" +cat > "$REPLACE_BLOCKER" <<'SH' +#!/usr/bin/env bash +printf 'started\n' > "$1" +shift +exec "$@" +SH +chmod +x "$REPLACE_BLOCKER" +pe_adapter "$HREPLACE" register endnow replace-src -- \ + "$REPLACE_BLOCKER" "$OLD_STARTED" "$BLOCKER" "$OLD_TRIGGER" "old terminal payload" >/dev/null pe_adapter "$HREPLACE" start replace-src > "$TMP_ROOT/replace-old.out" 2>&1 & replace_old_pid=$! -wait_for "$FM_PROCEVENT_CLAIM_ROOT/replace-src.claim" || fail "the old registration was never claimed" +wait_for "$OLD_STARTED" || fail "the old registration never started" pe_adapter "$HREPLACE" register openended replace-src -- /bin/echo "replacement payload" >/dev/null touch "$OLD_TRIGGER" wait "$replace_old_pid" || fail "the old terminal runner failed" @@ -552,7 +553,7 @@ pe_adapter "$HREPLACE" retire replace-src >/dev/null pass "terminal retirement preserves and releases a concurrently replaced registration" HRETFAIL="$TMP_ROOT/hretfail"; new_home "$HRETFAIL" -PE_TRACKED+=("$HRETFAIL|retire-fail-src") +fm_test_track_procevent_home "$HRETFAIL" FAIL_RM_BIN=$(fm_fakebin "$TMP_ROOT/retire-fail-bin") REAL_RM=$(command -v rm) export REAL_RM @@ -613,7 +614,7 @@ chmod +x "$LAVISH_BIN/lavish-axi" REVIEW_ART="$TMP_ROOT/review.html" printf '<h1>review</h1>\n' > "$REVIEW_ART" lavish_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$REVIEW_ART") -PE_TRACKED+=("$HLT|$lavish_id") +fm_test_track_procevent_home "$HLT" PATH="$LAVISH_BIN:$PATH" FM_HOME="$HLT" "$ROOT/bin/fm-procevent-lavish.sh" arm "$REVIEW_ART" >/dev/null for _ in $(seq 1 6); do PATH="$LAVISH_BIN:$PATH" pe "$HLT" reconcile >/dev/null @@ -653,7 +654,7 @@ chmod +x "$EMPTY_BIN/lavish-axi" QUIET_ART="$TMP_ROOT/quiet-board.html" printf '<h1>quiet</h1>\n' > "$QUIET_ART" quiet_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$QUIET_ART") -PE_TRACKED+=("$HEMPTY|$quiet_id") +fm_test_track_procevent_home "$HEMPTY" PATH="$EMPTY_BIN:$PATH" FM_HOME="$HEMPTY" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$QUIET_ART" >/dev/null quiet_out=$(PATH="$EMPTY_BIN:$PATH" pe "$HEMPTY" start "$quiet_id" 2>&1) @@ -699,7 +700,7 @@ chmod +x "$ANSWER_BIN/lavish-axi" ANSWER_ART="$TMP_ROOT/answered-board.html" printf '<h1>answered</h1>\n' > "$ANSWER_ART" answer_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$ANSWER_ART") -PE_TRACKED+=("$HANSWER|$answer_id") +fm_test_track_procevent_home "$HANSWER" PATH="$ANSWER_BIN:$PATH" FM_HOME="$HANSWER" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$ANSWER_ART" >/dev/null PATH="$ANSWER_BIN:$PATH" pe "$HANSWER" reconcile >/dev/null @@ -751,9 +752,27 @@ esac SH chmod +x "$LAVISH_SCRIPTED_BIN/lavish-axi" export LAVISH_COUNT LAVISH_SCRIPT + +DEFAULT_RATE_ART="$TMP_ROOT/default-rate-board.html" +printf '<h1>default rate</h1>\n' > "$DEFAULT_RATE_ART" +DEFAULT_RATE_COUNT="$TMP_ROOT/default-rate-count" +PATH="$LAVISH_SCRIPTED_BIN:$PATH" LAVISH_COUNT="$DEFAULT_RATE_COUNT" LAVISH_SCRIPT=interrupt \ + FM_LAVISH_POLL_RETRY_DELAY='' \ + "$ROOT/bin/fm-procevent-lavish.sh" poll "$DEFAULT_RATE_ART" >/dev/null 2>&1 & +DEFAULT_RATE_PID=$! +perl -MTime::HiRes=sleep -e 'sleep 6.2' +kill -TERM "$DEFAULT_RATE_PID" 2>/dev/null || true +wait "$DEFAULT_RATE_PID" 2>/dev/null || true +default_rate_count=$(cat "$DEFAULT_RATE_COUNT" 2>/dev/null || echo 0) +[ "$default_rate_count" -ge 2 ] \ + || fail "the default poll governor stopped an instantly returning source from making progress" +[ "$default_rate_count" -le 2 ] \ + || fail "the shipped poll governor allowed $default_rate_count iterations in 6.2 seconds" +pass "the shipped poll governor bounds an instantly returning source" + # A bounded test override keeps the retry policy's real bound under test without # making the suite wait out the production delay. -export FM_LAVISH_POLL_RETRY_DELAY=0 +export FM_LAVISH_POLL_RETRY_DELAY=1 # Two interruptions, then the captain's real feedback: the retries are silent and # only the feedback becomes a captured result and a check wake. @@ -761,7 +780,7 @@ HRETRY="$TMP_ROOT/hretry"; new_home "$HRETRY" RETRY_ART="$TMP_ROOT/retry-board.html" printf '<h1>retry</h1>\n' > "$RETRY_ART" retry_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$RETRY_ART") -PE_TRACKED+=("$HRETRY|$retry_id") +fm_test_track_procevent_home "$HRETRY" LAVISH_COUNT="$TMP_ROOT/retry-count"; LAVISH_SCRIPT="interrupt interrupt feedback" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HRETRY" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$RETRY_ART" >/dev/null @@ -785,7 +804,7 @@ HEXH="$TMP_ROOT/hexh"; new_home "$HEXH" EXH_ART="$TMP_ROOT/exhaust-board.html" printf '<h1>exhaust</h1>\n' > "$EXH_ART" exh_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$EXH_ART") -PE_TRACKED+=("$HEXH|$exh_id") +fm_test_track_procevent_home "$HEXH" LAVISH_COUNT="$TMP_ROOT/exhaust-count"; LAVISH_SCRIPT="interrupt" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HEXH" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$EXH_ART" >/dev/null @@ -808,7 +827,7 @@ HOTHER="$TMP_ROOT/hother"; new_home "$HOTHER" OTHER_ART="$TMP_ROOT/other-board.html" printf '<h1>other</h1>\n' > "$OTHER_ART" other_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$OTHER_ART") -PE_TRACKED+=("$HOTHER|$other_id") +fm_test_track_procevent_home "$HOTHER" LAVISH_COUNT="$TMP_ROOT/other-count"; LAVISH_SCRIPT="other-server-error" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HOTHER" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$OTHER_ART" >/dev/null @@ -828,9 +847,9 @@ HNEAR="$TMP_ROOT/hnear"; new_home "$HNEAR" NEAR_ART="$TMP_ROOT/near-board.html" printf '<h1>near</h1>\n' > "$NEAR_ART" near_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$NEAR_ART") -PE_TRACKED+=("$HNEAR|$near_id") +fm_test_track_procevent_home "$HNEAR" LAVISH_COUNT="$TMP_ROOT/near-count"; LAVISH_SCRIPT="near-interrupt feedback" -PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" FM_LAVISH_POLL_RETRY_DELAY=0 \ +PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" FM_LAVISH_POLL_RETRY_DELAY=1 \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$NEAR_ART" >/dev/null PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HNEAR" pe "$HNEAR" start "$near_id" >/dev/null [ "$(cat "$LAVISH_COUNT")" = 1 ] \ @@ -847,14 +866,14 @@ HINVALID="$TMP_ROOT/hinvalid"; new_home "$HINVALID" INVALID_ART="$TMP_ROOT/invalid-delay-board.html" printf '<h1>invalid delay</h1>\n' > "$INVALID_ART" invalid_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$INVALID_ART") -for invalid_delay in 61 invalid; do +for invalid_delay in 0 61 invalid; do invalid_status=0 invalid_out=$(PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HINVALID" \ FM_LAVISH_POLL_RETRY_DELAY="$invalid_delay" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$INVALID_ART" 2>&1) || invalid_status=$? [ "$invalid_status" -ne 0 ] \ || fail "arm accepted invalid retry delay: $invalid_delay" - assert_contains "$invalid_out" "must be whole seconds from 0 to 60" \ + assert_contains "$invalid_out" "must be whole seconds from 1 to 60" \ "arm explains the rejected retry delay" assert_absent "$HINVALID/state/procevent/$invalid_id.source" \ "arm publishes no source registration for an invalid retry delay" @@ -880,7 +899,7 @@ LAVISH_STREAM_RELEASE="$TMP_ROOT/stream-release" mkdir -p "$STREAM_TMPDIR" printf '<h1>stream</h1>\n' > "$STREAM_ART" stream_id=$("$ROOT/bin/fm-procevent-lavish.sh" source-id "$STREAM_ART") -PE_TRACKED+=("$HSTREAM|$stream_id") +fm_test_track_procevent_home "$HSTREAM" LAVISH_COUNT="$TMP_ROOT/stream-count"; LAVISH_SCRIPT="stream" PATH="$LAVISH_SCRIPTED_BIN:$PATH" FM_HOME="$HSTREAM" \ "$ROOT/bin/fm-procevent-lavish.sh" arm "$STREAM_ART" >/dev/null @@ -1010,15 +1029,18 @@ pass "retiring a never-completing source stops its runner and its blocked child" # reconcile must also stop a runner whose registration was removed out from under it. TRIG4="$TMP_ROOT/trigger-four" +ORPHAN_READY="$TMP_ROOT/orphan-ready" HZ="$TMP_ROOT/hz"; new_home "$HZ" pe_register "$HZ" lavish orphan-src -- "$BLOCKER" "$TRIG4" "orphan" >/dev/null -pe "$HZ" reconcile >/dev/null +BLOCKER_READY="$ORPHAN_READY" pe "$HZ" reconcile >/dev/null wait_for "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" \ || fail "orphan fixture runner did not publish its claim" orphan_pid=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" 2>/dev/null) if [ -z "$orphan_pid" ] || ! kill -0 "$orphan_pid" 2>/dev/null; then fail "orphan fixture runner did not start" fi +# A claim precedes launch-floor validation; exercise an already-running child. +wait_for "$ORPHAN_READY" || fail "orphan fixture child did not start" rm -f "$HZ/state/procevent/orphan-src.source" out=$(pe "$HZ" reconcile) assert_contains "$out" "stopped=1" "reconcile stops a runner whose registration was removed" @@ -1124,32 +1146,19 @@ kill -0 -"$orphan_leader" 2>/dev/null || fail "fixture invalid: the owned child orphan_out=$(pe "$HG" reconcile) kill -0 -"$orphan_leader" 2>/dev/null \ - && fail "reconcile left the crashed generation's process group alive: $orphan_out" -sleep 0.5 -assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the crashed generation remains alive" -case "$orphan_out" in - *"started=1"*) - # The replacement is detached: it records its own claim and execs its source - # after reconcile has already returned, so both effects must be waited for - # rather than snapshotted behind the settle window above. - wait_for "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" \ - || fail "a replacement runner started without recording its own claim" - wait_for_lines "$ORPHAN_LOG" 2 \ - || fail "the replacement runner never started its source: $(cat "$ORPHAN_LOG")" - [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 2 ] \ - || fail "reconcile did not start exactly one replacement source: $(cat "$ORPHAN_LOG")" - ;; - *"started=0"*) - [ -e "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" ] \ - || fail "refusing to replace must preserve the claim for retry: $orphan_out" - [ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ - || fail "reconcile started a source while refusing replacement: $(cat "$ORPHAN_LOG")" - ;; - *) fail "unexpected reconcile result for a crashed leader: $orphan_out" ;; -esac -: > "$ORPHAN_TRIGGER" + || fail "reconcile signalled an ambiguous leaderless process group: $orphan_out" +assert_contains "$orphan_out" "started=0" \ + "reconcile does not replace an ambiguous leaderless generation" +[ -e "$FM_PROCEVENT_CLAIM_ROOT/orphan-src.claim" ] \ + || fail "refusing ambiguous cleanup must preserve the claim" +[ "$(wc -l < "$ORPHAN_LOG" | tr -d ' ')" = 1 ] \ + || fail "reconcile started a source beside an ambiguous leaderless group" +assert_absent "$ORPHAN_OVERLAP" "no replacement source starts while the leaderless group remains" +kill -KILL -"$orphan_leader" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$orphan_leader" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$orphan_leader" 2>/dev/null && fail "could not clean up the leaderless fixture group" pe "$HG" retire orphan-src >/dev/null -pass "a crashed runner leader never lets a live owned group be reclaimed as stale" +pass "an ambiguous leaderless group is preserved without replacement" # Counterexample: a genuinely dead generation - no leader and no surviving # group - must still be reclaimable, or crash recovery would deadlock. @@ -1167,6 +1176,141 @@ wait_for "$DEAD_LOG" || fail "the replacement source never started for a truly d pe "$HG2" retire dead-gen-src >/dev/null pass "a truly dead generation with no surviving group is still safely reclaimed" +# --- a dead generation stays reclaimable when the state root cannot be +# revalidated ----------------------------------------------------------------- +# The reported wedge. Reclaiming a dead generation ran the claim's +# capture-reservation cleanup first, and that cleanup re-verifies the recorded +# state-root identity. Once that identity stopped matching, a claim naming a pid +# and a process group that were both provably gone could not be cleared: +# reconcile kept reporting a start while nothing ever attached, and retire +# refused with "cannot release source ownership". Reservation records are keyed +# by claim token and a replacement always claims a fresh one, so they can never +# collide with the generation that replaces them - they are hygiene, not an +# ownership invariant, and they must not veto an ownership move that the +# documented promise already grants. +# +# The claim below is the modern shape (it carries the state-root identity block +# a legacy claim does not have), which is why the existing dead-generation case +# above never reached this path. +HSR="$TMP_ROOT/hsr"; new_home "$HSR" +SR_TRIGGER="$TMP_ROOT/state-root-trigger" +SR_LOG="$TMP_ROOT/state-root-executions" +pe_register "$HSR" lavish state-root-src -- "$RACE_BLOCKER" "$SR_LOG" "$SR_TRIGGER" >/dev/null +pe "$HSR" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim" \ + || fail "state-root fixture never claimed its source" +wait_for "$SR_LOG" || fail "state-root fixture source never started" +sr_leader=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim") +case "$sr_leader" in ''|*[!0-9]*) fail "could not read the state-root fixture leader pid" ;; esac +[ -n "$(sed -n '8p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim")" ] \ + || fail "fixture invalid: the claim carries no state-root identity to invalidate" + +kill -KILL -"$sr_leader" 2>/dev/null || true +kill -KILL "$sr_leader" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$sr_leader" 2>/dev/null || break; sleep 0.1; done +kill -0 "$sr_leader" 2>/dev/null && fail "the state-root fixture leader survived SIGKILL" +kill -0 -"$sr_leader" 2>/dev/null && fail "fixture invalid: the owned group outlived the whole generation" +# Drift the live state root away from what the claim recorded. +chmod 750 "$HSR/state" || fail "could not drift the state-root identity" + +sr_out=$(pe "$HSR" reconcile) +assert_contains "$sr_out" "started=1" "a dead generation was not reclaimed after the state root drifted: $sr_out" +# Reporting a start is not the same fact as listening: the previous behavior +# reported exactly this while the replacement silently failed to claim. +wait_for_lines "$SR_LOG" 2 \ + || fail "reconcile reported a start but no replacement source ever ran: $(cat "$SR_LOG")" +sr_new=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/state-root-src.claim") +[ "$sr_new" != "$sr_leader" ] || fail "the dead generation's claim was never replaced" +kill -0 "$sr_new" 2>/dev/null || fail "the replacement runner did not take ownership" +: > "$SR_TRIGGER" +pe "$HSR" retire state-root-src >/dev/null +pass "reconcile reclaims a dead generation whose state-root identity no longer matches" + +# Retire must release the same wedged claim rather than refusing forever. +HSR2="$TMP_ROOT/hsr2"; new_home "$HSR2" +pe_register "$HSR2" lavish wedged-src -- /bin/echo recovered >/dev/null +sr2_identity=$(bash -c '. "$1/bin/fm-pr-lib.sh"; fm_pr_file_identity "$2"' _ \ + "$ROOT" "$HSR2/state/procevent/wedged-src.source") \ + || fail "could not read the wedged fixture registration identity" +{ + printf '%s\n%s\nwedged-token\nwedged-identity\n' "$HSR2" 999999 + printf '%s\n%s\nactive\n' "$HSR2/state/procevent" "$sr2_identity" + # A state-root identity that names the right directory with the wrong inode: + # exactly what a claim recorded before its home was re-created looks like. + printf '%s\n%s\n%s\n%s\n%s\n' "$HSR2/state" 1 1 "$(id -u)" 755 +} > "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" +chmod 0600 "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" +wedged_out=$(pe "$HSR2" retire wedged-src 2>&1) \ + || fail "retire refused to release a claim whose whole generation is gone: $wedged_out" +assert_contains "$wedged_out" "retired: wedged-src" "retire did not report releasing the wedged source: $wedged_out" +assert_absent "$FM_PROCEVENT_CLAIM_ROOT/wedged-src.claim" "retire left the dead generation owning the source" +assert_absent "$HSR2/state/procevent/wedged-src.source" "retire left the wedged source registered" +pass "retire releases a dead generation's claim instead of refusing forever" + +# The guard is not weakened in the other direction: the same unrevalidatable +# state root must NOT let anything take a source away from a live generation. +HSR3="$TMP_ROOT/hsr3"; new_home "$HSR3" +SR3_TRIGGER="$TMP_ROOT/state-root-live-trigger" +SR3_LOG="$TMP_ROOT/state-root-live-executions" +pe_register "$HSR3" lavish live-drift-src -- "$RACE_BLOCKER" "$SR3_LOG" "$SR3_TRIGGER" >/dev/null +pe "$HSR3" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim" \ + || fail "live-drift fixture never claimed its source" +wait_for "$SR3_LOG" || fail "live-drift fixture source never started" +sr3_leader=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim") +chmod 750 "$HSR3/state" || fail "could not drift the live owner's state-root identity" +sr3_out=$(pe "$HSR3" start live-drift-src) +assert_contains "$sr3_out" "already owned" "a live generation was displaced after its state root drifted: $sr3_out" +[ "$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/live-drift-src.claim")" = "$sr3_leader" ] \ + || fail "the live generation's claim was replaced" +kill -0 "$sr3_leader" 2>/dev/null || fail "the live owner was killed by a reclaim attempt" +sr3_reconcile=$(pe "$HSR3" reconcile) +assert_contains "$sr3_reconcile" "started=0" "reconcile started a second poller beside a live owner: $sr3_reconcile" +[ "$(wc -l < "$SR3_LOG" | tr -d ' ')" = 1 ] \ + || fail "a second source ran beside the live owner: $(cat "$SR3_LOG")" +: > "$SR3_TRIGGER" +pe "$HSR3" retire live-drift-src >/dev/null +pass "a live generation is never reclaimed, drifted state root or not" + +HSR4="$TMP_ROOT/hsr4"; new_home "$HSR4" +SR4_TRIGGER="$TMP_ROOT/state-root-reused-trigger" +SR4_LOG="$TMP_ROOT/state-root-reused-executions" +pe_register "$HSR4" lavish reused-group-src -- "$RACE_BLOCKER" "$SR4_LOG" "$SR4_TRIGGER" >/dev/null +pe "$HSR4" reconcile >/dev/null +wait_for "$FM_PROCEVENT_CLAIM_ROOT/reused-group-src.claim" \ + || fail "reused-group fixture never claimed its source" +wait_for "$SR4_LOG" || fail "reused-group fixture source never started" +sr4_claim="$FM_PROCEVENT_CLAIM_ROOT/reused-group-src.claim" +sr4_leader=$(sed -n '2p' "$sr4_claim") +sr4_identity=$(sed -n '4p' "$sr4_claim") +kill -0 -"$sr4_leader" 2>/dev/null \ + || fail "fixture invalid: the reused-pid process group is not alive" +awk 'NR == 4 { print "different-live-process-identity"; next } { print }' \ + "$sr4_claim" > "$sr4_claim.tmp" && mv "$sr4_claim.tmp" "$sr4_claim" +chmod 0600 "$sr4_claim" +chmod 750 "$HSR4/state" || fail "could not drift the reused-group state root" +sr4_out=$(pe "$HSR4" reconcile) +sleep 0.5 +[ "$(wc -l < "$SR4_LOG" | tr -d ' ')" = 1 ] \ + || fail "reconcile started a replacement beside a reused pid's live group: $sr4_out" +[ "$(sed -n '2p' "$sr4_claim")" = "$sr4_leader" ] \ + || fail "reconcile replaced the reused-pid generation's claim" +set +e +sr4_retire=$(pe "$HSR4" retire reused-group-src 2>&1) +sr4_rc=$? +set -e +[ "$sr4_rc" -ne 0 ] || fail "retire released a claim whose process group survives: $sr4_retire" +[ -e "$sr4_claim" ] || fail "retire removed the reused-pid generation's claim" +[ -e "$HSR4/state/procevent/reused-group-src.source" ] \ + || fail "retire removed the reused-pid generation's registration" +awk -v identity="$sr4_identity" 'NR == 4 { print identity; next } { print }' \ + "$sr4_claim" > "$sr4_claim.tmp" && mv "$sr4_claim.tmp" "$sr4_claim" +chmod 0600 "$sr4_claim" +chmod 755 "$HSR4/state" +: > "$SR4_TRIGGER" +pe "$HSR4" retire reused-group-src >/dev/null +pass "a reused pid never makes its surviving process group reclaimable" + HJ="$TMP_ROOT/hj"; new_home "$HJ" TORN_TRIGGER="$TMP_ROOT/torn-trigger" pe_register "$HJ" lavish torn-src -- "$BLOCKER" "$TORN_TRIGGER" "torn" >/dev/null @@ -1231,7 +1375,7 @@ kill -0 "$innocent_pid" 2>/dev/null || fail "retirement signaled a PID whose ide kill "$innocent_pid" 2>/dev/null || true wait "$innocent_pid" 2>/dev/null || true assert_absent "$FM_PROCEVENT_CLAIM_ROOT/reused-src.claim" "retirement releases the exact reused-pid claim" -pass "PID reuse cannot signal an unrelated process" +pass "detected PID reuse is refused before signalling" HL="$TMP_ROOT/hl"; new_home "$HL" IDENTITY_TRIGGER="$TMP_ROOT/identity-trigger" @@ -1266,6 +1410,19 @@ wait_for "$FM_PROCEVENT_CLAIM_ROOT/sweep-one.claim" || fail "home sweep fixture wait_for "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim" || fail "home sweep fixture two did not start" sweep_pid_one=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/sweep-one.claim") sweep_pid_two=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim") +# The claim-only case is an owned claim with no live runner, so build exactly +# that: kill the runner's group so it cannot run its own cleanup, confirm it is +# gone, and only then drop the registration. Deleting the registration out from +# under a LIVE runner no longer produces this case, because a superseded +# generation now observes the identity mismatch, self-retires, and releases its +# claim - so the sweep would race that exit and see one source or two depending +# on which won. +kill -KILL -"$sweep_pid_two" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 "$sweep_pid_two" 2>/dev/null || break; sleep 0.1; done +kill -0 "$sweep_pid_two" 2>/dev/null \ + && fail "the claim-only sweep fixture runner did not stop" +assert_present "$FM_PROCEVENT_CLAIM_ROOT/sweep-two.claim" \ + "a killed runner leaves its owned claim behind for the sweep" rm -f "$HM/state/procevent/sweep-two.source" out=$(pe "$HM" sweep-home --preflight) assert_contains "$out" "sweep preflight: ready" "home sweep preflight validates the full bounded snapshot" @@ -1393,6 +1550,64 @@ kill -0 "$noisy_child" 2>/dev/null && fail "TERM-resistant source child survived assert_absent "$staged" "retirement removes the tracked partial staging file" pass "live output stays bounded and retirement reaps the whole source group" +HPOST_TERM="$TMP_ROOT/post-term-reuse"; new_home "$HPOST_TERM" +POST_TERM_SOURCE="$TMP_ROOT/post-term-reuse-source.sh" +POST_TERM_PID="$TMP_ROOT/post-term-reuse.pid" +POST_TERM_MARKER="$TMP_ROOT/post-term-reuse.marker" +POST_TERM_COUNT="$TMP_ROOT/post-term-reuse.count" +cat > "$POST_TERM_SOURCE" <<'SH' +#!/usr/bin/env bash +trap '' TERM +printf '%s\n' "$$" > "$1" +while :; do sleep 1; done +SH +chmod +x "$POST_TERM_SOURCE" +POST_TERM_BIN=$(fm_fakebin "$TMP_ROOT/post-term-reuse-bin") +REAL_PS=$(command -v ps) || fail "the post-TERM reuse fixture requires ps" +cat > "$POST_TERM_BIN/ps" <<SH +#!/usr/bin/env bash +if [ -e "$POST_TERM_MARKER" ] && [ "\${1-}" = -p ] \ + && [ "\${3-}" = -o ] && [ "\${4-}" = lstart= ]; then + count=0 + [ ! -f "$POST_TERM_COUNT" ] || count=\$(cat "$POST_TERM_COUNT") + count=\$((count + 1)) + printf '%s\n' "\$count" > "$POST_TERM_COUNT" + if [ "\$count" -gt 1 ]; then + printf 'post-TERM reused identity\n' + exit 0 + fi +fi +exec "$REAL_PS" "\$@" +SH +chmod +x "$POST_TERM_BIN/ps" +pe_register "$HPOST_TERM" lavish post-term-src -- \ + "$POST_TERM_SOURCE" "$POST_TERM_PID" >/dev/null +FM_PROCEVENT_OWNER_CHECK_SECONDS=5 pe "$HPOST_TERM" reconcile >/dev/null +wait_for "$POST_TERM_PID" || fail "the post-TERM reuse fixture did not start" +wait_for "$FM_PROCEVENT_CLAIM_ROOT/post-term-src.claim" \ + || fail "the post-TERM reuse fixture did not claim its source" +POST_TERM_RUNNER=$(sed -n '2p' "$FM_PROCEVENT_CLAIM_ROOT/post-term-src.claim") +touch "$POST_TERM_MARKER" +post_term_status=0 +PATH="$POST_TERM_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-post-term-proc" \ + pe "$HPOST_TERM" retire post-term-src >/dev/null 2>&1 || post_term_status=$? +[ "$post_term_status" -ne 0 ] || fail "retirement escalated after runner identity became ambiguous" +# Which ambiguity the escalation meets here is platform-dependent, so this +# asserts the invariant both forms share rather than one form's internals. +# Where the runner leader keeps waiting on its TERM-ignoring source child the +# post-TERM check sees a live leader whose identity no longer matches, and +# where the leader dies promptly it sees a leaderless group carrying the same +# numeric id; fm_procevent_pid_state reaches the second verdict without +# consulting process identity at all, so counting identity lookups pins a +# timing- and platform-dependent internal rather than the behavior. +kill -0 -"$POST_TERM_RUNNER" 2>/dev/null \ + || fail "an ambiguous reused-PID group was killed during escalation" +kill -KILL -"$POST_TERM_RUNNER" 2>/dev/null || true +for _ in $(seq 1 50); do kill -0 -"$POST_TERM_RUNNER" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$POST_TERM_RUNNER" 2>/dev/null && fail "could not clean up the post-TERM fixture group" +pe "$HPOST_TERM" retire post-term-src >/dev/null +pass "cleanup aborts escalation after runner identity becomes ambiguous" + HBAD="$TMP_ROOT/hbad"; new_home "$HBAD" pe_register "$HBAD" lavish bad-limit -- /bin/true >/dev/null bad_limit_status=0 @@ -1784,4 +1999,531 @@ assert_not_contains "$runner_help" "exactly-once" \ "the runner's help claims no exactly-once delivery" pass "the published interfaces state the loss limitation and claim no lossless delivery" +# --- launch pacing and guard startup ---------------------------------------- + +FAST_SOURCE="$TMP_ROOT/fast-source.sh" +cat > "$FAST_SOURCE" <<'SH' +#!/usr/bin/env bash +perl -MTime::HiRes=time -e 'printf "%.6f\n", time' >> "$1" +exit 1 +SH +chmod +x "$FAST_SOURCE" + +STORM_SOURCE="$TMP_ROOT/storm-source.sh" +cat > "$STORM_SOURCE" <<'SH' +#!/usr/bin/env bash +perl -MTime::HiRes=time -e 'printf "%.6f\n", time' >> "$1" +FM_HOME="$2" perl -MPOSIX=setsid -e ' + my @command = @ARGV; + defined(my $pid = fork) or exit 1; + exit 0 if $pid; + setsid() >= 0 or exit 1; + open STDIN, "<", "/dev/null" or exit 1; + open STDOUT, ">", "/dev/null" or exit 1; + open STDERR, ">", "/dev/null" or exit 1; + select undef, undef, undef, 0.2; + exec @command; +' "$3/bin/fm-procevent.sh" reconcile +exit 1 +SH +chmod +x "$STORM_SOURCE" + +HFLOOR="$TMP_ROOT/launch-floor"; new_home "$HFLOOR" +fm_test_track_procevent_home "$HFLOOR" +pe_register "$HFLOOR" lavish floor-src -- \ + "$STORM_SOURCE" "$TMP_ROOT/launch-times" "$HFLOOR" "$ROOT" +FM_PROCEVENT_OWNER_LEASE_SECONDS=4 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HFLOOR" reconcile >/dev/null +floor_deadline=$((SECONDS + 12)) +while :; do + floor_count=0 + [ ! -f "$TMP_ROOT/launch-times" ] \ + || floor_count=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + [ "$floor_count" -ge 3 ] && break + [ "$SECONDS" -lt "$floor_deadline" ] \ + || fail "the orphan-storm fixture did not relaunch its source command" + sleep 0.1 +done +launch_count=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') +launch_span=$(perl -e '@t=<>; printf "%.3f", $t[-1] - $t[0]' "$TMP_ROOT/launch-times") +perl -e 'exit($ARGV[0] >= ($ARGV[1] - 1) * 0.8 ? 0 : 1)' "$launch_span" "$launch_count" \ + || fail "an orphaned source launched $launch_count times in only ${launch_span}s" +[ "$launch_count" -le 6 ] \ + || fail "an orphaned source stormed $launch_count launches during its owner-dead grace window" +pass "an orphaned source command obeys the launch floor during its grace window" + +HPACE="$TMP_ROOT/registration-pacing"; new_home "$HPACE" +fm_test_track_procevent_home "$HPACE" +PACE_LOG="$TMP_ROOT/registration-pacing.log" +pe_register "$HPACE" lavish pace-src -- "$FAST_SOURCE" "$PACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HPACE" start pace-src >/dev/null +pe "$HPACE" retire pace-src >/dev/null +pe_register "$HPACE" lavish pace-src -- "$FAST_SOURCE" "$PACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HPACE" start pace-src > "$TMP_ROOT/replacement-pacing.out" 2>&1 & +PACE_START_PID=$! +pace_deadline=$((SECONDS + 4)) +while kill -0 "$PACE_START_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$pace_deadline" ]; then + pe "$HPACE" retire pace-src >/dev/null 2>&1 || true + wait "$PACE_START_PID" 2>/dev/null || true + fail "a replacement registration inherited the prior launch floor" + fi + sleep 0.1 +done +wait "$PACE_START_PID" || fail "the replacement registration failed" +[ "$(wc -l < "$PACE_LOG" | tr -d ' ')" = 2 ] \ + || fail "a replacement registration did not launch immediately" +PACE_STAMPS=$(find "$HPACE/state/procevent" -maxdepth 1 -type f \ + -name 'pace-src.*.last-launch' | wc -l | tr -d ' ') +[ "$PACE_STAMPS" = 1 ] || fail "replacement registrations accumulated stale pacing state" +pass "a replacement registration starts with one fresh launch floor" + +HPACE_RACE="$TMP_ROOT/registration-pacing-race"; new_home "$HPACE_RACE" +fm_test_track_procevent_home "$HPACE_RACE" +PACE_RACE_LOG="$TMP_ROOT/registration-pacing-race.log" +pe_register "$HPACE_RACE" lavish pace-race-src -- "$FAST_SOURCE" "$PACE_RACE_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 pe "$HPACE_RACE" start pace-race-src >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 \ + pe "$HPACE_RACE" start pace-race-src > "$TMP_ROOT/registration-pacing-race.out" 2>&1 & +PACE_RACE_PID=$! +wait_for "$FM_PROCEVENT_CLAIM_ROOT/pace-race-src.claim" \ + || fail "the superseded pacing fixture did not claim its registration" +[ "$(wc -l < "$PACE_RACE_LOG" | tr -d ' ')" = 1 ] \ + || fail "the superseded pacing fixture was not waiting on its launch floor" +pe_register "$HPACE_RACE" lavish pace-race-src -- "$FAST_SOURCE" "$PACE_RACE_LOG" >/dev/null +wait "$PACE_RACE_PID" || fail "the superseded paced runner failed" +[ "$(wc -l < "$PACE_RACE_LOG" | tr -d ' ')" = 1 ] \ + || fail "the superseded paced runner invoked its stale command" +# The runner marker is written before the launch floor is waited on, and a home +# sweep counts a marker with no owned claim as a preflight failure. A superseded +# generation that exits without clearing its marker therefore makes the whole +# home refuse to sweep, so assert the marker is gone and the sweep still runs. +assert_absent "$HPACE_RACE/state/procevent/pace-race-src.runner" \ + "a superseded paced runner leaves no runner marker behind" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3 pe "$HPACE_RACE" start pace-race-src >/dev/null +PACE_RACE_STAMPS=$(find "$HPACE_RACE/state/procevent" -maxdepth 1 -type f \ + -name 'pace-race-src.*.last-launch' | wc -l | tr -d ' ') +[ "$PACE_RACE_STAMPS" = 1 ] \ + || fail "a superseded sleeping runner recreated stale pacing state" +pass "a superseded sleeping runner cannot recreate stale pacing state" + +HCOMMIT="$TMP_ROOT/registration-commit"; new_home "$HCOMMIT" +fm_test_track_procevent_home "$HCOMMIT" +COMMIT_LOG="$TMP_ROOT/registration-commit.log" +mkdir -p "$HCOMMIT/state/procevent/commit-src.1-2.last-launch" +pe_register "$HCOMMIT" lavish commit-src -- "$FAST_SOURCE" "$COMMIT_LOG" >/dev/null \ + || fail "post-commit pacing cleanup made registration report failure" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HCOMMIT" start commit-src >/dev/null \ + || fail "a successfully published registration was not executable" +[ "$(wc -l < "$COMMIT_LOG" | tr -d ' ')" = 1 ] \ + || fail "the committed registration did not invoke its source" +pass "post-commit pacing cleanup cannot veto registration publication" + +HROLLBACK="$TMP_ROOT/rollback-pacing"; new_home "$HROLLBACK" +fm_test_track_procevent_home "$HROLLBACK" +ROLLBACK_LOG="$TMP_ROOT/rollback-pacing.log" +pe_register "$HROLLBACK" lavish rollback-src -- "$FAST_SOURCE" "$ROLLBACK_LOG" >/dev/null +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=1 pe "$HROLLBACK" start rollback-src >/dev/null +ROLLBACK_STAMP= +for candidate in "$HROLLBACK/state/procevent"/rollback-src.*.last-launch; do + [ -f "$candidate" ] && ROLLBACK_STAMP=$candidate +done +[ -n "$ROLLBACK_STAMP" ] || fail "the first launch did not persist its pacing state" +printf '%s\n' "$(( $(date +%s) + 3600 ))" > "$ROLLBACK_STAMP" +FM_PROCEVENT_LAUNCH_FLOOR_SECONDS=3600 pe "$HROLLBACK" start rollback-src > "$TMP_ROOT/rollback.out" 2>&1 & +ROLLBACK_START_PID=$! +rollback_deadline=$((SECONDS + 4)) +while kill -0 "$ROLLBACK_START_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$rollback_deadline" ]; then + pe "$HROLLBACK" retire rollback-src >/dev/null 2>&1 || true + wait "$ROLLBACK_START_PID" 2>/dev/null || true + fail "a pre-reboot monotonic stamp delayed the first launch" + fi + sleep 0.1 +done +wait "$ROLLBACK_START_PID" || fail "the rollback-paced source failed" +[ "$(wc -l < "$ROLLBACK_LOG" | tr -d ' ')" = 2 ] \ + || fail "the rollback-paced source did not invoke twice" +pass "a pre-reboot monotonic stamp is treated as expired" + +storm_deadline=$((SECONDS + 15)) +while :; do + storm_before=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + sleep 2 + storm_after=$(wc -l < "$TMP_ROOT/launch-times" | tr -d ' ') + [ "$storm_before" = "$storm_after" ] && break + [ "$SECONDS" -lt "$storm_deadline" ] \ + || fail "an orphaned self-relaunching source survived its expired owner lease" +done +pass "an expired owner lease stops a self-relaunching source generation" + +HRECREATED="$TMP_ROOT/recreated-owner"; new_home "$HRECREATED" +fm_test_track_procevent_home "$HRECREATED" +RECREATED_TRIGGER="$TMP_ROOT/recreated-owner.trigger" +pe_register "$HRECREATED" lavish recreated-src -- \ + "$BLOCKER" "$RECREATED_TRIGGER" "recreated payload" >/dev/null +FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HRECREATED" reconcile >/dev/null +wait_for "$HRECREATED/state/procevent/recreated-src.runner" \ + || fail "the recreated-path fixture never launched its source" +RECREATED_RUNNER_PID=$(cat "$HRECREATED/state/procevent/recreated-src.runner") +mv "$HRECREATED/state" "$TMP_ROOT/recreated-owner-old-state" +pe_register "$HRECREATED" lavish recreated-src -- \ + "$BLOCKER" "$RECREATED_TRIGGER" "replacement payload" >/dev/null +FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HRECREATED" reconcile >/dev/null +recreated_deadline=$((SECONDS + 8)) +while kill -0 "$RECREATED_RUNNER_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$recreated_deadline" ] \ + || fail "a fresh lease at a recreated state path preserved the old runner" + sleep 0.1 +done +pass "a recreated state path does not preserve the old runner" + +HGUARDFAIL="$TMP_ROOT/guard-failure"; new_home "$HGUARDFAIL" +fm_test_track_procevent_home "$HGUARDFAIL" +pe_register "$HGUARDFAIL" lavish guard-fail-src -- "$FAST_SOURCE" "$TMP_ROOT/unguarded-launches" +guard_fail_status=0 +guard_fail_out=$(FM_PROCEVENT_OWNER_LEASE_SECONDS=invalid \ + pe "$HGUARDFAIL" start guard-fail-src 2>&1) || guard_fail_status=$? +[ "$guard_fail_status" -ne 0 ] || fail "a runner continued after its owner guard failed to initialize" +assert_contains "$guard_fail_out" "cannot start the runner's owner guard" \ + "guard initialization failure is reported at the runner boundary" +assert_absent "$TMP_ROOT/unguarded-launches" \ + "a source command ran without a successfully initialized owner guard" +pass "a runner fails closed when its owner guard cannot initialize" + +HATTACHED="$TMP_ROOT/attached-owner"; new_home "$HATTACHED" +fm_test_track_procevent_home "$HATTACHED" +ATTACHED_TRIGGER="$TMP_ROOT/attached.trigger" +pe_register "$HATTACHED" lavish attached-src -- "$BLOCKER" "$ATTACHED_TRIGGER" "attached payload" +FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HATTACHED" start attached-src > "$TMP_ROOT/attached.out" 2>&1 & +ATTACHED_START_PID=$! +wait_for "$HATTACHED/state/procevent/attached-src.runner" \ + || fail "the attached start never launched its source" +sleep 4 +kill -0 "$ATTACHED_START_PID" 2>/dev/null \ + || fail "a foreground start lost its owner lease while its caller remained attached" +touch "$ATTACHED_TRIGGER" +wait "$ATTACHED_START_PID" || fail "the attached start did not complete after its source returned" +assert_contains "$(cat "$TMP_ROOT/attached.out")" "captured:" \ + "the attached source result was not captured" +pass "a foreground start refreshes its lease while its caller remains attached" + +HCLOCK="$TMP_ROOT/lease-clock"; new_home "$HCLOCK" +fm_test_track_procevent_home "$HCLOCK" +CLOCK_TRIGGER="$TMP_ROOT/lease-clock.trigger" +CLOCK_STATE="$TMP_ROOT/lease-clock-state" +CLOCK_BIN=$(fm_fakebin "$TMP_ROOT/lease-clock-bin") +REAL_DATE=$(command -v date) || fail "the lease clock fixture requires date" +cat > "$CLOCK_BIN/date" <<SH +#!/usr/bin/env bash +if [ "\${1-}" = +%s ]; then + while ! mkdir "$CLOCK_STATE.lock" 2>/dev/null; do sleep 0.01; done + value=0 + [ ! -f "$CLOCK_STATE" ] || value=\$(cat "$CLOCK_STATE") + value=\$((value + 10000)) + printf '%s\n' "\$value" > "$CLOCK_STATE" + rmdir "$CLOCK_STATE.lock" + printf '%s\n' "\$value" + exit 0 +fi +exec "$REAL_DATE" "\$@" +SH +chmod +x "$CLOCK_BIN/date" +pe_register "$HCLOCK" lavish lease-clock-src -- \ + "$BLOCKER" "$CLOCK_TRIGGER" "clock payload" >/dev/null +PATH="$CLOCK_BIN:$PATH" FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HCLOCK" start lease-clock-src > "$TMP_ROOT/lease-clock.out" 2>&1 & +CLOCK_START_PID=$! +wait_for "$HCLOCK/state/procevent/lease-clock-src.runner" \ + || fail "the clock-shift fixture never launched its source" +sleep 4 +kill -0 "$CLOCK_START_PID" 2>/dev/null \ + || fail "wall-clock corrections expired a live foreground owner" +touch "$CLOCK_TRIGGER" +wait "$CLOCK_START_PID" || fail "the clock-shift fixture did not complete" +pass "wall-clock corrections do not alter owner lease age" + +HDETACHED="$TMP_ROOT/detached-attached-owner"; new_home "$HDETACHED" +fm_test_track_procevent_home "$HDETACHED" +DETACHED_TRIGGER="$TMP_ROOT/detached-attached.trigger" +pe_register "$HDETACHED" lavish detached-attached-src -- \ + "$BLOCKER" "$DETACHED_TRIGGER" "detached attached payload" +FM_PROCEVENT_OWNER_LEASE_SECONDS=1 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 FM_HOME="$HDETACHED" \ + perl -MPOSIX=setsid -e 'setsid() >= 0 or exit 1; exec @ARGV' \ + "$ROOT/bin/fm-procevent.sh" start detached-attached-src \ + > "$TMP_ROOT/detached-attached.out" 2>&1 & +DETACHED_START_PID=$! +wait_for "$HDETACHED/state/procevent/detached-attached-src.runner" \ + || fail "the detachable foreground start never launched its source" +DETACHED_RUNNER_PID=$(cat "$HDETACHED/state/procevent/detached-attached-src.runner") +kill "$DETACHED_START_PID" +wait "$DETACHED_START_PID" 2>/dev/null || true +detached_deadline=$((SECONDS + 8)) +while kill -0 "$DETACHED_RUNNER_PID" 2>/dev/null; do + if [ "$SECONDS" -ge "$detached_deadline" ]; then + pe "$HDETACHED" retire detached-attached-src >/dev/null 2>&1 || true + fail "an orphaned attached-start keeper preserved its owner's lease" + fi + sleep 0.1 +done +pass "an attached-start keeper stops refreshing after its parent exits" + +HREUSED_GROUP="$TMP_ROOT/reused-runner-group"; new_home "$HREUSED_GROUP" +fm_test_track_procevent_home "$HREUSED_GROUP" +REUSED_GROUP_MARKER="$TMP_ROOT/reused-runner-group.marker" +REUSED_GROUP_TRIGGER="$TMP_ROOT/reused-runner-group.trigger" +REUSED_GROUP_BIN=$(fm_fakebin "$TMP_ROOT/reused-runner-group-bin") +REAL_PS=$(command -v ps) || fail "the reused-group fixture requires ps" +cat > "$REUSED_GROUP_BIN/ps" <<SH +#!/usr/bin/env bash +if [ -e "$REUSED_GROUP_MARKER" ] && [ "\${1-}" = -p ] \ + && [ "\${3-}" = -o ] && [ "\${4-}" = lstart= ]; then + printf 'reused runner identity\n' + exit 0 +fi +exec "$REAL_PS" "\$@" +SH +chmod +x "$REUSED_GROUP_BIN/ps" +pe_register "$HREUSED_GROUP" lavish reused-runner-group-src -- \ + "$BLOCKER" "$REUSED_GROUP_TRIGGER" "reused group payload" >/dev/null +PATH="$REUSED_GROUP_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-reused-group-proc" \ + FM_PROCEVENT_OWNER_LEASE_SECONDS=30 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + pe "$HREUSED_GROUP" reconcile >/dev/null +wait_for "$HREUSED_GROUP/state/procevent/reused-runner-group-src.runner" \ + || fail "the reused-group fixture runner did not start" +REUSED_GROUP_RUNNER=$(cat "$HREUSED_GROUP/state/procevent/reused-runner-group-src.runner") +touch "$REUSED_GROUP_MARKER" +sleep 3 +kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null \ + || fail "the guard killed a process group after its runner identity became ambiguous" +# Retiring here must read identity from the source this runner was recorded +# under, so the override stays in place: without it the read falls back to +# /proc where that exists, which is a different source than the recorded ps +# identity, and the guard would refuse this retirement on Linux while accepting +# it on macOS. Clearing the marker restores the matching identity, so this also +# asserts the complementary guarantee - once the ambiguity is gone, retirement +# reaps the whole group rather than leaving it behind. +rm -f "$REUSED_GROUP_MARKER" +PATH="$REUSED_GROUP_BIN:$PATH" FM_PROC_ROOT_OVERRIDE="$TMP_ROOT/no-reused-group-proc" \ + pe "$HREUSED_GROUP" retire reused-runner-group-src >/dev/null +for _ in $(seq 1 50); do kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null || break; sleep 0.1; done +kill -0 -"$REUSED_GROUP_RUNNER" 2>/dev/null \ + && fail "retirement left the group alive once runner identity was unambiguous" +pass "a detected ambiguous reused-PID group is not signalled" + +# --- an accidentally orphaned runner is bounded by its owner ---------------- +# +# Reproduces the shape that wedged a host: a listener detached into its own +# process group, reparented to init when its session ended, and left running for +# a day with its blocking child - and everything that child spawned - still +# executing. The cost was not the runner itself but the process churn under it, +# which is why this asserts the whole descendant tree stops, not just the leader. +# +# Scope is asserted alongside it, in the same run and against the same stub: a +# home whose session is still there keeps its runner. Reaping that keyed on the +# script or process name instead of the owning session would take both. + +ORPHAN_STUB="$TMP_ROOT/orphan-stub.sh" +cat > "$ORPHAN_STUB" <<'SH' +#!/usr/bin/env bash +# A blocking source whose child keeps spawning processes, which is what a poll +# stub waiting on a trigger file actually does. The spawn rate is what turned a +# leftover listener into a host-wide storm, so the tick log is the evidence that +# the storm stopped and not merely that one pid went away. +marker=$1 +( while [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ]; do + printf 'tick\n' >> "$marker.ticks" + sleep 0.1 + done ) & +printf '%s\n' "$!" > "$marker.descendant" +while [ ! -e "$marker.trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.1 +done +printf 'orphan payload\n' +SH +chmod +x "$ORPHAN_STUB" + +# The same shape without the spawn churn, for the home that exercises explicit +# retirement rather than the storm. Retirement refuses instead of signalling +# when it cannot confirm the runner's identity, that identity is read through +# `ps`, and the churning stub above starves that read often enough to make a +# single retirement attempt a race. The storm itself is already covered against +# the churning stub by the owner-loss reaping above, which asserts the tick log +# stops, so this home only needs a reparented listener holding a real +# descendant in its group. +QUIET_STUB="$TMP_ROOT/quiet-stub.sh" +cat > "$QUIET_STUB" <<'SH' +#!/usr/bin/env bash +marker=$1 +( sleep "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ) & +printf '%s\n' "$!" > "$marker.descendant" +while [ ! -e "$marker.trigger" ]; do + [ "$SECONDS" -lt "${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120}" ] || exit 75 + sleep 0.1 +done +printf 'orphan payload\n' +SH +chmod +x "$QUIET_STUB" + +# Short enough to observe, and driven through the same environment a real home +# uses, so the bound under test is the shipped one rather than a test-only path. +orphan_pe() { # <home> <command...> + local home=$1 + shift + FM_PROCEVENT_OWNER_LEASE_SECONDS=2 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_HOME="$home" "$ROOT/bin/fm-procevent.sh" "$@" +} + +wait_gone() { # <pid-or-group-spec> [tries] + local spec=$1 n=${2:-160} + for _ in $(seq 1 "$n"); do + kill -0 "$spec" 2>/dev/null || return 0 + sleep 0.1 + done + return 1 +} + +HORPHAN="$TMP_ROOT/orphan-dead-owner"; new_home "$HORPHAN" +fm_test_track_procevent_home "$HORPHAN" +HKEEP="$TMP_ROOT/orphan-live-owner"; new_home "$HKEEP" +fm_test_track_procevent_home "$HKEEP" +orphan_pe "$HORPHAN" register lavish orphan-src -- "$ORPHAN_STUB" "$TMP_ROOT/orphan-dead" >/dev/null +orphan_pe "$HKEEP" register lavish keep-src -- "$QUIET_STUB" "$TMP_ROOT/orphan-live" >/dev/null +orphan_pe "$HORPHAN" reconcile >/dev/null +orphan_pe "$HKEEP" reconcile >/dev/null + +wait_for "$HORPHAN/state/procevent/orphan-src.runner" \ + || fail "the dead-owner listener never recorded its runner" +wait_for "$HKEEP/state/procevent/keep-src.runner" \ + || fail "the live-owner listener never recorded its runner" +wait_for "$TMP_ROOT/orphan-dead.descendant" \ + || fail "the dead-owner listener's child never spawned its own descendant" +ORPHAN_PID=$(cat "$HORPHAN/state/procevent/orphan-src.runner") +KEEP_PID=$(cat "$HKEEP/state/procevent/keep-src.runner") +ORPHAN_DESCENDANT=$(cat "$TMP_ROOT/orphan-dead.descendant") + +# The reproduction condition itself: the listener is already an orphan in the +# kernel's sense before anything is asserted about reaping it. +orphan_ppid=$(ps -o ppid= -p "$ORPHAN_PID" 2>/dev/null | tr -d '[:space:]') +[ "$orphan_ppid" = 1 ] \ + || fail "the listener under test was not reparented away from its session (ppid $orphan_ppid)" +kill -0 -"$ORPHAN_PID" 2>/dev/null \ + || fail "the listener's process group was not running" +kill -0 "$ORPHAN_DESCENDANT" 2>/dev/null \ + || fail "the listener's descendant was not running" +pass "a detached listener starts reparented, with a live descendant tree under it" + +# Only the second home's session stays present, on the same short bound, so the +# owning session is the single difference between the two listeners. +keep_owner_present() { orphan_pe "$HKEEP" reconcile >/dev/null 2>&1 || true; sleep 0.25; } + +deadline=$((SECONDS + 40)) +while kill -0 -"$ORPHAN_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "a listener whose owning session was gone kept its process group running" + keep_owner_present +done +deadline=$((SECONDS + 20)) +while kill -0 "$ORPHAN_DESCENDANT" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "a listener whose owning session was gone left a descendant running" + keep_owner_present +done +pass "a listener whose owning session is gone stops itself and its whole process group" + +keep_owner_present +before=$(wc -l < "$TMP_ROOT/orphan-dead.ticks" | tr -d ' ') +deadline=$((SECONDS + 2)) +while [ "$SECONDS" -lt "$deadline" ]; do keep_owner_present; done +after=$(wc -l < "$TMP_ROOT/orphan-dead.ticks" | tr -d ' ') +[ "$before" = "$after" ] \ + || fail "the reaped listener's descendant kept spawning processes ($before then $after)" +pass "reaping the listener stops the process churn under it" + +keep_owner_present +kill -0 -"$KEEP_PID" 2>/dev/null \ + || fail "an identical listener in a home whose session is still there was reaped too" +pass "an identical listener in a home whose session is still there is untouched" + +# Retirement remains the explicit path, and it must reach a listener that has +# already reparented, along with everything under it. +wait_for "$TMP_ROOT/orphan-live.descendant" \ + || fail "the live-owner listener's child never spawned its own descendant" +KEEP_DESCENDANT=$(cat "$TMP_ROOT/orphan-live.descendant") +keep_owner_present +orphan_pe "$HKEEP" retire keep-src >/dev/null +wait_gone "-$KEEP_PID" \ + || fail "retiring a source left its reparented listener's process group running" +wait_gone "$KEEP_DESCENDANT" \ + || fail "retiring a source left a descendant of its listener running" +pass "retiring a source reaps its reparented listener and every descendant under it" + +# --- an expired runner's guard retries unproved cleanup --------------------- +# +# A stop the guard cannot PROVE must not end the guard. A descendant still +# finishing uninterruptible work outlives even the group KILL, and a guard that +# gave up after one attempt would walk away from a still-running expired runner. +# +# The unprovable attempt is injected through the signal the real path actually +# reads: Perl answers ONE process-group query for the runner with a group it +# does not lead, which is exactly how a stop that cannot be proved is reported. +# Every other Perl call, and every later query, runs the real interpreter. + +RETRY_HOME="$TMP_ROOT/stop-retry"; new_home "$RETRY_HOME" +fm_test_track_procevent_home "$RETRY_HOME" +RETRY_STATE="$TMP_ROOT/stop-retry-state"; mkdir -p "$RETRY_STATE" +RETRY_BIN=$(fm_fakebin "$TMP_ROOT/stop-retry-bin") +REAL_PERL=$(command -v perl) || fail "this host has no perl to build the retry fixture on" +cat > "$RETRY_BIN/perl" <<SH +#!/usr/bin/env bash +if [ "\$#" -eq 3 ] && [ "\$1" = -we ] \\ + && [ -s "\$STOP_RETRY_STATE/target" ] \\ + && [ "\$3" = "\$(cat "\$STOP_RETRY_STATE/target")" ] \\ + && [ ! -e "\$STOP_RETRY_STATE/spent" ]; then + case "\$2" in + *getpgrp*) + : > "\$STOP_RETRY_STATE/spent" + printf ' 999999\n' + exit 0 + ;; + esac +fi +exec "$REAL_PERL" "\$@" +SH +chmod +x "$RETRY_BIN/perl" + +retry_pe() { # <command...> + PATH="$RETRY_BIN:$PATH" STOP_RETRY_STATE="$RETRY_STATE" \ + FM_PROCEVENT_OWNER_LEASE_SECONDS=2 FM_PROCEVENT_OWNER_CHECK_SECONDS=1 \ + FM_HOME="$RETRY_HOME" "$ROOT/bin/fm-procevent.sh" "$@" +} + +retry_pe register lavish retry-src -- "$ORPHAN_STUB" "$TMP_ROOT/stop-retry-marker" >/dev/null +retry_pe reconcile >/dev/null +wait_for "$RETRY_HOME/state/procevent/retry-src.runner" \ + || fail "the retry listener never recorded its runner" +RETRY_PID=$(cat "$RETRY_HOME/state/procevent/retry-src.runner") +# Armed only now: the runner already proved its own process group at startup, +# and arming earlier would fail that assertion instead of the stop under test. +printf '%s\n' "$RETRY_PID" > "$RETRY_STATE/target" +wait_for "$TMP_ROOT/stop-retry-marker.descendant" \ + || fail "the retry listener's child never spawned its own descendant" +RETRY_DESCENDANT=$(cat "$TMP_ROOT/stop-retry-marker.descendant") + +deadline=$((SECONDS + 60)) +while kill -0 -"$RETRY_PID" 2>/dev/null; do + [ "$SECONDS" -lt "$deadline" ] \ + || fail "the guard gave up on an expired runner after a stop it could not prove" + sleep 0.5 +done +[ -e "$RETRY_STATE/spent" ] \ + || fail "the unprovable stop attempt this test injects never happened" +wait_gone "$RETRY_DESCENDANT" \ + || fail "the guard stopped retrying before the expired runner's descendant was reaped" +pass "a stop the guard cannot prove is retried until the expired runner is reaped" + printf '\nall procevent tests passed\n' diff --git a/tests/fm-quota-array-dispatch-live-e2e.test.sh b/tests/fm-quota-array-dispatch-live-e2e.test.sh index 417aeef86ca..5c9ace3e7fd 100755 --- a/tests/fm-quota-array-dispatch-live-e2e.test.sh +++ b/tests/fm-quota-array-dispatch-live-e2e.test.sh @@ -8,10 +8,10 @@ # fall back without the call log catching it. set -u -if [ "${FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E=1 to run the credentialed Pi dispatch-selection regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E pi python3 ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" OWNER="$ROOT/.agents/skills/quota-array-dispatch/SKILL.md" @@ -21,8 +21,6 @@ fail() { exit 1 } -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v python3 >/dev/null 2>&1 || fail "python3 not found" [ -f "$OWNER" ] || fail "quota-array-dispatch skill not found" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-quota-array-dispatch-live.XXXXXX") diff --git a/tests/fm-quota-choose.test.sh b/tests/fm-quota-choose.test.sh index 50dee72a117..88ae72fbdbb 100755 --- a/tests/fm-quota-choose.test.sh +++ b/tests/fm-quota-choose.test.sh @@ -215,6 +215,24 @@ fi [ "$out" = "none" ] || fail "specific scope: expected 'none', got '$out'" ok "specific model scope bounds generic quota" +if out=$(call_choose --snapshot "$LAB/captured.json" --candidate omp:openai-codex/codex_bengalfox 2>/dev/null); then + fail "omp prefix: the bare codex model scope did not veto, got exit 0 with '$out'" +fi +[ "$out" = "none" ] || fail "omp prefix: expected 'none' from the exhausted codex model scope, got '$out'" +out=$(call_choose --snapshot "$LAB/captured.json" --candidate omp:openai-codex/codex_other) +[ "$out" = "omp openai-codex/codex_other" ] || fail "omp prefix: expected the provider-wide codex quota to select the prefixed model, got '$out'" +ok "omp openai-codex prefix matches the bare codex model scope" + +if err=$(call_choose --snapshot "$LAB/captured.json" --candidate omp:ollama/qwen3:8b --candidate claude:claude-3-5-sonnet 2>&1); then + fail "unmapped omp prefix unexpectedly selected a later candidate" +fi +[ "$err" = "error: omp quota mapping covers only the openai-codex and claude-bridge prefixes: ollama/qwen3:8b" ] || fail "unmapped omp prefix returned: $err" +if err=$(call_choose --snapshot "$LAB/captured.json" --candidate omp 2>&1); then + fail "bare omp candidate unexpectedly selected" +fi +[ "$err" = "error: omp quota mapping covers only the openai-codex and claude-bridge prefixes: default" ] || fail "bare omp candidate returned: $err" +ok "omp without a mapped prefix fails closed" + out=$(call_choose --snapshot "$LAB/captured.json" --candidate codex:default) [ "$out" = "codex default" ] || fail "default scope: expected provider-wide quota, got '$out'" ok "default model uses provider-wide quota" diff --git a/tests/fm-reconcile-validation.test.mjs b/tests/fm-reconcile-validation.test.mjs new file mode 100644 index 00000000000..4d25220034d --- /dev/null +++ b/tests/fm-reconcile-validation.test.mjs @@ -0,0 +1,170 @@ +import assert from "node:assert/strict"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import { spawnSync } from "node:child_process"; +import { fileURLToPath } from "node:url"; +import { setTimeout as delay } from "node:timers/promises"; +import test from "node:test"; + +const root = path.resolve(path.dirname(fileURLToPath(import.meta.url)), ".."); +const controller = path.join(root, "skills", "reconcile-firstmate-upstream", "scripts", "validate-local.mjs"); +const bash = process.env.FM_RECONCILE_TEST_BASH ?? + (process.platform === "win32" ? path.join(process.env.ProgramFiles ?? "C:\\Program Files", "Git", "bin", "bash.exe") : "bash"); +const context = { fork: "1".repeat(40), upstream: "2".repeat(40) }; + +function fixture(t) { + const temporary = fs.mkdtempSync(path.join(os.tmpdir(), "fm-reconcile-validation-")); + const repo = path.join(temporary, "repo"); + const state = path.join(temporary, "state"); + const planFile = path.join(temporary, "plan.json"); + fs.mkdirSync(path.join(repo, "bin"), { recursive: true }); + fs.copyFileSync(path.join(root, "bin", "fm-timeout-lib.sh"), path.join(repo, "bin", "fm-timeout-lib.sh")); + fs.writeFileSync(path.join(repo, "input.txt"), "initial"); + fs.writeFileSync(path.join(repo, "unrelated.md"), "documentation"); + t.after(() => fs.rmSync(temporary, { recursive: true, force: true })); + const marker = path.join(temporary, "executions"); + const command = [process.execPath, "-e", + 'require("node:fs").appendFileSync(process.argv[1], "run\\n")', marker]; + function check(id = "fixture", extra = {}) { + return { id, kind: "test", command, requires: [], ci: "fixture-ci", ...extra }; + } + function run(checks, options = {}) { + fs.writeFileSync(planFile, JSON.stringify({ context: options.context ?? context, checks })); + const result = spawnSync(process.execPath, [controller, "--plan", planFile, "--root", repo, + "--state-dir", state, "--bash", bash, "--budget-seconds", String(options.budget ?? 30), + ...(options.resume ? ["--resume"] : []), ...(options.preflight ? ["--preflight-only"] : [])], + { encoding: "utf8", timeout: 60000, env: { ...process.env, ...options.env } }); + assert.equal(result.error, undefined, result.error?.message); + const artifact = path.join(state, "result.json"); + return { ...result, document: fs.existsSync(artifact) ? JSON.parse(fs.readFileSync(artifact, "utf8")) : null }; + } + return { repo, state, marker, temporary, check, run }; +} + +test("preflight defers a missing prerequisite without executing the check", (t) => { + const f = fixture(t); + const result = f.run([f.check("missing", { requires: [["fm-nonexistent-prerequisite-5645cf38", "--version"]] })]); + assert.equal(result.status, 75, result.stdout + result.stderr); + assert.equal(result.document.results[0].status, "deferred"); + assert.equal(result.document.results[0].exit, 127); + assert.equal(result.document.results[0].ci, "fixture-ci"); + assert.equal(fs.existsSync(f.marker), false); +}); + +test("preflight-only never certifies or executes a check", (t) => { + const f = fixture(t); + const result = f.run([f.check("probe", { requires: [[process.execPath, "--version"]] })], { preflight: true }); + assert.equal(result.status, 75, result.stdout + result.stderr); + assert.equal(result.document.results[0].status, "preflight-only"); + assert.equal(fs.existsSync(f.marker), false); +}); + +test("resume reuses exact inputs, ignores unrelated prose, and invalidates changed inputs", (t) => { + const f = fixture(t); + const checks = [f.check("cache", { cache: true, inputs: ["input.txt"] })]; + const first = f.run(checks); + assert.equal(first.status, 0, first.stdout + first.stderr); + const originalLog = first.document.results[0].log; + fs.writeFileSync(path.join(f.repo, "unrelated.md"), "updated documentation"); + const cached = f.run(checks, { resume: true }); + assert.equal(cached.status, 0, cached.stdout + cached.stderr); + assert.equal(cached.document.results[0].status, "cached"); + assert.equal(cached.document.results[0].reused.log, originalLog); + assert.equal(fs.readFileSync(f.marker, "utf8"), "run\n"); + fs.writeFileSync(path.join(f.repo, "input.txt"), "new behavior"); + const changed = f.run(checks, { resume: true }); + assert.equal(changed.status, 0, changed.stdout + changed.stderr); + assert.equal(changed.document.results[0].status, "passed"); + assert.notEqual(changed.document.results[0].log, originalLog); + assert.equal(fs.readFileSync(f.marker, "utf8"), "run\nrun\n"); +}); + +test("cache identity includes environment, frozen context, and prerequisite versions", (t) => { + const f = fixture(t); + const versionFile = path.join(f.temporary, "tool-version"); + fs.writeFileSync(versionFile, "v1"); + const checks = [f.check("identity", { cache: true, inputs: ["input.txt"], + requires: [[process.execPath, "-e", 'process.stdout.write(require("node:fs").readFileSync(process.argv[1]))', versionFile]] })]; + assert.equal(f.run(checks).status, 0); + assert.equal(f.run(checks, { resume: true, env: { FM_RECONCILE_FIXTURE: "changed" } }).document.results[0].status, "passed"); + assert.equal(f.run(checks, { resume: true, context: { ...context, upstream: "3".repeat(40) } }).document.results[0].status, "passed"); + fs.writeFileSync(versionFile, "v2"); + assert.equal(f.run(checks, { resume: true }).document.results[0].status, "passed"); + assert.equal(fs.readFileSync(f.marker, "utf8"), "run\n".repeat(4)); +}); + +test("failures and successful exits with skips never enter the pass cache", (t) => { + const f = fixture(t); + const checks = [ + f.check("failure", { cache: true, inputs: ["input.txt"], command: [process.execPath, "-e", "process.exit(3)"] }), + f.check("skip", { cache: true, inputs: ["input.txt"], command: [process.execPath, "-e", 'console.log("x".repeat(70000)); console.log("skip: optional fixture capability absent")'] }), + ]; + for (let attempt = 0; attempt < 2; attempt++) { + const result = f.run(checks, { resume: true }); + assert.equal(result.status, 1, result.stdout + result.stderr); + assert.deepEqual(result.document.results.map((item) => item.status), ["failed", "deferred"]); + assert.equal(fs.existsSync(path.join(f.state, "cache.json")), false); + } +}); + +test("two timeouts persist partial results and defer additional work", (t) => { + const f = fixture(t); + const slow = { command: [process.execPath, "-e", "setInterval(() => {}, 1000)"], timeoutSeconds: 2 }; + const result = f.run([f.check("timeout-one", slow), f.check("timeout-two", slow), f.check("not-started")]); + assert.equal(result.status, 75, result.stdout + result.stderr); + assert.deepEqual(result.document.results.map((item) => item.status), ["timeout", "timeout", "deferred"]); + assert.equal(result.document.results[2].reason, "timeout circuit breaker"); + assert.equal(result.document.timeouts, 2); + assert.equal(fs.existsSync(f.marker), false); +}); + +test("the wall budget cancels a running command and prevents its delayed result", async (t) => { + const f = fixture(t); + const started = path.join(f.temporary, "started"); + const late = path.join(f.temporary, "late"); + const command = [process.execPath, "-e", + 'const fs=require("node:fs"); fs.writeFileSync(process.argv[1],"started"); setTimeout(()=>fs.writeFileSync(process.argv[2],"late"),8000)', + started, late]; + const begin = Date.now(); + const result = f.run([f.check("budget", { command }), f.check("not-started")], { budget: 3 }); + assert.equal(result.status, 75, result.stdout + result.stderr); + assert.equal(fs.existsSync(started), true, "budget test must reach the real command"); + assert.equal(result.document.results[0].status, "timeout"); + assert.equal(result.document.results[1].reason, "local budget exhausted"); + assert.ok(Date.now() - begin < 15000, "the budget must interrupt rather than wait for a command to finish"); + await delay(Math.max(0, 10000 - (Date.now() - begin))); + assert.equal(fs.existsSync(late), false, "timed-out command survived and wrote its delayed result"); + assert.equal(fs.existsSync(f.marker), false); +}); + +test("caching refuses missing or out-of-scope dependency inputs", (t) => { + const f = fixture(t); + for (const inputs of [[], ["."], ["../"], ["absent"]]) { + const result = f.run([f.check("invalid", { cache: true, inputs })]); + assert.equal(result.status, 2, result.stdout + result.stderr); + assert.equal(fs.existsSync(f.marker), false); + } +}); + +test("registered shell cases preserve default order, select exactly one, and refuse typos", (t) => { + const f = fixture(t); + const library = path.join(root, "tests", "lib.sh").replaceAll("\\", "/"); + const script = `. "$1"\ntest_first() { echo first; }\ntest_second() { echo second; }\nfm_test_run_cases test_first test_second`; + const run = (env) => spawnSync(bash, ["-c", script, "_", library], + { encoding: "utf8", timeout: 30000, env: { ...process.env, TMPDIR: f.temporary.replaceAll("\\", "/"), ...env } }); + assert.match(run({ FM_TEST_ONLY: "", FM_TEST_LIST_CASES: "0" }).stdout, /first\nsecond/); + const selected = run({ FM_TEST_ONLY: "test_second", FM_TEST_LIST_CASES: "0" }); + assert.equal(selected.status, 0, selected.stdout + selected.stderr); + assert.equal(selected.stdout.trim(), "second"); + const listed = run({ FM_TEST_ONLY: "", FM_TEST_LIST_CASES: "1" }); + assert.equal(listed.status, 0, listed.stderr); + assert.equal(listed.stdout.trim(), "test_first\ntest_second"); + const typo = run({ FM_TEST_ONLY: "test_typo", FM_TEST_LIST_CASES: "0" }); + assert.notEqual(typo.status, 0); + assert.match(typo.stderr, /unknown case: test_typo/); + assert.equal(typo.stdout, ""); + const invalidListing = run({ FM_TEST_ONLY: "", FM_TEST_LIST_CASES: "typo" }); + assert.notEqual(invalidListing.status, 0); + assert.equal(invalidListing.stdout, ""); +}); diff --git a/tests/fm-reconcile-validation.test.sh b/tests/fm-reconcile-validation.test.sh new file mode 100755 index 00000000000..24986a16896 --- /dev/null +++ b/tests/fm-reconcile-validation.test.sh @@ -0,0 +1,6 @@ +#!/usr/bin/env bash +# Portable behavior checks for bounded reconciliation plans and case selection. +set -eu +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +command -v node >/dev/null 2>&1 || { echo "not ok - node is required for reconciliation validation tests" >&2; exit 1; } +exec node --test "$ROOT/tests/fm-reconcile-validation.test.mjs" diff --git a/tests/fm-remote-job-orphan-reap.test.sh b/tests/fm-remote-job-orphan-reap.test.sh index 0c52a4c9012..667be925aa3 100755 --- a/tests/fm-remote-job-orphan-reap.test.sh +++ b/tests/fm-remote-job-orphan-reap.test.sh @@ -61,6 +61,30 @@ wait_child() { # <pid> <seconds> return 1 } +# True when <pid>'s parent is a reaper for orphaned processes: init itself, or +# a subreaper systemd registers one hop below init (PR_SET_CHILD_SUBREAPER, +# e.g. `systemd --user`) - a live host's per-user manager adopts orphans there +# instead of letting them reach real init, and that is just as orphaned for +# this fixture's purpose. +is_orphaned() { # <pid> + local parent + parent=$(ppid_of "$1") + case "$parent" in ''|*[!0-9]*) return 1 ;; esac + [ "$parent" = 1 ] && return 0 + [ "$(ppid_of "$parent")" = 1 ] +} + +# Wait up to <seconds> for <pid> to be reparented to an orphan reaper (see +# is_orphaned) after its launching shell exits; 0 when it does. +wait_orphaned() { # <pid> <seconds> + local pid=$1 deadline=$(( $(date +%s) + $2 )) + while [ "$(date +%s)" -lt "$deadline" ]; do + is_orphaned "$pid" && return 0 + sleep 0.1 + done + return 1 +} + # --- a real worker fixture, launched exactly the way fm-on's Linux start does - # build_remote_root <dir>: a minimal but genuine Firstmate code root carrying @@ -123,7 +147,7 @@ SERVE=$(pgrep -P "$WORKER" | head -n 1) fail "the serving child is outside the worker's process group" pass "the Linux start path puts the whole worker tree in its own process group" -[ "$(ppid_of "$WORKER")" = 1 ] || +wait_orphaned "$WORKER" 5 || fail "the fixture worker is not orphaned to init, so this case does not reproduce the leak" # The exact teardown shape that leaked in production: a fixture cleanup removes @@ -137,7 +161,7 @@ kill -KILL "$SERVE" 2>/dev/null || true wait_gone "$SERVE" 10 || fail "the recorded serving child did not stop" alive "$WORKER" || fail "the fixture supervisor did not survive a lone child kill, so this case no longer covers the leak" wait_child "$WORKER" 15 || fail "the supervisor did not respawn after its recorded child pid was killed" -pass "removing the state root and killing the recorded worker pid leaves the tree running at ppid 1" +pass "removing the state root and killing the recorded worker pid leaves the tree running, orphaned" # A worker whose code root is intact is never a reap candidate, which is what # keeps the account's healthy LaunchAgent worker out of scope. diff --git a/tests/fm-rovo-harness.test.sh b/tests/fm-rovo-harness.test.sh new file mode 100644 index 00000000000..bdfd17d2f26 --- /dev/null +++ b/tests/fm-rovo-harness.test.sh @@ -0,0 +1,464 @@ +#!/usr/bin/env bash +# Behavior tests for the verified Rovo CLI crewmate/scout adapter. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# bin/fm-harness.sh checks verified ENV markers before ancestry. A suite run +# from inside Cursor, Claude, Pi, or Grok inherits those markers, which outrank +# the fake ancestry the detection cases set up. Drop the ambient markers so the +# asserted verdict does not depend on which harness launched the suite. +unset CLAUDECODE PI_CODING_AGENT FM_PI_HARNESS GROK_AGENT CURSOR_AGENT CURSOR_INVOKED_AS \ + ATLASSIAN_AGENT_TYPE ROVODEV_CLI + +SPAWN="$ROOT/bin/fm-spawn.sh" +TMP_ROOT=$(fm_test_tmproot fm-rovo-harness) + +# A stateful fake tmux for rovo's launch-then-send shape (the same shape kimi +# uses): a positional brief is dead-on-arrival, so rovo launches BARE and only +# receives an absolute brief pointer after a readiness gate, then a delivery +# gate. This fake renders a rovo-shaped screen that advances through +# launched -> ready -> pointer-typed -> delivered as the real spawn drives it, so +# the launch command, the typed pointer, and both gates are exercised through +# their real code paths rather than asserted from static text. +make_rovo_fakebin() { + local dir=$1 fakebin + fakebin=$(fm_fakebin "$dir") + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +set -u +printf '%s\n' "$*" >> "$FM_FAKE_TMUX_CALL_LOG" +state=$(cat "$FM_FAKE_ROVO_STATE" 2>/dev/null || true) +fake_screen() { + case "$state" in + ready) + printf 'Welcome to Rovo!\nContext: | 0.0%% 0/922K\n? for shortcuts.\n╭────────────────────────────────╮\n│ > │\n╰────────────────────────────────╯\n' + ;; + pointer-typed) + printf 'Context: | 0.0%% 0/922K\n╭────────────────────────────────╮\n│ > Read the brief and follow it │\n│ │\n╰────────────────────────────────╯\n' + ;; + delivered) + printf 'Read the brief at %s and follow it exactly.\nContext: | 3.3%% 30.1K/922K\n╭────────────────────────────────╮\n│ > │\n╰────────────────────────────────╯\n' "$FM_FAKE_BRIEF_REAL" + ;; + *) + printf 'shell starting\n$ \n' + ;; + esac +} +fake_cursor_y() { + case "$state" in + pointer-typed) printf '4\n' ;; + ready|delivered) printf '4\n' ;; + *) printf '1\n' ;; + esac +} +case "$*" in + *"#{pane_current_path}"*) printf '%s\n' "$FM_FAKE_PANE_PATH"; exit 0 ;; + *"#{cursor_y}"*) fake_cursor_y; exit 0 ;; +esac +case "${1:-}" in + display-message) printf 'firstmate\n'; exit 0 ;; + list-windows) exit 0 ;; + has-session|new-session|new-window|kill-window) exit 0 ;; + send-keys) + prev= + literal= + for arg in "$@"; do + if [ "$prev" = -l ]; then literal=$arg; break; fi + prev=$arg + done + if [ -n "$literal" ]; then + case "$literal" in + *'run --yolo'*) + printf '%s\n' "$literal" >> "$FM_FAKE_LAUNCH_LOG" + printf 'launched\n' > "$FM_FAKE_ROVO_STATE" + ;; + *) + printf '%s\n' "$literal" >> "$FM_FAKE_POINTER_LOG" + printf 'pointer-typed\n' > "$FM_FAKE_ROVO_STATE" + ;; + esac + exit 0 + fi + case " $* " in + *' Enter '*) + case "$state" in + launched) + if [ "${FM_FAKE_ROVO_READY:-yes}" = yes ]; then + printf 'ready\n' > "$FM_FAKE_ROVO_STATE" + fi + ;; + pointer-typed) + if [ "${FM_FAKE_ROVO_DELIVERY:-yes}" = yes ]; then + printf 'delivered\n' > "$FM_FAKE_ROVO_STATE" + else + printf 'ready\n' > "$FM_FAKE_ROVO_STATE" + fi + ;; + esac + ;; + esac + exit 0 + ;; + capture-pane) + start= end= prev= + for arg in "$@"; do + case "$prev" in + -S) start=$arg ;; + -E) end=$arg ;; + esac + case "$arg" in -S|-E) prev=$arg ;; *) prev= ;; esac + done + case "$start:$end" in + *[!0-9:]*|'':*|*:'') fake_screen ;; + *) fake_screen | awk -v start="$start" -v end="$end" \ + 'NR - 1 >= start && NR - 1 <= end' ;; + esac + exit 0 + ;; +esac +exit 0 +SH + chmod +x "$fakebin/tmux" + fm_fake_exit0 "$fakebin" treehouse gh-axi gh + fm_fake_exit0 "$fakebin" rovo + printf '%s\n' "$fakebin" +} + +make_spawn_case() { + local name=$1 id=$2 case_dir home proj wt fakebin + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + proj="$case_dir/project" + wt="$case_dir/wt" + fakebin=$(make_rovo_fakebin "$case_dir/fake") + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + cat > "$home/data/$id/brief.md" <<'EOF' +# Task +## Captain's intent +Exercise Rovo dispatch. + +## Firstmate spec +Verify launch and delivery behavior. +EOF + printf 'rovo\n' > "$home/config/crew-harness" + fm_git_worktree "$proj" "$wt" "wt-$name" + touch "$home/state/.last-watcher-beat" + : > "$case_dir/launch.log" + : > "$case_dir/pointer.log" + : > "$case_dir/rovo.state" + : > "$case_dir/tmux-calls.log" + printf '%s\n' "$case_dir|$home|$proj|$wt|$fakebin" +} + +read_spawn_record() { + IFS='|' read -r CASE_DIR HOME_DIR PROJ_DIR WT_DIR FAKEBIN_DIR <<EOF +$1 +EOF +} + +BASE_PATH=${FM_TEST_BASE_PATH:-/usr/bin:/bin:/usr/sbin:/sbin} + +run_spawn() { + local case_dir=$1 home=$2 proj=$3 wt=$4 fakebin=$5 id=$6 + shift 6 + HOME="$home" FM_ROOT_OVERRIDE='' FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" FM_DATA_OVERRIDE="$home/data" \ + FM_PROJECTS_OVERRIDE="$home/projects" FM_CONFIG_OVERRIDE="$home/config" \ + FM_SPAWN_NO_GUARD=1 FM_FAKE_PANE_PATH="$wt" TMUX="fake,1,0" \ + FM_FAKE_LAUNCH_LOG="$case_dir/launch.log" \ + FM_FAKE_POINTER_LOG="$case_dir/pointer.log" \ + FM_FAKE_ROVO_STATE="$case_dir/rovo.state" \ + FM_FAKE_TMUX_CALL_LOG="$case_dir/tmux-calls.log" \ + FM_FAKE_BRIEF_REAL="$(cd "$home/data/$id" && pwd -P)/launch-brief.md" \ + FM_FAKE_ROVO_READY="${FM_FAKE_ROVO_READY:-yes}" \ + FM_FAKE_ROVO_DELIVERY="${FM_FAKE_ROVO_DELIVERY:-yes}" \ + FM_ROVO_READY_POLLS=3 FM_ROVO_DELIVERY_POLLS=3 FM_ROVO_POLL_INTERVAL=0 \ + PATH="$fakebin:$BASE_PATH" \ + "$SPAWN" "$id" "$proj" --harness rovo --mode no-mistakes --yolo off "$@" 2>&1 +} + +test_rovo_launch_then_send_is_verified() { + local id rec out rc launch pointer brief_real meta data_real state_real + id="rovo-success-z1-$$" + rec=$(make_spawn_case success "$id") + read_spawn_record "$rec" + out=$(run_spawn \ + "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" \ + --model auto --effort high) + rc=$? + expect_code 0 "$rc" "verified rovo launch-then-send should succeed" + assert_contains "$out" "spawned $id harness=rovo" "rovo spawn did not report success" + + launch=$(cat "$CASE_DIR/launch.log") + assert_contains "$launch" "$FAKEBIN_DIR/rovo' run --yolo" \ + "rovo launch did not use the resolved binary with the bare launch-then-send shape" + assert_not_contains "$launch" "--startup-receipt" "rovo launch used the incompatible --startup-receipt flag" + assert_not_contains "$launch" "encode launch-brief" "rovo launch carried a positional brief instead of launching bare" + assert_not_contains "$launch" "brief for rovo" "rovo launch embedded the brief body as a positional argument" + assert_contains "$launch" "--model 'auto'" "rovo launch omitted the requested model" + assert_contains "$launch" "env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT -u FM_PI_HARNESS" \ + "rovo launch did not clear foreign primary markers" + assert_contains "$launch" "env -u CURSOR_AGENT -u CURSOR_INVOKED_AS" \ + "rovo launch did not clear cursor's markers via the shared outer wrap" + assert_not_contains "$launch" "turn-ended" "rovo launch embedded a turn-end path it does not own" + + brief_real="$(cd "$HOME_DIR/data/$id" && pwd -P)/launch-brief.md" + pointer=$(cat "$CASE_DIR/pointer.log") + [ "$pointer" = "Read the brief at $brief_real and follow it exactly." ] \ + || fail "rovo pointer was not the exact absolute-path-only instruction: $pointer" + + meta="$HOME_DIR/state/$id.meta" + assert_grep 'model=auto' "$meta" "rovo meta lost the requested model" + assert_grep 'effort=high' "$meta" "rovo meta lost the requested effort" + assert_not_contains "$(cat "$CASE_DIR/tmux-calls.log")" "kill-window" \ + "a successful rovo spawn must never tear down the endpoint it just delivered into" + + # rovo confines every file-tool operation to its worktree by default + # (confirmed live), so the launch must grant allowedExternalPaths covering + # this task's brief directory, steering inbox, and status file - otherwise + # the standard instructions/steering/status/report loop cannot work. + data_real=$(cd "$HOME_DIR/data/$id" && pwd -P) + state_real=$(cd "$HOME_DIR/state" && pwd -P) + assert_contains "$launch" "allowedExternalPaths" \ + "rovo launch did not grant allowedExternalPaths for this task's home paths" + assert_contains "$launch" "$data_real" \ + "rovo launch's allowedExternalPaths grant omitted the brief directory" + assert_contains "$launch" "$state_real/$id.inbox" \ + "rovo launch's allowedExternalPaths grant omitted the steering inbox directory" + assert_contains "$launch" "$state_real/$id.status" \ + "rovo launch's allowedExternalPaths grant omitted the status file" + pass "fm-spawn: rovo launches bare, waits for readiness, and delivers its brief pointer" +} + +test_rovo_effort_xhigh_is_recorded_but_omitted() { + local id rec out rc launch meta + id="rovo-xhigh-z2-$$" + rec=$(make_spawn_case xhigh "$id") + read_spawn_record "$rec" + out=$(run_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" --effort xhigh) + rc=$? + expect_code 0 "$rc" "rovo spawn with an unsupported effort should still succeed" + launch=$(cat "$CASE_DIR/launch.log") + assert_not_contains "$launch" "efficiencyLevel" "rovo launch set agent.efficiencyLevel for an unsupported effort value" + # --config-override still carries the mandatory allowedExternalPaths grant + # even when the requested effort is unsupported and omitted: the two ride + # one merged JSON object because rovo's --config-override is single-value + # (a second occurrence silently discards the first, confirmed live). + assert_contains "$launch" "allowedExternalPaths" "rovo launch dropped its allowedExternalPaths grant when effort was unsupported" + meta="$HOME_DIR/state/$id.meta" + assert_grep 'effort=xhigh' "$meta" "rovo meta did not retain the unsupported effort axis" + pass "fm-spawn: rovo omits efficiencyLevel for xhigh but keeps its allowedExternalPaths grant, recording xhigh in task metadata" +} + +test_rovo_effort_high_sets_config_override() { + local id rec out rc launch override_count + id="rovo-effort-z6-$$" + rec=$(make_spawn_case effort-high "$id") + read_spawn_record "$rec" + out=$(run_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id" --effort high) + rc=$? + expect_code 0 "$rc" "rovo spawn with a supported effort should succeed" + launch=$(cat "$CASE_DIR/launch.log") + assert_contains "$launch" 'efficiencyLevel' "rovo launch did not set agent.efficiencyLevel via --config-override" + assert_contains "$launch" '"high"' "rovo launch did not carry the requested efficiency level" + # Exactly one --config-override: rovo silently discards a second occurrence + # (confirmed live), so efficiencyLevel and the allowedExternalPaths grant + # must ride the same merged JSON object rather than two flags. + override_count=$(printf '%s' "$launch" | grep -o -- '--config-override' | wc -l | tr -d ' ') + [ "$override_count" = 1 ] \ + || fail "rovo launch emitted $override_count --config-override occurrences; a second one would silently discard the first" + assert_contains "$launch" 'allowedExternalPaths' "rovo launch's single --config-override dropped the allowedExternalPaths grant when merging in efficiencyLevel" + pass "fm-spawn: rovo's supported effort values merge into the same --config-override as its allowedExternalPaths grant" +} + +test_rovo_readiness_gate_precedes_pointer() { + local id rec out rc + id="rovo-not-ready-z3-$$" + rec=$(make_spawn_case not-ready "$id") + read_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_ROVO_READY=no run_spawn \ + "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "rovo spawn without a ready signal should fail" + assert_contains "$out" "rovo did not show a verified ready signal" \ + "rovo readiness failure lacked a loud diagnostic" + assert_grep 'failed: rovo did not show a verified ready signal' "$HOME_DIR/state/$id.status" \ + "rovo readiness failure did not leave a supervisor-visible failure" + [ ! -s "$CASE_DIR/pointer.log" ] || fail "rovo pointer was sent before an observable ready signal" + grep -q "kill-window.*fm-$id" "$CASE_DIR/tmux-calls.log" \ + || fail "a failed rovo readiness gate must tear down the exact endpoint it created instead of leaking an orphaned --yolo process" + pass "fm-spawn: rovo never sends the brief pointer before an observable ready signal, and tears down the created endpoint on failure" +} + +test_rovo_unconfirmed_delivery_fails_loudly() { + local id rec out rc pointer + id="rovo-drop-z7-$$" + rec=$(make_spawn_case drop "$id") + read_spawn_record "$rec" + rc=0 + out=$(FM_FAKE_ROVO_DELIVERY=no run_spawn \ + "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "an unconfirmed rovo delivery should fail" + # The pointer was typed (readiness passed) but its delivery never confirmed. + pointer=$(cat "$CASE_DIR/pointer.log") + [ -n "$pointer" ] || fail "rovo never typed the pointer before the delivery gate" + assert_contains "$out" "rovo brief pointer delivery was not confirmed" \ + "unconfirmed rovo delivery lacked a loud diagnostic" + assert_grep 'failed: rovo brief pointer delivery was not confirmed' "$HOME_DIR/state/$id.status" \ + "unconfirmed rovo delivery did not leave a supervisor-visible failure" + grep -q "kill-window.*fm-$id" "$CASE_DIR/tmux-calls.log" \ + || fail "an unconfirmed rovo delivery must tear down the exact endpoint it created instead of leaking an orphaned --yolo process" + pass "fm-spawn: rovo treats a silent pointer drop as a failed spawn, and tears down the created endpoint" +} + +test_rovo_missing_binary_refuses_before_pane_creation() { + local id rec out rc + id="rovo-missing-z4-$$" + rec=$(make_spawn_case missing "$id") + read_spawn_record "$rec" + rm "$FAKEBIN_DIR/rovo" + rc=0 + out=$(run_spawn "$CASE_DIR" "$HOME_DIR" "$PROJ_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$id") || rc=$? + [ "$rc" -ne 0 ] || fail "missing rovo executable should refuse the spawn" + assert_contains "$out" "searched PATH for 'rovo'" "missing rovo diagnostic omitted PATH search" + [ -s "$CASE_DIR/launch.log" ] && fail "missing rovo executable created a launch command" || true + pass "fm-spawn: missing rovo executable refuses before pane creation" +} + +test_rovo_secondmate_is_refused() { + local id rec out rc + id="rovo-secondmate-z5-$$" + rec=$(make_spawn_case secondmate-refuse "$id") + read_spawn_record "$rec" + rc=0 + out=$(HOME="$HOME_DIR" FM_ROOT_OVERRIDE='' FM_HOME="$HOME_DIR" \ + FM_STATE_OVERRIDE="$HOME_DIR/state" FM_DATA_OVERRIDE="$HOME_DIR/data" \ + FM_PROJECTS_OVERRIDE="$HOME_DIR/projects" FM_CONFIG_OVERRIDE="$HOME_DIR/config" \ + FM_SPAWN_NO_GUARD=1 PATH="$FAKEBIN_DIR:$BASE_PATH" \ + "$SPAWN" "$id" --secondmate rovo 2>&1) || rc=$? + [ "$rc" -ne 0 ] || fail "a rovo secondmate spawn should be refused" + assert_contains "$out" "rovo is a verified crewmate/scout adapter only" \ + "rovo secondmate refusal lacked its concrete reason" + pass "fm-spawn: rovo cannot be launched as a secondmate" +} + +test_rovo_detection_precedence_and_ancestry() { + local dir fakebin cfg out + dir="$TMP_ROOT/detection" + fakebin=$(fm_fakebin "$dir") + cfg="$dir/config" + mkdir -p "$cfg" + cat > "$fakebin/ps" <<'SH' +#!/usr/bin/env bash +set -u +field= +pid= +prev= +for arg in "$@"; do + [ "$prev" = -o ] && field=$arg + [ "$prev" = -p ] && pid=$arg + prev=$arg +done +case "$field:$pid" in + comm=:4242) printf 'rovo\n' ;; + comm=:*) printf '/bin/bash\n' ;; + ppid=:4242) printf '1\n' ;; + ppid=:*) printf '4242\n' ;; + args=:*) printf 'bash\n' ;; +esac +SH + chmod +x "$fakebin/ps" + + out=$(env -u CLAUDECODE -u PI_CODING_AGENT -u FM_PI_HARNESS -u GROK_AGENT \ + -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u ATLASSIAN_AGENT_TYPE -u ROVODEV_CLI \ + PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + [ "$out" = rovo ] || fail "rovo ancestry detection returned '$out'" + + out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + ATLASSIAN_AGENT_TYPE=rovo CLAUDECODE=1 \ + PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + [ "$out" = rovo ] || fail "rovo's ATLASSIAN_AGENT_TYPE marker did not outrank an inherited CLAUDECODE, got '$out'" + + out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + ROVODEV_CLI=1 CLAUDECODE=1 \ + PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + [ "$out" = rovo ] || fail "rovo's ROVODEV_CLI marker did not outrank an inherited CLAUDECODE, got '$out'" + + out=$(env -u CURSOR_AGENT -u CURSOR_INVOKED_AS \ + CLAUDECODE=1 PATH="$fakebin:$BASE_PATH" FM_CONFIG_OVERRIDE="$cfg" "$ROOT/bin/fm-harness.sh") + [ "$out" = claude ] || fail "verified env-marker precedence changed, got '$out'" + pass "fm-harness: rovo's markers outrank an inherited CLAUDECODE, and markerless ancestry still resolves rovo" +} + +test_rovo_control_lib_table() { + # shellcheck source=/dev/null + . "$ROOT/bin/fm-control-lib.sh" + [ "$(fm_control_interrupt_key rovo)" = Escape ] || fail "rovo interrupt key is not Escape" + [ "$(fm_control_interrupt_repeat rovo)" = 1 ] || fail "rovo interrupt repeat is not 1" + [ -z "$(fm_control_interrupt_clear_key rovo)" ] || fail "rovo should need no interrupt clear key" + [ "$(fm_control_interrupt_ack_source rovo)" = none ] || fail "rovo interrupt ack source is not none" + [ "$(fm_control_exit_command rovo)" = /exit ] || fail "rovo exit command is not /exit" + [ "$(fm_control_harness_family rovo-anything)" = rovo ] || fail "rovo harness family prefix match failed" + fm_control_harness_supports_kind rovo ship || fail "rovo should support ship tasks" + fm_control_harness_supports_kind rovo scout || fail "rovo should support scout tasks" + if fm_control_harness_supports_kind rovo secondmate; then + fail "rovo should never support secondmate tasks" + fi + pass "fm-control-lib: rovo's lifecycle table matches its verified facts" +} + +test_rovo_busy_regex_isolated() { + # shellcheck source=/dev/null + . "$ROOT/bin/fm-busy-lib.sh" + printf 'Enter to queue, Ctrl+Enter to steer\n⬢ Rovo is thinking...\n' | fm_busy_rovo_tail_busy \ + || fail "rovo's real busy line was not recognized as busy" + printf 'Context: 5.3%% 48.7K/922K\n? for shortcuts.\n' | fm_busy_rovo_tail_busy \ + && fail "an idle rovo composer footer was misread as busy" + printf 'Ctrl+c:cancel\n' | fm_busy_rovo_tail_busy \ + && fail "Grok's exact busy token leaked into rovo's harness-scoped matcher" + printf 'Rovo is thinking\n' | fm_busy_grok_tail_busy \ + && fail "rovo's busy line leaked into Grok's harness-scoped matcher" + + local out + out=$(fm_busy_classify tmux fake:0 rovo taskid /nonexistent-state '⬢ Rovo is thinking...') + [ "$out" = "busy rovo-regex" ] || fail "fm_busy_classify did not read a real rovo busy tail as busy rovo-regex, got '$out'" + # A tail with no busy marker at all is "can't tell," never definitive idle: + # rovo's fallback is best-effort, so absence of the marker must not let + # supervision conclude a still-working worker went idle. + out=$(fm_busy_classify tmux fake:0 rovo taskid /nonexistent-state 'Context: 1% 2K/900K') + [ "$out" = "unknown rovo-regex" ] || fail "fm_busy_classify misread a marker-absent rovo tail as definitive idle instead of unknown, got '$out'" + pass "busy detection: rovo's rendered busy line classifies through its own isolated fallback" +} + +test_rovo_busy_marker_scrolled_out_of_tail_is_unknown() { + # shellcheck source=/dev/null + . "$ROOT/bin/fm-busy-lib.sh" + local tail40 out i + # fm_busy_rovo_tail_busy only inspects the last 12 nonblank lines. Build a + # captured tail where the busy marker is present but pushed out of that + # window by a long active turn's own output, so a naive "marker absent" + # check would misread this still-busy worker as idle. + tail40='⬢ Rovo is thinking...' + for i in $(seq 1 20); do + tail40="$tail40 +tool output line $i" + done + out=$(fm_busy_classify tmux fake:0 rovo taskid /nonexistent-state "$tail40") + [ "$out" = "unknown rovo-regex" ] \ + || fail "an active rovo turn whose busy marker scrolled out of the tail must classify unknown, not idle; got '$out'" + pass "busy detection: a rovo busy marker scrolled out of the tail classifies unknown, never idle" +} + +test_rovo_launch_then_send_is_verified +test_rovo_effort_xhigh_is_recorded_but_omitted +test_rovo_effort_high_sets_config_override +test_rovo_readiness_gate_precedes_pointer +test_rovo_unconfirmed_delivery_fails_loudly +test_rovo_missing_binary_refuses_before_pane_creation +test_rovo_secondmate_is_refused +test_rovo_detection_precedence_and_ancestry +test_rovo_control_lib_table +test_rovo_busy_regex_isolated +test_rovo_busy_marker_scrolled_out_of_tail_is_unknown diff --git a/tests/fm-rovo-signals-live-e2e.test.sh b/tests/fm-rovo-signals-live-e2e.test.sh new file mode 100644 index 00000000000..3d9896849e8 --- /dev/null +++ b/tests/fm-rovo-signals-live-e2e.test.sh @@ -0,0 +1,358 @@ +#!/usr/bin/env bash +# Live guard for the real, installed Rovo CLI (bin/fm-test-run.sh's +# live-harness-optin family). Env-gated and self-skipping: it drives the real +# binary through a raw PTY (the same TTY contract tmux/herdr allocate) rather +# than requiring tmux, so it runs on hosts without tmux installed. It exercises +# the EXACT production launch-then-send shape - bare `rovo run --yolo` with NO +# positional brief, a readiness gate on the `Welcome to Rovo!` banner, a typed +# absolute brief pointer, then a delivery gate - and proves the +# harness-dependent facts bin/fm-busy-lib.sh and bin/fm-control-lib.sh encode for +# rovo: that the "Rovo is thinking" busy line renders for a real tool call, that +# an Escape sent mid-tool-call prints "Agent cancelled" without wedging the +# session, and that /exit then exits cleanly with rovo's resume hint. +# +# A positional brief is deliberately NOT used: it is dead-on-arrival (rovo loads, +# never enters a working state, and drops back to an idle shell within ~10-15s; +# confirmed live four times over raw PTY and once under real tmux with the exact +# send-keys shape). The launch-then-send shape is what fm-spawn.sh actually +# places, so it is what this guard drives. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +ROVO_BIN=$(command -v rovo 2>/dev/null || true) +[ -x "${ROVO_BIN:-}" ] || ROVO_BIN="$HOME/.local/bin/rovo" + +fail() { + printf 'not ok - %s\n' "$1" >&2 + exit 1 +} + +pass() { + printf 'ok - %s\n' "$1" +} + +fm_live_gate opt-in FM_ROVO_SIGNALS_LIVE python3 + +[ -x "$ROVO_BIN" ] || fail "FM_ROVO_SIGNALS_LIVE=1 but no real rovo executable is installed" + +VERSION_OUT=$("$ROVO_BIN" --version 2>&1) || fail "rovo --version failed: $VERSION_OUT" +echo "BOOTSTRAP_INFO: live rovo version: $VERSION_OUT" + +# shellcheck source=bin/fm-busy-lib.sh +. "$ROOT/bin/fm-busy-lib.sh" + +LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-rovo-signals.XXXXXX") || fail "could not create the isolated Rovo lab" +cleanup() { rm -rf -- "$LAB"; } +trap cleanup EXIT +mkdir -p "$LAB/workspace" +git -C "$LAB/workspace" init -q || fail "could not initialize the isolated Rovo workspace" +WORKSPACE=$(cd "$LAB/workspace" && pwd -P) || fail "could not resolve the isolated Rovo workspace" +TRANSCRIPT="$LAB/transcript.log" + +# The brief the typed pointer will reference. It triggers a slow bash tool call +# so the busy line and the interrupt window are both long enough to observe. +BRIEF="$LAB/brief.md" +printf 'Run this exact bash command and nothing else: sleep 25\n' > "$BRIEF" +BRIEF_REAL=$(cd "$LAB" && pwd -P)/brief.md + +# Drive the real binary over a raw PTY through the exact production +# launch-then-send shape: launch bare, wait for the readiness banner, type the +# absolute brief pointer, confirm delivery, observe the busy line, interrupt with +# Escape, and exit cleanly. Bytes are dumped raw to TRANSCRIPT for the shell-side +# substring checks below. +python3 - "$ROVO_BIN" "$WORKSPACE" "$TRANSCRIPT" "$BRIEF_REAL" <<'PY' || fail "the PTY driver reported a failure" +import os +import pty +import select +import subprocess +import sys +import time + +rovo_bin, workspace, transcript_path, brief_real = sys.argv[1:5] + +pointer = "Read the brief at %s and follow it exactly." % brief_real + +pid, fd = pty.fork() +if pid == 0: + os.chdir(workspace) + # Bare launch: NO positional brief. A message is typed in after readiness. + os.execvp(rovo_bin, [rovo_bin, "run", "--yolo"]) + os._exit(127) + +transcript = open(transcript_path, "wb") + +def pump(timeout, want=None): + deadline = time.time() + timeout + buf = b"" + while time.time() < deadline: + r, _, _ = select.select([fd], [], [], 0.5) + if fd in r: + try: + chunk = os.read(fd, 65536) + except OSError: + break + if not chunk: + break + transcript.write(chunk) + transcript.flush() + buf += chunk + if want and want in buf: + return buf + return buf + +# 1. Readiness gate: wait for rovo's fresh-launch welcome banner. +ready = pump(60, want=b"Welcome to Rovo!") +if b"Welcome to Rovo!" not in ready: + sys.exit("rovo never rendered its 'Welcome to Rovo!' readiness banner") + +# 2. Type the absolute brief pointer and submit it, then let the brief drive the +# slow bash tool call. The busy line proves both delivery and the working state. +os.write(fd, pointer.encode() + b"\r") +busy = pump(90, want=b"Rovo is thinking") +if b"Rovo is thinking" not in busy: + sys.exit("rovo never rendered its busy line after the typed brief pointer") + +# 3. Interrupt with Escape while the tool call is genuinely in flight. Real rovo +# prints "Agent cancelled" and stays controllable (confirmed live here and under +# real tmux 3.6a); hold the guard to that render. The exact instant the interrupt +# lands is timing-sensitive over a raw PTY - a single fixed-timer Escape can fall +# between states - so send Escape across the tool-call window until the cancel +# renders. This is a deterministic way to reproduce a timing-sensitive interrupt, +# not a weakening of the claim: a session that never rendered the cancel would +# exhaust every attempt and fail. +cancelled = b"" +for _ in range(15): + os.write(fd, b"\x1b") + cancelled = pump(2, want=b"Agent cancelled") + if b"Agent cancelled" in cancelled: + break +if b"Agent cancelled" not in cancelled: + sys.exit("rovo did not print 'Agent cancelled' after a mid-tool-call Escape") + +# 4. Exit cleanly, proving Escape left the session controllable. +time.sleep(1) +os.write(fd, b"/exit\r") +for _ in range(60): + try: + done_pid, status = os.waitpid(pid, os.WNOHANG) + except ChildProcessError: + done_pid = pid + status = 0 + if done_pid == pid: + break + pump(1) +else: + subprocess.run(["kill", "-9", str(pid)]) + sys.exit("rovo did not exit after /exit, sent after a mid-tool-call Escape") + +transcript.close() +PY + +grep -aFq 'Welcome to Rovo!' "$TRANSCRIPT" \ + || fail "real rovo never rendered its readiness banner" +pass "real rovo launches bare and renders its 'Welcome to Rovo!' readiness banner" + +grep -aFq 'Rovo is thinking' "$TRANSCRIPT" \ + || fail "real rovo never rendered its busy line from the typed brief pointer" +printf '%s\n' "Rovo is thinking..." | fm_busy_rovo_tail_busy \ + || fail "fm_busy_rovo_tail_busy did not classify the real busy line as busy" +pass "real rovo delivers the typed brief pointer and renders its busy line" + +grep -aFq 'Agent cancelled' "$TRANSCRIPT" \ + || fail "real rovo did not print 'Agent cancelled' on a mid-tool-call Escape" +pass "real rovo's session prints 'Agent cancelled' and survives a mid-tool-call Escape" + +grep -aFq 'resume your' "$TRANSCRIPT" \ + || fail "real rovo did not print its resume hint after /exit" +pass "real rovo exits cleanly on /exit" + +# --- allowedExternalPaths: prove the standard crewmate flow's brief-read and +# status/report-write against real files OUTSIDE the worktree, mirroring +# fm-spawn.sh's rovo_config_override_flag grant, and that the same access is +# blocked by default (AGENTS.md fm-rovo-external-paths finding). rovo's +# --config-override must be set at launch; there is no live escalation once +# the process is running. +EXT_DATA="$LAB/external/data/task" +EXT_STATE="$LAB/external/state" +mkdir -p "$EXT_DATA" "$EXT_STATE" +EXT_DATA_REAL=$(cd "$EXT_DATA" && pwd -P) || fail "could not resolve the external brief directory" +EXT_STATE_REAL=$(cd "$EXT_STATE" && pwd -P) || fail "could not resolve the external state directory" +EXT_STATUS="$EXT_STATE_REAL/task.status" +printf 'existing: seed line\n' > "$EXT_STATUS" +EXT_TOKEN="ROVO_EXT_WRITE_$$_$RANDOM" +cat > "$EXT_DATA_REAL/brief.md" <<EOF +Append the exact line "working: $EXT_TOKEN" to the file $EXT_STATUS, +preserving its existing content exactly as-is, then reply with exactly the +word CONFIRMED and nothing else. +EOF +CONFIG_OVERRIDE_JSON='{"toolPermissions":{"allowedExternalPaths":["'"$EXT_DATA_REAL"'","'"$EXT_STATUS"'"]}}' + +mkdir -p "$LAB/workspace2" +git -C "$LAB/workspace2" init -q || fail "could not initialize the second isolated workspace" +WORKSPACE2=$(cd "$LAB/workspace2" && pwd -P) || fail "could not resolve the second isolated workspace" +TRANSCRIPT2="$LAB/transcript2.log" + +python3 - "$ROVO_BIN" "$WORKSPACE2" "$TRANSCRIPT2" "$EXT_DATA_REAL/brief.md" "$CONFIG_OVERRIDE_JSON" <<'PY' || fail "the allowedExternalPaths grant PTY driver reported a failure" +import os +import pty +import select +import subprocess +import sys +import time + +rovo_bin, workspace, transcript_path, brief_real, config_override = sys.argv[1:6] + +pointer = "Read the brief at %s and follow it exactly." % brief_real + +pid, fd = pty.fork() +if pid == 0: + os.chdir(workspace) + os.execvp(rovo_bin, [rovo_bin, "run", "--yolo", "--config-override", config_override]) + os._exit(127) + +transcript = open(transcript_path, "wb") + +def pump(timeout, want=None): + deadline = time.time() + timeout + buf = b"" + while time.time() < deadline: + r, _, _ = select.select([fd], [], [], 0.5) + if fd in r: + try: + chunk = os.read(fd, 65536) + except OSError: + break + if not chunk: + break + transcript.write(chunk) + transcript.flush() + buf += chunk + if want and want in buf: + return buf + return buf + +ready = pump(60, want=b"Welcome to Rovo!") +if b"Welcome to Rovo!" not in ready: + sys.exit("rovo (grant session) never rendered its 'Welcome to Rovo!' readiness banner") + +os.write(fd, pointer.encode() + b"\r") +reply = pump(90, want=b"CONFIRMED") +if b"CONFIRMED" not in reply: + sys.exit("rovo (grant session) never confirmed the external brief-read/status-write") + +time.sleep(1) +os.write(fd, b"/exit\r") +for _ in range(60): + try: + done_pid, status = os.waitpid(pid, os.WNOHANG) + except ChildProcessError: + done_pid = pid + status = 0 + if done_pid == pid: + break + pump(1) +else: + subprocess.run(["kill", "-9", str(pid)]) + sys.exit("rovo (grant session) did not exit after /exit") + +transcript.close() +PY + +grep -aFq 'CONFIRMED' "$TRANSCRIPT2" \ + || fail "real rovo with the allowedExternalPaths grant never confirmed the external brief-read/status-write" +grep -aFq "working: $EXT_TOKEN" "$EXT_STATUS" \ + || fail "real rovo with the grant did not actually append to the external status file" +grep -aFq 'existing: seed line' "$EXT_STATUS" \ + || fail "real rovo with the grant clobbered the external status file's existing content instead of appending" +pass "real rovo with the allowedExternalPaths grant reads an external brief and appends to an external status file" + +# Negative control: the exact same shape, launched WITHOUT the grant, proves +# the standard flow is genuinely blocked by default rather than merely +# untested - the concrete problem this fix closes. +NOGRANT_TOKEN="ROVO_NOGRANT_$$_$RANDOM" +cat > "$EXT_DATA_REAL/brief-nogrant.md" <<EOF +Append the exact line "working: $NOGRANT_TOKEN" to the file $EXT_STATUS, +preserving its existing content exactly as-is, then reply with exactly the +word CONFIRMED and nothing else. +EOF + +mkdir -p "$LAB/workspace3" +git -C "$LAB/workspace3" init -q || fail "could not initialize the third isolated workspace" +WORKSPACE3=$(cd "$LAB/workspace3" && pwd -P) || fail "could not resolve the third isolated workspace" +TRANSCRIPT3="$LAB/transcript3.log" + +python3 - "$ROVO_BIN" "$WORKSPACE3" "$TRANSCRIPT3" "$EXT_DATA_REAL/brief-nogrant.md" <<'PY' || fail "the no-grant control PTY driver reported a failure" +import os +import pty +import select +import subprocess +import sys +import time + +rovo_bin, workspace, transcript_path, brief_real = sys.argv[1:5] + +pointer = "Read the brief at %s and follow it exactly." % brief_real + +pid, fd = pty.fork() +if pid == 0: + os.chdir(workspace) + os.execvp(rovo_bin, [rovo_bin, "run", "--yolo"]) + os._exit(127) + +transcript = open(transcript_path, "wb") + +def pump(timeout, want=None): + deadline = time.time() + timeout + buf = b"" + while time.time() < deadline: + r, _, _ = select.select([fd], [], [], 0.5) + if fd in r: + try: + chunk = os.read(fd, 65536) + except OSError: + break + if not chunk: + break + transcript.write(chunk) + transcript.flush() + buf += chunk + if want and want in buf: + return buf + return buf + +ready = pump(60, want=b"Welcome to Rovo!") +if b"Welcome to Rovo!" not in ready: + sys.exit("rovo (no-grant session) never rendered its 'Welcome to Rovo!' readiness banner") + +os.write(fd, pointer.encode() + b"\r") +# No grant: rovo cannot read the external brief at all, so it settles back to +# an idle composer instead of ever confirming. Wait long enough for the +# refusal to render, then move on regardless. +pump(30) + +time.sleep(1) +os.write(fd, b"/exit\r") +for _ in range(60): + try: + done_pid, status = os.waitpid(pid, os.WNOHANG) + except ChildProcessError: + done_pid = pid + status = 0 + if done_pid == pid: + break + pump(1) +else: + subprocess.run(["kill", "-9", str(pid)]) + sys.exit("rovo (no-grant session) did not exit after /exit") + +transcript.close() +PY + +grep -aFq "working: $NOGRANT_TOKEN" "$EXT_STATUS" \ + && fail "real rovo without the allowedExternalPaths grant still wrote to the external status file - the confinement this fix relies on is gone" +grep -aiFq 'outside the' "$TRANSCRIPT3" \ + || fail "real rovo without the grant did not visibly refuse the external brief/status access" +pass "real rovo without the allowedExternalPaths grant cannot read the external brief or write the external status file" diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index 97601351de4..19f03ed68f1 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -769,7 +769,7 @@ test_spawn_secondmate_harness_model_token() { [ "$(meta_field "$meta" model)" = opus ] || fail "model-token: meta model not opus (got '$(meta_field "$meta" model)')" [ "$(meta_field "$meta" effort)" = default ] || fail "model-token: meta effort not default (got '$(meta_field "$meta" effort)')" launch=$(cat "$launchlog") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'opus'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus'" \ "model-token: launch did not carry --model opus" assert_not_contains "$launch" "--effort" "model-token: launch must not carry an --effort flag" pass "C3 spawn: config/secondmate-harness's model token threads --model into the launch and meta" @@ -791,7 +791,7 @@ test_spawn_secondmate_harness_model_and_effort_tokens() { [ "$(meta_field "$meta" model)" = opus ] || fail "model-effort-tokens: meta model not opus" [ "$(meta_field "$meta" effort)" = high ] || fail "model-effort-tokens: meta effort not high (got '$(meta_field "$meta" effort)')" launch=$(cat "$launchlog") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'opus' --effort 'high'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'opus' --effort 'high'" \ "model-effort-tokens: launch did not carry both --model opus and --effort high" pass "C4 spawn: config/secondmate-harness's model+effort tokens thread into the launch and meta" } @@ -1065,6 +1065,10 @@ if [ -n "${FM_FAKE_TMUX_LOG:-}" ]; then printf '%s\n' "$*" >> "$FM_FAKE_TMUX_LOG" fi case "$*" in + list-windows*) + sed -n 's/^window=[^:]*://p' "${FM_HOME:?}"/state/*.meta + exit 0 + ;; *display-message*'#{pane_current_command}'*) printf '%s\n' codex; exit 0 ;; *display-message*'#{pane_id}'*) printf '%s\n' '%1'; exit 0 ;; *display-message*'#{cursor_y}'*) printf '%s\n' 0; exit 0 ;; diff --git a/tests/fm-secondmate-lifecycle-e2e.test.sh b/tests/fm-secondmate-lifecycle-e2e.test.sh index 0bb9b44f2fa..bec284dbef1 100755 --- a/tests/fm-secondmate-lifecycle-e2e.test.sh +++ b/tests/fm-secondmate-lifecycle-e2e.test.sh @@ -137,8 +137,10 @@ phase_send() { : > "$LOG" printf '❯\n' > "$PANE" # The meta window (firstmate:fm-design) must win over a foreign same-named - # window returned by list-windows. - PATH="$FAKEBIN:$PATH" FM_HOME="$HOME_DIR" FM_FAKE_TMUX_WINDOW="other-session:fm-design" \ + # window returned by list-windows. Include the recorded endpoint in the fake + # inventory so the recovery-grade liveness check can verify it exists. + PATH="$FAKEBIN:$PATH" FM_HOME="$HOME_DIR" FM_FAKE_TMUX_WINDOW="firstmate:fm-design +other-session:fm-design" \ FM_FAKE_TMUX_LOG="$LOG" FM_FAKE_TMUX_CAPTURE="$PANE" \ "$ROOT/bin/fm-send.sh" fm-design 'route this work' >/dev/null 2>&1 \ || fail "fm-send failed for a bare firstmate window with home metadata" @@ -150,7 +152,7 @@ phase_send() { body=$(bash -c '. "$1"; fm_task_inbox_body "$2"' _ "$ROOT/bin/fm-task-inbox-lib.sh" "$record") assert_contains "$body" '[fm-from-firstmate]' "the inbox request was not marked as from-firstmate" assert_contains "$body" 'route this work' "the original request text did not survive the marker" - assert_grep 'send-keys -t firstmate:fm-design -l Firstmate instruction waiting:' "$LOG" "send did not ring the window recorded in this home's meta" + assert_grep 'send-keys -t firstmate:fm-design -l : Firstmate instruction waiting:' "$LOG" "send did not ring the window recorded in this home's meta" assert_no_grep 'route this work' "$LOG" "send typed the payload instead of only the doorbell" assert_no_grep 'send-keys -t other-session:fm-design' "$LOG" "send targeted a foreign same-named window" pass "send: a bare fm-<id> secondmate enqueues a marked request and rings the meta window" diff --git a/tests/fm-secondmate-reconcile.test.sh b/tests/fm-secondmate-reconcile.test.sh index 7cc148b2dff..0d33ccd5092 100755 --- a/tests/fm-secondmate-reconcile.test.sh +++ b/tests/fm-secondmate-reconcile.test.sh @@ -545,6 +545,7 @@ SH while [ ! -f "$signal" ]; do sleep 0.01; done ( + # shellcheck source=/dev/null . "$ROOT/bin/fm-wake-lib.sh" fm_lock_acquire_wait "$home/state/.control-mate.lock" fm_lock_acquire_wait "$home/state/.meta-mate.lock" @@ -563,6 +564,7 @@ META fm_lock_release "$home/state/.control-mate.lock" : > "$lifecycle_done" ) & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. lifecycle_pid=$! sleep 0.1 @@ -663,12 +665,14 @@ SH rc=0 FM_RECONCILE_RACE_SIGNAL="$signal" FM_RECONCILE_RACE_RELEASE="$release" \ run_remote_notify "$home" "$fakebin" "$snap" > "$home/notify.out" 2>&1 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. notify_pid=$! while [ ! -f "$signal" ]; do kill -0 "$notify_pid" 2>/dev/null || fail "reconcile exited before entering fm-send" sleep 0.01 done + # shellcheck source=/dev/null . "$ROOT/bin/fm-wake-lib.sh" fm_lock_acquire_wait "$home/state/.control-remote-send-race-mate.lock" fm_lock_acquire_wait "$home/state/.meta-remote-send-race-mate.lock" diff --git a/tests/fm-secondmate-restart.test.sh b/tests/fm-secondmate-restart.test.sh index a514c5430ce..8f764de2505 100755 --- a/tests/fm-secondmate-restart.test.sh +++ b/tests/fm-secondmate-restart.test.sh @@ -68,10 +68,10 @@ case "${1:-}" in if [ -e "$D/remote-relaunch-start" ] && [ ! -e "$D/remote-relaunch-end" ]; then : > "$D/local-relaunch-during-remote" fi - printf 'zsh' > "$D/command" + printf 'zsh' > "$D/command.$target" ;; - *'encode launch-brief'*) cat "$D/becomes" > "$D/command" ;; - 'Firstmate instruction waiting: list '*) + *'encode launch-brief'*) cat "$D/becomes" > "$D/command.$target" ;; + ': Firstmate instruction waiting: list '*) printf 'doorbell\n' >> "$D/rings" if [ -x "$D/on-doorbell" ]; then "$D/on-doorbell" "$payload" @@ -95,12 +95,18 @@ case "${1:-}" in fi exit 0 ;; display-message) + target= + prev= for a in "$@"; do + if [ "$prev" = -t ]; then target=$a; fi case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; - *pane_current_command*) cat "$D/command"; printf '\n'; exit 0 ;; + *pane_current_command*) + if [ -f "$D/command.$target" ]; then cat "$D/command.$target"; else cat "$D/command"; fi + printf '\n'; exit 0 ;; *pane_current_path*) cat "$D/cwd"; printf '\n'; exit 0 ;; esac + prev=$a done printf 'fakepane\n'; exit 0 ;; capture-pane) printf '> \n'; exit 0 ;; @@ -159,7 +165,7 @@ add_local_mate() { echo "home=$smhome" [ -z "$backend" ] || echo "backend=$backend" } > "$home/state/$id.meta" - printf '%s\n' "fm-$id" > "$dir/fake/windows" + printf '%s\n' "fm-$id" >> "$dir/fake/windows" printf '%s' "$smhome" > "$dir/fake/cwd" } @@ -284,7 +290,7 @@ test_persist_precedes_restart() { assert_contains "$out" "summary: 1 of 1 restarted, 0 nudged, 0 unreached" "the summary should report the reload" # The pane transcript orders the two phases: the instruction doorbell first, # the harness exit command only after it. - doorbell_line=$(grep -n '^Firstmate instruction waiting: ' "$dir/fake/literal" | head -1 | cut -d: -f1) + doorbell_line=$(grep -n '^: Firstmate instruction waiting: ' "$dir/fake/literal" | head -1 | cut -d: -f1) exit_line=$(grep -n '^/exit$' "$dir/fake/literal" | head -1 | cut -d: -f1) [ -n "$doorbell_line" ] || fail "the persist request never reached the mate" [ -n "$exit_line" ] || fail "the mate was never stopped, so it was not restarted" @@ -310,6 +316,43 @@ test_arrived_answer_precedes_deadline_check() { pass "T2b an arrived persist answer is resolved before timeout" } +# --- T2c: an answer arriving between resolution and timeout wins ------------- +test_answer_between_resolution_and_timeout_wins() { + local dir out rc + dir=$(new_case answer-at-timeout-decision) + add_local_mate "$dir" sm1 + + # Delay the modelled answer until the first resolution attempt has completed + # its unsuccessful status scan. The real pending-reply machinery publishes + # that scan signature with mv; this wrapper appends the correlated answer only + # after that publication, reproducing the boundary race deterministically. + cat > "$dir/fakebin/mv" <<'SH' +#!/usr/bin/env bash +set -u +/bin/mv "$@" || exit $? +target=${!#} +case "$target" in + "${FM_FAKE_DIR%/fake}"/home/state/pending-replies/*) + if [ ! -e "$FM_FAKE_DIR/answer-after-scan" ] \ + && grep -q '^parent_status_scan_signature=.' "$target"; then + : > "$FM_FAKE_DIR/answer-after-scan" + corr=${target##*/} + status=$(sed -n 's/^parent_status=//p' "$target") + printf 'done [corr=%s]: open records written down\n' "$corr" >> "$status" + fi + ;; +esac +SH + chmod +x "$dir/fakebin/mv" + + out=$(FM_TEST_PERSIST_WAIT=0 run_restart "$dir" sm1); rc=$? + + expect_code 0 "$rc" "an answer already on disk at the timeout decision must release the gate"$'\n'"$out" + assert_contains "$out" "restarted: sm1" "the reply that raced the timeout was ignored" + assert_not_contains "$out" "nudged: sm1" "a confirmed mate must not take the timeout fallback" + pass "T2c a reply between the preliminary scan and timeout decision wins" +} + # --- T3: a runtime that cannot prove a restart never gets one ---------------- test_unprovable_runtime_falls_back() { local dir out rc @@ -553,7 +596,7 @@ test_persist_waits_are_polled_together() { expect_code 3 "$rc" "the unanswered mate should fall back after the confirmed mate restarts"$'\n'"$out" exit_line=$(grep -n '^/exit$' "$dir/fake/literal" | head -1 | cut -d: -f1) - nudge_line=$(grep -n '^Firstmate instruction waiting: ' "$dir/fake/literal" | tail -1 | cut -d: -f1) + nudge_line=$(grep -n '^: Firstmate instruction waiting: ' "$dir/fake/literal" | tail -1 | cut -d: -f1) [ -n "$exit_line" ] && [ -n "$nudge_line" ] && [ "$exit_line" -lt "$nudge_line" ] \ || fail "the first mate's timeout held the confirmed second mate behind it: $out" pass "T10 pending persist answers are polled as one fleet" @@ -712,7 +755,7 @@ test_already_current_mate_restarts_end_to_end() { assert_contains "$out" "summary: 1 of 1 restarted, 0 nudged, 0 unreached" \ "the pass must report the reload it performed" # Persist strictly before replace, read off the pane transcript. - doorbell_line=$(grep -n '^Firstmate instruction waiting: ' "$dir/fake/literal" | head -1 | cut -d: -f1) + doorbell_line=$(grep -n '^: Firstmate instruction waiting: ' "$dir/fake/literal" | head -1 | cut -d: -f1) exit_line=$(grep -n '^/exit$' "$dir/fake/literal" | head -1 | cut -d: -f1) [ -n "$doorbell_line" ] || fail "the persist request never reached the already-current mate" [ -n "$exit_line" ] || fail "the already-current mate was never stopped, so it was not restarted" @@ -764,6 +807,7 @@ test_already_current_unprovable_mate_stays_on_the_nudge_path() { test_persist_gates_and_asks_only_for_open_records test_persist_precedes_restart test_arrived_answer_precedes_deadline_check +test_answer_between_resolution_and_timeout_wins test_unprovable_runtime_falls_back test_unknown_mate_is_accounted_for test_refused_restart_falls_back_without_claiming_a_reload diff --git a/tests/fm-secondmate-safety.test.sh b/tests/fm-secondmate-safety.test.sh index 9b97b21e568..7a69e15fe86 100755 --- a/tests/fm-secondmate-safety.test.sh +++ b/tests/fm-secondmate-safety.test.sh @@ -1936,15 +1936,71 @@ EOF pass "secondmate force teardown discards child work" } +test_secondmate_force_teardown_refuses_duplicated_child_slot() { + local home subhome childproj childwt fakebin log err rc + home="$TMP_ROOT/force-duplicate-slot-home" + subhome="$TMP_ROOT/force-duplicate-slot-subhome" + childproj="$subhome/projects/alpha" + childwt="$TMP_ROOT/force-duplicate-slot-pool/1/alpha" + err="$TMP_ROOT/force-duplicate-slot.err" + mkdir -p "$home/state" "$home/data" "$subhome/state" "$(dirname "$childwt")" + fm_git_worktree "$childproj" "$childwt" duplicate-child + printf '{"worktrees":[{"name":"1","path":"%s"}]}\n' "$childwt" \ + > "$TMP_ROOT/force-duplicate-slot-pool/treehouse-state.json" + printf 'domain\n' > "$subhome/.fm-secondmate-home" + cat > "$home/state/domain.meta" <<EOF +window=firstmate:fm-domain +worktree=$subhome +project=$subhome +harness=echo +kind=secondmate +mode=secondmate +yolo=off +home=$subhome +projects=alpha +EOF + printf '%s\n' '- domain - design domain (home: '"$subhome"'; scope: design domain; projects: alpha; added 2026-06-22)' > "$home/data/secondmates.md" + for child in stale-child live-child; do + cat > "$subhome/state/$child.meta" <<EOF +window=firstmate:fm-$child +worktree=$childwt +project=$childproj +harness=echo +kind=ship +mode=no-mistakes +yolo=off +EOF + done + fakebin=$(make_fake_tmux "$TMP_ROOT/force-duplicate-slot-fake") + log="$TMP_ROOT/force-duplicate-slot-fake/tmux.log" + + set +e + PATH="$fakebin:$PATH" FM_HOME="$home" FM_FAKE_TMUX_LOG="$log" \ + FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/force-duplicate-slot-fake/pane.txt" \ + "$ROOT/bin/fm-teardown.sh" domain --force >/dev/null 2>"$err" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "forced secondmate teardown returned a duplicated child slot" + [ -d "$childwt" ] || fail "forced secondmate teardown removed the duplicated child slot" + [ -e "$subhome/state/stale-child.meta" ] || fail "forced secondmate teardown removed the stale child record" + [ -e "$subhome/state/live-child.meta" ] || fail "forced secondmate teardown removed the live child record" + grep -F 'kill-window' "$log" >/dev/null && fail "forced secondmate teardown killed a child before detecting its slot collision" + grep -F 'live-child' "$err" >/dev/null || grep -F 'stale-child' "$err" >/dev/null \ + || fail "forced secondmate teardown did not identify the duplicated child slot" + pass "forced secondmate teardown refuses duplicated descendant pool slots" +} + test_secondmate_force_teardown_preserves_child_on_unproven_lock() { local home subhome childproj childwt fakebin log err rc lock home="$TMP_ROOT/force-lock-home" subhome="$TMP_ROOT/force-lock-subhome" childproj="$subhome/projects/alpha" - childwt="$TMP_ROOT/force-lock-child-worktree" + childwt="$TMP_ROOT/force-lock-child-pool/1/alpha" err="$TMP_ROOT/force-lock-child.err" - mkdir -p "$home/state" "$home/data" "$subhome/state" + mkdir -p "$home/state" "$home/data" "$subhome/state" "$(dirname "$childwt")" fm_git_worktree "$childproj" "$childwt" force-child-lock + printf '{"worktrees":[{"name":"1","path":"%s"}]}\n' "$childwt" \ + > "$TMP_ROOT/force-lock-child-pool/treehouse-state.json" printf 'domain\n' > "$subhome/.fm-secondmate-home" cat > "$home/state/domain.meta" <<EOF window=firstmate:fm-domain @@ -2330,6 +2386,7 @@ EOF task_set_lock_path() { # <state-dir> local state=$1 + # shellcheck source=/dev/null ( . "$ROOT/bin/fm-wake-lib.sh"; fm_task_set_lock_path "$state" ) } @@ -2350,6 +2407,7 @@ hold_task_set_lock() { # <state-dir> -> echoes "<holder-pid> <lock-path>" fm_lock_try_acquire "$lock" || exit 1 sleep 30 ) >/dev/null 2>&1 & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. holder=$! while [ ! -e "$lock" ] && [ "$i" -lt 100 ]; do sleep 0.1 @@ -2464,6 +2522,7 @@ SH XDG_STATE_HOME="$TMP_ROOT/taskset-state-absent-xdg" \ FM_TASK_SET_TEST_READY="$ready" FM_TASK_SET_TEST_RELEASE="$release" \ "$ROOT/bin/fm-teardown.sh" domain --force >/dev/null 2>"$err" & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. pid=$! while [ ! -e "$ready" ] && kill -0 "$pid" 2>/dev/null && [ "$i" -lt 200 ]; do sleep 0.05 @@ -2737,6 +2796,7 @@ EOF out="$TMP_ROOT/watch-fake/watch.out" PATH="$fakebin:$PATH" FM_HOME="$home" FM_FAKE_TMUX_WINDOW="$window" FM_FAKE_TMUX_LOG="$TMP_ROOT/watch-fake/tmux.log" FM_FAKE_TMUX_CAPTURE="$TMP_ROOT/watch-fake/pane.txt" \ FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$ROOT/bin/fm-watch.sh" > "$out" & + # shellcheck disable=SC2031 # The background PID is captured immediately in this shell. pid=$! if ! wait_live "$pid" 25; then wait "$pid" || true @@ -2950,6 +3010,7 @@ test_secondmate_force_teardown_preserves_nested_restore_status test_secondmate_teardown_refuses_failed_leased_home_return test_secondmate_teardown_removes_plain_clone_home_without_treehouse_return test_secondmate_force_teardown_discards_child_work +test_secondmate_force_teardown_refuses_duplicated_child_slot test_secondmate_force_teardown_preserves_child_on_unproven_lock test_secondmate_force_teardown_allows_non_state_operational_dir_symlinks_inside_home test_secondmate_force_teardown_refuses_operational_dir_symlink_outside_home diff --git a/tests/fm-send-inbox-doorbell-live-e2e.test.sh b/tests/fm-send-inbox-doorbell-live-e2e.test.sh index e6a5d696fe7..36de34f56b4 100644 --- a/tests/fm-send-inbox-doorbell-live-e2e.test.sh +++ b/tests/fm-send-inbox-doorbell-live-e2e.test.sh @@ -28,14 +28,13 @@ # unready state and correctly fails that harness's check. set -u +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" -if [ "${FM_SEND_INBOX_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SEND_INBOX_LIVE_E2E=1 to run the live steering-inbox doorbell guard" - exit 0 -fi +fm_live_gate opt-in FM_SEND_INBOX_LIVE_E2E tmux -command -v tmux >/dev/null 2>&1 || { echo "not ok - FM_SEND_INBOX_LIVE_E2E=1 but tmux is not installed" >&2; exit 1; } unset NO_MISTAKES_GATE SOCKET="fm-inbox-live-$$" diff --git a/tests/fm-send-inbox.test.sh b/tests/fm-send-inbox.test.sh index 783e489e1ab..0046f4c149c 100644 --- a/tests/fm-send-inbox.test.sh +++ b/tests/fm-send-inbox.test.sh @@ -72,7 +72,7 @@ case "${1:-}" in printf '╭────╮\n│ │\n╰────╯\n' fi exit 0 ;; - list-windows) exit 0 ;; + list-windows) printf 'fm-t1\n'; exit 0 ;; esac exit 0 SH @@ -124,7 +124,7 @@ test_text_steer_rides_inbox() { body=$(record_body _ "$rec") [ "$body" = "please rebase onto main" ] || fail "the recorded body differs: $body" typed=$(cat "$dir/send.log") - assert_contains "$typed" "Firstmate instruction waiting: list $dir/home/state/t1.inbox/*.msg" \ + assert_contains "$typed" "Firstmate instruction waiting: list '$dir/home/state/t1.inbox'/*.msg" \ "the doorbell should direct the worker to drain the inbox" case "$typed" in *"please rebase onto main"*) fail "the payload must never be typed:"$'\n'"$typed" ;; diff --git a/tests/fm-send-popup-settle.test.sh b/tests/fm-send-popup-settle.test.sh index 8860c6bfa27..3cbbab24b42 100755 --- a/tests/fm-send-popup-settle.test.sh +++ b/tests/fm-send-popup-settle.test.sh @@ -57,7 +57,7 @@ case "${1:-}" in for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done printf 'fakepane\n'; exit 0 ;; capture-pane) printf '╭────╮\n│ │\n╰────╯\n'; exit 0 ;; - list-windows) exit 0 ;; + list-windows) printf 'win\n'; exit 0 ;; esac exit 0 SH diff --git a/tests/fm-send-resolve-key.test.sh b/tests/fm-send-resolve-key.test.sh index f395504559c..fc51a123a55 100755 --- a/tests/fm-send-resolve-key.test.sh +++ b/tests/fm-send-resolve-key.test.sh @@ -72,7 +72,9 @@ case "${1:-}" in for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done printf 'fakepane\n'; exit 0 ;; capture-pane) printf '╭────╮\n│ │\n╰────╯\n'; exit 0 ;; - list-windows) exit 0 ;; + list-windows) + printf '%s\n' fm-t1 fm-t2 fm-t3 fm-t4 fm-t5 fm-t6 fm-t7 fm-t8 fm-t9 fm-mate + exit 0 ;; esac exit 0 SH diff --git a/tests/fm-send-secondmate-marker-herdr-e2e.test.sh b/tests/fm-send-secondmate-marker-herdr-e2e.test.sh index 5e28b9aaf8d..bdddfa2334e 100755 --- a/tests/fm-send-secondmate-marker-herdr-e2e.test.sh +++ b/tests/fm-send-secondmate-marker-herdr-e2e.test.sh @@ -22,14 +22,7 @@ set -u # shellcheck source=/dev/null . "$ROOT/bin/fm-backend.sh" -if [ "${FM_SEND_MARKER_HERDR_E2E:-0}" != 1 ]; then - echo "skip: set FM_SEND_MARKER_HERDR_E2E=1 to run the real Pi/Herdr secondmate-marker regression" - exit 0 -fi - -for tool in git herdr jq pi; do - command -v "$tool" >/dev/null 2>&1 || { echo "skip: $tool not found"; exit 0; } -done +fm_live_gate opt-in FM_SEND_MARKER_HERDR_E2E git herdr jq pi LAB_HELPER=${HERDR_LAB_HELPER:-$ROOT/bin/fm-herdr-lab.sh} SESSION=$("$LAB_HELPER" name fm-send-secondmate-marker-v7) diff --git a/tests/fm-send-strict.test.sh b/tests/fm-send-strict.test.sh index 216e281368b..82fe77fdff2 100755 --- a/tests/fm-send-strict.test.sh +++ b/tests/fm-send-strict.test.sh @@ -68,7 +68,7 @@ case "${1:-}" in printf '╭────╮\n│ │\n╰────╯\n' exit 0 ;; list-windows) - printf 'foreign:%s\n' "${FM_FAKE_TMUX_WINDOW:-fm-lost}" + printf 'foreign:%s\nfm-mpf-lane-m8\nfm-lane-ok\n' "${FM_FAKE_TMUX_WINDOW:-fm-lost}" exit 0 ;; esac exit 0 @@ -109,7 +109,7 @@ test_exact_lane_id_send_still_works() { "$SEND" mpf-lane-m8 "lost dispatch" >/dev/null 2>"$err"; rc=$? expect_code 0 "$rc" "exact task id send should succeed when metadata exists" got=$(cat "$log") - assert_contains "$got" "target=sess:fm-mpf-lane-m8 literal=1 arg=Firstmate instruction waiting" \ + assert_contains "$got" "target=sess:fm-mpf-lane-m8 literal=1 arg=: Firstmate instruction waiting" \ "exact id should ring the doorbell at the meta target" assert_contains "$got" "target=sess:fm-mpf-lane-m8 literal=0 arg=Enter" "exact id should submit the doorbell with Enter" grep -qF 'lost dispatch' "$home/state/mpf-lane-m8.inbox/001.msg" \ @@ -202,7 +202,7 @@ test_healthy_fm_id_send_still_works() { "$SEND" fm-lane-ok "hello captain" >/dev/null 2>"$err"; rc=$? expect_code 0 "$rc" "healthy fm-id send should succeed" got=$(cat "$log") - assert_contains "$got" "target=sess:fm-lane-ok literal=1 arg=Firstmate instruction waiting" \ + assert_contains "$got" "target=sess:fm-lane-ok literal=1 arg=: Firstmate instruction waiting" \ "healthy send should ring the doorbell at the meta target" assert_contains "$got" "target=sess:fm-lane-ok literal=0 arg=Enter" "healthy send should submit the doorbell with Enter" grep -qF 'hello captain' "$home/state/lane-ok.inbox/001.msg" \ diff --git a/tests/fm-session-start.test.sh b/tests/fm-session-start.test.sh index 24b0d67a50a..d62157274f6 100755 --- a/tests/fm-session-start.test.sh +++ b/tests/fm-session-start.test.sh @@ -293,7 +293,7 @@ SH } make_fake_ps_pi_holder() { - local fakebin=$1 holder_pid=$2 + local fakebin=$1 holder_pid=$2 harness=${3:-pi} cat > "$fakebin/ps" <<SH #!/usr/bin/env bash set -u @@ -306,7 +306,7 @@ done case "\$*" in *"comm="*) if [ "\$pid" = "$holder_pid" ]; then - printf '/usr/local/bin/pi\n' + printf '/usr/local/bin/$harness\n' else printf '/bin/zsh\n' fi @@ -314,7 +314,7 @@ case "\$*" in ;; *"args="*) if [ "\$pid" = "$holder_pid" ]; then - printf 'pi\n' + printf '$harness\n' else printf 'zsh\n' fi @@ -777,6 +777,21 @@ write_pi_loaded_markers() { write_pi_turnend_loaded_marker "$home" "$root" "$pid" } +install_omp_extension_fixtures() { + local root=$1 + mkdir -p "$root/.omp/extensions" + cp "$ROOT/.omp/extensions/fm-primary-omp-watch.ts" "$root/.omp/extensions/fm-primary-omp-watch.ts" + cp "$ROOT/.omp/extensions/fm-primary-turnend-guard.ts" "$root/.omp/extensions/fm-primary-turnend-guard.ts" +} + +write_omp_loaded_markers() { + local home=$1 root=$2 pid=$3 version + version=$(hash_file_for_test "$root/.omp/extensions/fm-primary-omp-watch.ts") + printf '%s\n%s\n' "$version" "$pid" > "$home/state/.omp-watch-extension-loaded" + version=$(hash_file_for_test "$root/.omp/extensions/fm-primary-turnend-guard.ts") + printf '%s\n%s\n' "$version" "$pid" > "$home/state/.omp-turnend-extension-loaded" +} + # --- context digest: absent vs empty vs present ----------------------------- test_context_digest_absent_empty_present() { @@ -1065,7 +1080,7 @@ EOF printf 'window=fm-sess:w1\nkind=ship\n' > "$home/state/task-a.meta" printf 'Captain memory that may be truncated away safely.\n' > "$home/data/captain.md" - out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + out=$(run_session_start "$home" "$root" "$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)") lock_line=$(printf '%s\n' "$out" | grep -n '^LOCK$' | head -1 | cut -d: -f1) boot_line=$(printf '%s\n' "$out" | grep -n '^BOOTSTRAP$' | head -1 | cut -d: -f1) @@ -1474,7 +1489,7 @@ EOF printf 'needs-decision: pick a library\n' > "$home/state/task-z.status" append_wake "$home/state" signal task-z.status "needs-decision: pick a library" - out=$(run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + out=$(run_session_start "$home" "$root" "$fakebin:$(fm_test_base_path_sans "$BASE_PATH" node)") # fm-lock.sh's own exact success text. assert_contains "$out" "lock acquired: harness pid" "fm-lock.sh's real output did not appear (composition, not reimplementation)" @@ -2709,6 +2724,51 @@ EOF pass "session start accepts current Pi markers written before lock acquisition" } +test_omp_supervision_block_and_diagnostic() { + local rec root home fakebin out block_count + rec=$(new_world omp-supervision-block) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + make_fake_ps_harness "$fakebin" omp + + out=$(FM_FAKE_HARNESS=omp run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + + block_count=$(printf '%s\n' "$out" | grep -c '^SUPERVISION OPERATING INSTRUCTIONS - primary harness:') + [ "$block_count" -eq 1 ] || fail "expected exactly one supervision block, got $block_count" + assert_contains "$out" "SUPERVISION OPERATING INSTRUCTIONS - primary harness: omp" "omp supervision block missing" + assert_contains "$out" "Mode: omp (Oh My Pi) extension background wake." "omp snippet missing from session start" + assert_contains "$out" "OMP_WATCH_EXTENSION: not loaded" "omp extension load diagnostic missing" + assert_contains "$out" "so $root/.omp/extensions/fm-primary-turnend-guard.ts and $root/.omp/extensions/fm-primary-omp-watch.ts auto-load" "omp diagnostic omits the two tracked extension paths" + assert_not_contains "$out" "PI_WATCH_EXTENSION" "omp primary must not receive the Pi diagnostic" + assert_not_contains "$out" "project trust" "omp diagnostic must not carry Pi's trust prerequisite" + pass "session start emits the omp block and reports omp extension load state" +} + +test_omp_diagnostic_accepts_prelock_loaded_marker() { + local rec root home fakebin out holder_pid + rec=$(new_world omp-prelock-loaded-marker) + IFS='|' read -r root home fakebin <<EOF +$rec +EOF + make_fake_toolchain "$fakebin" + + sleep 300 & + holder_pid=$! + make_fake_ps_pi_holder "$fakebin" "$holder_pid" omp + install_omp_extension_fixtures "$root" + write_omp_loaded_markers "$home" "$root" "$holder_pid" + + out=$(FM_FAKE_HARNESS=omp run_session_start "$home" "$root" "$fakebin:$BASE_PATH") + kill "$holder_pid" 2>/dev/null || true + wait "$holder_pid" 2>/dev/null || true + + assert_contains "$out" "primary harness: omp" "omp holder ancestry was not detected as omp" + assert_not_contains "$out" "OMP_WATCH_EXTENSION: not loaded" "omp diagnostic rejected a current pre-lock loaded marker" + pass "session start accepts current omp markers written before lock acquisition" +} + test_pi_diagnostic_rejects_missing_turnend_guard_marker() { local rec root home fakebin out holder_pid rec=$(new_world pi-missing-turnend-marker) @@ -2799,6 +2859,8 @@ test_supervision_block_exactly_one_and_pi_diagnostic test_pi_signed_primary_uses_pi_extensions_without_identity_normalization test_pi_diagnostic_rejects_stale_loaded_marker test_pi_diagnostic_accepts_prelock_loaded_marker +test_omp_supervision_block_and_diagnostic +test_omp_diagnostic_accepts_prelock_loaded_marker test_pi_diagnostic_rejects_missing_turnend_guard_marker test_pi_diagnostic_rejects_previous_session_loaded_marker test_runtime_bound_truncates_loudly_and_exits_zero diff --git a/tests/fm-sessionstart-hook-live-e2e.test.sh b/tests/fm-sessionstart-hook-live-e2e.test.sh index f15c9edb806..75428e970fa 100755 --- a/tests/fm-sessionstart-hook-live-e2e.test.sh +++ b/tests/fm-sessionstart-hook-live-e2e.test.sh @@ -40,11 +40,10 @@ # FM_PI_SESSIONSTART_RACE_LIVE_E2E=1 tests/fm-sessionstart-hook-live-e2e.test.sh set -u -if [ "${FM_SESSIONSTART_HOOK_LIVE_E2E:-0}" != 1 ] && \ - [ "${FM_PI_SESSIONSTART_RACE_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SESSIONSTART_HOOK_LIVE_E2E=1 for the cross-harness guard or FM_PI_SESSIONSTART_RACE_LIVE_E2E=1 for the offline Pi /new race regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_SESSIONSTART_HOOK_LIVE_E2E,FM_PI_SESSIONSTART_RACE_LIVE_E2E tmux ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" unset NO_MISTAKES_GATE @@ -56,8 +55,6 @@ fail() { pass() { printf 'ok - %s\n' "$1"; } note() { printf '# %s\n' "$1"; } -command -v tmux >/dev/null 2>&1 || fail "tmux not found; the context-reset checks drive real interactive harnesses" - # Outside the repo on purpose: each lab is its own git repo, and nesting one # inside the checkout would show up as an embedded repository in a working tree # a maintainer may be committing from while this guard runs. diff --git a/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh b/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh index 0ab68bc2cec..42819acc1bb 100755 --- a/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh +++ b/tests/fm-sessionstart-instruction-refresh-live-e2e.test.sh @@ -24,10 +24,10 @@ # This costs real Pi model turns and requires its normal authenticated profile. set -u -if [ "${FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E:-0}" != 1 ]; then - echo "skip: set FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E=1 to run the isolated real-Pi instruction-refresh regression" - exit 0 -fi +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +fm_live_gate opt-in FM_SESSIONSTART_INSTRUCTION_REFRESH_LIVE_E2E pi tmux git ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" TMUX_SOCKET="fm-sessionstart-instruction-refresh-$$" @@ -107,10 +107,6 @@ cleanup() { } trap cleanup EXIT INT TERM -command -v pi >/dev/null 2>&1 || fail "pi not found" -command -v tmux >/dev/null 2>&1 || fail "tmux not found" -command -v git >/dev/null 2>&1 || fail "git not found" - mkdir -p "$LAB" git clone --quiet --no-hardlinks "$ROOT" "$PROJECT" || fail "could not create isolated Firstmate checkout" git -C "$PROJECT" checkout -q -B main "$TEST_COMMIT" \ diff --git a/tests/fm-spawn-batch.test.sh b/tests/fm-spawn-batch.test.sh index 7f8311077a3..ab447397c0e 100755 --- a/tests/fm-spawn-batch.test.sh +++ b/tests/fm-spawn-batch.test.sh @@ -84,12 +84,14 @@ test_projects_path_scoping() { home="$TMP_ROOT/$id home" projects="$TMP_ROOT/$id projects" mkdir -p "$home/data" "$projects/alpha" + git -C "$projects/alpha" init -q || fail "$label: could not initialize project fixture" if [ "$use_override" = yes ]; then out=$(FM_ROOT_OVERRIDE='' FM_STATE_OVERRIDE='' FM_DATA_OVERRIDE='' FM_CONFIG_OVERRIDE='' \ FM_HOME="$home" FM_PROJECTS_OVERRIDE="$projects" FM_SPAWN_NO_GUARD=1 \ "$SPAWN" "$id" projects/alpha codex --mode no-mistakes --yolo off 2>&1) else mkdir -p "$home/projects/alpha" + git -C "$home/projects/alpha" init -q || fail "$label: could not initialize home project fixture" out=$(FM_ROOT_OVERRIDE='' FM_STATE_OVERRIDE='' FM_DATA_OVERRIDE='' FM_PROJECTS_OVERRIDE='' FM_CONFIG_OVERRIDE='' \ FM_HOME="$home" FM_SPAWN_NO_GUARD=1 \ "$SPAWN" "$id" projects/alpha codex --mode no-mistakes --yolo off 2>&1) diff --git a/tests/fm-spawn-dispatch-profile.test.sh b/tests/fm-spawn-dispatch-profile.test.sh index 3d9894e3e14..1717d1fe028 100755 --- a/tests/fm-spawn-dispatch-profile.test.sh +++ b/tests/fm-spawn-dispatch-profile.test.sh @@ -131,7 +131,7 @@ test_no_profile_keeps_claude_profile_defaults() { assert_meta_profile "$HOME_DIR/state/$id.meta" claude default default launch=$(cat "$LAUNCH_LOG") - expected="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI -u COPILOT_CLI -u COPILOT_AGENT_SESSION_ID -u COPILOT_LOADER_PID CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$HOME_DIR/data/$id/launch-brief.md')\"" + expected="env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI -u COPILOT_CLI -u COPILOT_AGENT_SESSION_ID -u COPILOT_LOADER_PID CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' \"\$('${ROOT}/bin/fm-operational-input.sh' encode launch-brief < '$HOME_DIR/data/$id/launch-brief.md')\"" [ "$launch" = "$expected" ] || fail "no-profile claude launch did not use the canonical launch kind"$'\n'"expected: $expected"$'\n'"actual: $launch" pass "no --model/--effort records defaults and types the claude launch instructions" } @@ -435,7 +435,7 @@ test_claude_threads_model_and_effort() { expect_code 0 "$status" "claude spawn with profile flags should succeed" assert_meta_profile "$HOME_DIR/state/$id.meta" claude sonnet high launch=$(cat "$LAUNCH_LOG") - assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}' --model 'sonnet' --effort 'high'" \ + assert_contains "$launch" "claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}' --model 'sonnet' --effort 'high'" \ "claude launch did not thread model and effort flags" assert_not_contains "$launch" "--tui-mode" "non-Pi launches must not receive Pi's TUI mode override" pass "claude receives --model and --effort profile flags" @@ -735,6 +735,8 @@ test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity() { sm="$CASE_DIR/secondmate-home" make_seeded_secondmate_home "$sm" "$id" sm=$(cd "$sm" && pwd -P) + cp "$ROOT/AGENTS.md" "$sm/AGENTS.md" + cp "$sm/data/charter.md" "$CASE_DIR/charter-before" out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$sm" --secondmate) status=$? @@ -742,9 +744,19 @@ test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity() { assert_contains "$out" "spawned $id harness=pi-signed kind=secondmate" \ "pi-signed secondmate spawn did not preserve its runtime identity" assert_meta_profile "$HOME_DIR/state/$id.meta" pi-signed default default + cmp -s "$ROOT/AGENTS.md" "$sm/AGENTS.md" || fail "secondmate launch rewrote the supervisor contract" + cmp -s "$CASE_DIR/charter-before" "$sm/data/charter.md" || fail "secondmate launch rewrote the charter" + assert_absent "$HOME_DIR/data/$id/launch-brief.md" "secondmate launch received a worker overlay" launch=$(cat "$LAUNCH_LOG") + assert_contains "$launch" "< '$sm/data/charter.md'" "secondmate launch lost its original charter" assert_contains "$launch" "FM_PI_HARNESS=pi-signed '$FAKEBIN_DIR/pi-signed' --tui-mode regular -e '$sm/.pi/extensions/fm-primary-turnend-guard.ts' -e '$sm/.pi/extensions/fm-primary-pi-watch.ts'" \ "pi-signed secondmate did not force the regular TUI with Pi's primary extension launch shape" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# evidence begin: persistent secondmate\n%s\n' "$out" + printf 'launch command:\n%s\noriginal charter:\n' "$launch" + cat "$sm/data/charter.md" + printf 'supervisor AGENTS.md and charter remain byte-identical; no worker overlay created\n# evidence end\n' + fi pass "pi-signed is a distinct persistent secondmate runtime with shared Pi supervision semantics" } @@ -812,7 +824,7 @@ test_claude_forwards_firstmate_config_dir_when_set() { status=$? expect_code 0 "$status" "claude spawn with CLAUDE_CONFIG_DIR set should succeed" launch=$(cat "$LAUNCH_LOG") - assert_contains "$launch" "CLAUDE_CONFIG_DIR='$CASE_DIR/claude-work' env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI -u COPILOT_CLI -u COPILOT_AGENT_SESSION_ID -u COPILOT_LOADER_PID CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\"}'" \ + assert_contains "$launch" "CLAUDE_CONFIG_DIR='$CASE_DIR/claude-work' env -u CURSOR_AGENT -u CURSOR_INVOKED_AS -u GEMINI_CLI -u COPILOT_CLI -u COPILOT_AGENT_SESSION_ID -u COPILOT_LOADER_PID CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false CLAUDE_CODE_SEND_FEEDBACK=0 claude --dangerously-skip-permissions --settings '{\"feedbackDrafts\":\"off\",\"attribution\":{\"commit\":\"\",\"pr\":\"\",\"sessionUrl\":false}}'" \ "claude launch did not forward firstmate's CLAUDE_CONFIG_DIR to the crewmate pane" pass "claude forwards firstmate's CLAUDE_CONFIG_DIR so the crewmate uses the same credential store" } @@ -850,6 +862,49 @@ test_non_claude_harness_ignores_config_dir() { pass "non-claude harnesses do not receive the claude CLAUDE_CONFIG_DIR prefix" } +# The captain's attribution policy lives in the `user` settings scope, which a +# spawned worker's settings sources are not guaranteed to load. Every claude +# launch must therefore carry the policy itself, or a spawned worker writes +# Co-Authored-By and Claude-Session trailers into commits and PR bodies. +assert_attribution_policy() { # <launch-command> <what> + local launch=$1 what=$2 + assert_contains "$launch" '"attribution":' "$what launch carries no attribution policy" + assert_contains "$launch" '"commit":""' "$what launch does not silence the commit trailer" + assert_contains "$launch" '"pr":""' "$what launch does not silence the PR-body attribution" + assert_contains "$launch" '"sessionUrl":false' "$what launch does not silence the session URL" +} + +test_claude_crewmate_launch_carries_the_attribution_policy() { + local rec id out status launch + id=profile-claude-attribution-z22 + rec=$(make_spawn_case profile-claude-attribution claude "$id") + read_case_record "$rec" + + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 0 "$status" "claude crewmate spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_attribution_policy "$launch" "claude crewmate" + pass "a claude crewmate launch carries the attribution-off policy in its own settings" +} + +test_claude_secondmate_launch_carries_the_attribution_policy() { + local rec id sm out status launch + id=profile-secondmate-attribution-z23 + rec=$(make_spawn_case profile-secondmate-attribution claude "$id") + read_case_record "$rec" + sm="$CASE_DIR/secondmate-home" + make_seeded_secondmate_home "$sm" "$id" + + out=$(FM_TEST_CLAUDE_CONFIG_DIR="$CASE_DIR/claude-work" \ + run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$sm" --secondmate) + status=$? + expect_code 0 "$status" "secondmate claude spawn should succeed"$'\n'"$out" + launch=$(cat "$LAUNCH_LOG") + assert_attribution_policy "$launch" "claude secondmate" + pass "a claude secondmate launch carries the attribution-off policy too" +} + test_active_dispatch_profile_does_not_block_secondmate_launch() { local rec id sm out status id=profile-secondmate-z16 @@ -868,38 +923,324 @@ test_active_dispatch_profile_does_not_block_secondmate_launch() { pass "active crew-dispatch profile does not block secondmate launches" } -test_no_profile_keeps_claude_profile_defaults -test_non_cursor_launch_clears_inherited_cursor_markers -test_copilot_threads_model_effort_and_hooks -test_relative_home_overrides_launch_with_absolute_cross_process_paths -test_home_defaults_preserve_absolute_or_resolve_relative_paths -test_absolute_override_spelling_is_preserved_in_launch_paths -test_unresolvable_relative_overrides_fail_loudly -test_active_dispatch_profile_requires_explicit_harness_for_ship -test_active_dispatch_profile_requires_explicit_harness_for_scout -test_active_dispatch_profile_allows_explicit_harness -test_active_dispatch_profile_allows_positional_harness -test_active_dispatch_profile_allows_raw_launch_command -test_claude_threads_model_and_effort -test_codex_threads_model_and_effort -test_codex_omits_invalid_max_effort -test_grok_threads_model_and_reasoning_effort -test_grok_omits_invalid_max_reasoning_effort -test_grok_omits_invalid_xhigh_reasoning_effort -test_cursor_threads_model_workspace_and_omits_effort_axis -test_cursor_refuses_model_absent_from_live_catalog -test_cursor_failed_catalog_probe_does_not_block_spawn -test_opencode_threads_model_and_ignores_effort_axis -test_pi_threads_model_and_max_effort -test_pi_tui_mode_probe_is_safe_for_old_and_new_pi -test_pi_signed_threads_shared_pi_profile_and_preserves_identity -test_pi_signed_missing_binary_refuses_before_endpoint_or_metadata -test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity -test_copilot_secondmate_uses_semantic_submission_hooks -test_batch_forwards_shared_profile_flags -test_claude_forwards_firstmate_config_dir_when_set -test_claude_omits_config_dir_prefix_when_unset -test_non_claude_harness_ignores_config_dir -test_active_dispatch_profile_does_not_block_secondmate_launch - -echo "# all fm-spawn-dispatch-profile tests passed" +# Execute the actual emitted command in a synthetic pane environment: the +# fake backend records delivery, while real shells exercise the env boundary. +# No developer environment or credential values are inspected by these probes. +test_launch_environment_allowlist() { + local setting rec id out status probe result expected launch value pane_shell pane_path + # shellcheck disable=SC2016 + value='synthetic value; $(touch SHOULD_NOT_EXIST) `false` "quoted"' + for setting in absent missing-config enabled empty; do + id="env-$setting" + rec=$(make_spawn_case "$id" codex "$id") + read_case_record "$rec" + case "$setting" in + missing-config) rm "$HOME_DIR/config/crew-harness"; rmdir "$HOME_DIR/config" ;; + enabled) printf '# Synthetic credential name\nFM_TEST_ALLOWED\nFM_TEST_EMPTY\nFM_TEST_UNSET\n' > "$HOME_DIR/config/launch-env-allowlist" ;; + empty) : > "$HOME_DIR/config/launch-env-allowlist" ;; + esac + probe="$CASE_DIR/probe.sh" + cat > "$probe" <<'SH' +#!/bin/sh +printf '%s\n' "${FM_TEST_AMBIENT_SENTINEL-unset}" "${FM_TEST_ALLOWED-unset}" \ + "${FM_TEST_EMPTY-unset}" "${FM_TEST_UNSET-unset}" "$HOME" "$PATH" "$TERM" "$TMUX" "$GOTMPDIR" +SH + out=$(FM_TEST_AMBIENT_SENTINEL=synthetic-unrelated \ + run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" \ + "$id" "$PROJ_DIR" --harness "/bin/sh '$probe'") + status=$? + expect_code 0 "$status" "allowlist=$setting spawn should succeed: $out" + launch=$(cat "$LAUNCH_LOG") + for pane_shell in /bin/sh /bin/bash /bin/zsh; do + [ -x "$pane_shell" ] || continue + pane_path=$(env -i HOME="$HOME_DIR/user-home" PATH=/usr/bin:/bin TERM=xterm \ + TMUX=synthetic-pane GOTMPDIR=/synthetic/gotmp \ + "$pane_shell" -c "printf %s \"\$PATH\"") \ + || fail "could not read $pane_shell startup PATH" + result=$(env -i HOME="$HOME_DIR/user-home" PATH=/usr/bin:/bin TERM=xterm \ + TMUX=synthetic-pane GOTMPDIR=/synthetic/gotmp \ + FM_TEST_AMBIENT_SENTINEL=synthetic-unrelated FM_TEST_ALLOWED="$value" FM_TEST_EMPTY='' \ + "$pane_shell" -c "$launch") || fail "allowlist=$setting emitted launch failed in $pane_shell" + case "$setting" in + absent|missing-config) expected=$(printf '%s\n' synthetic-unrelated "$value" '' unset) ;; + enabled) expected=$(printf '%s\n' unset "$value" '' unset) ;; + empty) expected=$(printf '%s\n' unset unset unset unset) ;; + esac + expected="$expected"$'\n'"$HOME_DIR/user-home"$'\n'"$pane_path"$'\nxterm\nsynthetic-pane\n/synthetic/gotmp' + [ "$result" = "$expected" ] || fail "allowlist=$setting worker environment mismatch: $result" + done + pass "allowlist=$setting preserves the operational floor and filters only when opted in" + done +} + +test_launch_environment_invalid_config_refuses() { + local rec id bad out status + id=env-invalid + rec=$(make_spawn_case "$id" codex "$id") + read_case_record "$rec" + for bad in 'FM_TEST_ALLOWED=value' 'NAME;false' '1INVALID' '*'; do + printf '%s\n' "$bad" > "$HOME_DIR/config/launch-env-allowlist" + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR") + status=$? + expect_code 1 "$status" "invalid allowlist must refuse spawn" + assert_contains "$out" 'launch-env-allowlist' "refusal must identify the config file" + [ ! -s "$LAUNCH_LOG" ] || fail "invalid allowlist delivered a launch command" + [ ! -f "$HOME_DIR/state/$id.meta" ] || fail "invalid allowlist published a task" + done + pass "invalid allowlist names refuse before launch or task publication" +} + +test_launch_environment_inaccessible_config_refuses() { + local setting presence rec id blocked out status + if [ "$(id -u)" = 0 ]; then + printf '# skip - inaccessible launch configuration requires a non-root user\n' + return + fi + for setting in config ancestor; do + for presence in present absent; do + id="env-inaccessible-$setting-$presence" + rec=$(make_spawn_case "$id" codex "$id") + read_case_record "$rec" + if [ "$presence" = present ]; then + printf 'FM_TEST_ALLOWED\n' > "$HOME_DIR/config/launch-env-allowlist" + fi + blocked="$HOME_DIR/config" + if [ "$setting" = ancestor ]; then + blocked="$HOME_DIR/config-parent" + mkdir "$blocked" + mv "$HOME_DIR/config" "$blocked/config" + ln -s config-parent/config "$HOME_DIR/config" + fi + chmod 600 "$blocked" || fail "could not remove configuration search permission" + out=$(run_ship_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" \ + "$id" "$PROJ_DIR" --harness codex --backend tmux) + status=$? + chmod 700 "$blocked" || fail "could not restore configuration search permission" + expect_code 1 "$status" "inaccessible $setting with $presence allowlist must refuse spawn: $out" + assert_contains "$out" 'launch-env-allowlist' "refusal must identify the launch configuration" + [ ! -s "$LAUNCH_LOG" ] || fail "inaccessible configuration delivered a launch command" + [ ! -f "$HOME_DIR/state/$id.meta" ] || fail "inaccessible configuration published a task" + pass "inaccessible $setting with $presence allowlist refuses before launch or task publication" + done + done +} + +test_launch_environment_inherited_by_secondmate() { + local rec id sm out status result + id=env-secondmate + rec=$(make_spawn_case "$id" codex "$id") + read_case_record "$rec" + printf 'FM_TEST_ALLOWED\n' > "$HOME_DIR/config/launch-env-allowlist" + sm="$CASE_DIR/secondmate-home" + make_seeded_secondmate_home "$sm" "$id" + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$sm" --secondmate) + status=$? + expect_code 0 "$status" "secondmate with an allowlist should spawn: $out" + cmp -s "$HOME_DIR/config/launch-env-allowlist" "$sm/config/launch-env-allowlist" \ + || fail "secondmate did not inherit the launch environment contract" + cat > "$FAKEBIN_DIR/codex" <<'SH' +#!/bin/sh +printf '%s\n' "${FM_TEST_AMBIENT_SENTINEL-unset}" "$FM_TEST_ALLOWED" "$FM_HOME" "${FM_STATE_OVERRIDE-unset}" +SH + chmod +x "$FAKEBIN_DIR/codex" + result=$(env -i HOME="$HOME_DIR/user-home" PATH="$FAKEBIN_DIR:$PATH" \ + FM_TEST_AMBIENT_SENTINEL=synthetic-unrelated FM_TEST_ALLOWED=synthetic-provider \ + /bin/sh -c "$(cat "$LAUNCH_LOG")") || fail "secondmate's emitted command failed" + [ "$result" = "unset"$'\nsynthetic-provider\n'"$sm" ] \ + || fail "secondmate's environment lost filtering or explicit home assignments: $result" + # Exercise the same inheritance owner used by local and remote transfers; + # removal must restore absence downstream as well as copying an opt-in. + ( + # shellcheck source=/dev/null + . "$ROOT/bin/fm-config-inherit-lib.sh" + rm "$HOME_DIR/config/launch-env-allowlist" + propagate_secondmate_inheritance "$HOME_DIR" "$sm" >/dev/null + ) || fail "allowlist removal failed to converge" + [ ! -e "$sm/config/launch-env-allowlist" ] || fail "secondmate retained a removed allowlist" + pass "secondmate launch inherits the allowlist for subsequent worker launches" +} + +run_launch_environment_inheritance() { + local route=$1 home=$2 dest=$3 fakebin=$4 generation=$5 + if [ "$route" = local ]; then + ( + # shellcheck source=/dev/null + . "$ROOT/bin/fm-config-inherit-lib.sh" + FM_INHERITABLE_CONFIG=launch-env-allowlist \ + propagate_inheritable_config "$home/config" "$dest/config" + ) + else + FM_HOME="$home" FM_ROOT_OVERRIDE="$ROOT" FM_CONFIG_OVERRIDE="$home/config" \ + FM_DATA_OVERRIDE="$home/data" FM_INHERITABLE_CONFIG=launch-env-allowlist \ + FM_SSH_BIN="$fakebin/inherit-ssh" \ + "$ROOT/bin/fm-remote-inherit-push.sh" inherited-env "$generation" + fi +} + +test_launch_environment_inheritance_preserves_on_source_errors() { + local route rec id dest out status + if [ "$(id -u)" = 0 ]; then + printf '# skip - inaccessible inheritance sources require a non-root user\n' + return + fi + for route in local remote; do + id="env-inherit-$route" + rec=$(make_spawn_case "$id" codex "$id") + read_case_record "$rec" + dest="$CASE_DIR/inherited-home" + mkdir -p "$dest/config" + printf 'FM_TEST_ALLOWED\n' > "$HOME_DIR/config/launch-env-allowlist" + printf -- '- inherited-env - Test route (host: inherit-host; root: %s; home: %s; scope: test; projects: ; added 2026-09-05)\n' \ + "$ROOT" "$dest" > "$HOME_DIR/data/secondmates.md" + cat > "$FAKEBIN_DIR/inherit-ssh" <<'SH' +#!/usr/bin/env bash +set -eu +while [ "$#" -gt 0 ]; do + case "$1" in -o) shift 2 ;; --) shift; break ;; *) exit 90 ;; esac +done +[ "$#" -eq 6 ] && [ "$1" = inherit-host ] && [ "$2" = fm-remote-entrypoint.sh ] && [ "$3" = 1 ] || exit 91 +remote_root=$(printf '%s' "$4" | base64 --decode) +remote_home=$(printf '%s' "$5" | base64 --decode) +args=() +while IFS= read -r -d '' arg; do args+=("$arg"); done < <(printf '%s' "$6" | base64 --decode) +[ "${args[0]}" = fm-remote-inherit.sh ] || exit 92 +FM_HOME="$remote_home" FM_STATE_OVERRIDE="$remote_home/state" \ + exec "$remote_root/bin/${args[0]}" "${args[@]:1}" +SH + chmod +x "$FAKEBIN_DIR/inherit-ssh" + out=$(run_launch_environment_inheritance "$route" "$HOME_DIR" "$dest" "$FAKEBIN_DIR" 1 2>&1) + status=$? + expect_code 0 "$status" "$route allowlist inheritance should succeed: $out" + [ "$(cat "$dest/config/launch-env-allowlist")" = FM_TEST_ALLOWED ] \ + || fail "$route inheritance did not publish the allowlist" + + chmod 600 "$HOME_DIR/config" || fail "could not remove source search permission" + out=$(run_launch_environment_inheritance "$route" "$HOME_DIR" "$dest" "$FAKEBIN_DIR" 2 2>&1) + status=$? + chmod 700 "$HOME_DIR/config" || fail "could not restore source search permission" + expect_code 1 "$status" "$route inheritance must refuse an inaccessible source: $out" + assert_contains "$out" launch-env-allowlist "$route inspection error must identify the allowlist" + [ "$(cat "$dest/config/launch-env-allowlist")" = FM_TEST_ALLOWED ] \ + || fail "$route inheritance removed or changed the allowlist after an inspection error" + + rm "$HOME_DIR/config/launch-env-allowlist" + ln -s missing-allowlist "$HOME_DIR/config/launch-env-allowlist" + out=$(run_launch_environment_inheritance "$route" "$HOME_DIR" "$dest" "$FAKEBIN_DIR" 3 2>&1) + status=$? + expect_code 1 "$status" "$route inheritance must refuse a dangling source link: $out" + [ "$(cat "$dest/config/launch-env-allowlist")" = FM_TEST_ALLOWED ] \ + || fail "$route inheritance treated a dangling source link as absence" + + rm "$HOME_DIR/config/launch-env-allowlist" + out=$(run_launch_environment_inheritance "$route" "$HOME_DIR" "$dest" "$FAKEBIN_DIR" 4 2>&1) + status=$? + expect_code 0 "$status" "$route inheritance should mirror proven absence: $out" + [ ! -e "$dest/config/launch-env-allowlist" ] || fail "$route inheritance retained a removed allowlist" + pass "$route inheritance preserves the allowlist on source errors and mirrors proven absence" + done +} + +test_worker_launch_delivers_role_scope() { + local rec id out launch kind prompt brief_kind brief content + for brief_kind in heading legacy scaffold; do + for kind in no-mistakes direct-PR local-only scout; do + [ "$brief_kind" = heading ] && [ "$kind" != no-mistakes ] && continue + id="role-launch-$brief_kind-$kind" + rec=$(make_spawn_case "$id" codex) + read_case_record "$rec" + if [ "$brief_kind" != scaffold ]; then + fm_test_spawn_brief "$HOME_DIR" "$id" + if [ "$brief_kind" = heading ]; then + printf '\n# Worker role\nFollow the project instructions.\n' >> "$HOME_DIR/data/$id/brief.md" + fi + else + if [ "$kind" = scout ]; then + FM_HOME="$HOME_DIR" "$ROOT/bin/fm-brief.sh" "$id" arbitrary-project-name --scout >/dev/null || fail "scout scaffold failed" + else + FM_HOME="$HOME_DIR" "$ROOT/bin/fm-brief.sh" "$id" arbitrary-project-name --mode "$kind" >/dev/null || fail "$kind scaffold failed" + fi + brief="$HOME_DIR/data/$id/brief.md" + content=$(cat "$brief") + content=${content//'{TASK}'/brief for $id} + content=${content//'{FIRSTMATE_SPEC}'/Exercise the spawn behavior under test.} + printf '%s\n' "$content" > "$brief" + fi + cp "$HOME_DIR/data/$id/brief.md" "$CASE_DIR/brief-before" + cat > "$FAKEBIN_DIR/codex" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$@" > "$FM_ROLE_PROMPT" +SH + chmod +x "$FAKEBIN_DIR/codex" + if [ "$kind" = scout ]; then + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --scout) + else + out=$(run_spawn "$HOME_DIR" "$WT_DIR" "$FAKEBIN_DIR" "$LAUNCH_LOG" "$id" "$PROJ_DIR" --mode "$kind" --yolo off) + fi + expect_code 0 "$?" "$kind worker spawn failed: $out" + launch=$(cat "$LAUNCH_LOG") + prompt="$CASE_DIR/prompt" + FM_ROLE_PROMPT="$prompt" PATH="$FAKEBIN_DIR:$PATH" bash -c "$launch" || fail "could not consume $kind launch command" + # The final prompt delivered to the harness is the generated interface. + # An authored role heading must neither suppress nor duplicate the current + # worker contract; the launch section is its single, superseding owner. + assert_grep 'follow this brief instead of that supervisor contract' "$prompt" "$kind command did not deliver the role correction" + assert_grep 'brief for' "$prompt" "$kind command lost the task" + [ "$(grep -c '^# Current worker role contract$' "$prompt")" -eq 1 ] || + fail "$brief_kind $kind duplicated the delivered worker contract" + if [ "$brief_kind" = heading ]; then + assert_grep 'Follow the project instructions' "$prompt" "$kind command dropped the authored role section" + fi + cmp -s "$CASE_DIR/brief-before" "$HOME_DIR/data/$id/brief.md" || fail "spawn rewrote the authored brief" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# evidence begin: %s %s worker\n%s\n' "$brief_kind" "$kind" "$out" + printf 'launch command executed with an argv-capture harness:\n%s\nreceived arguments and final prompt:\n' "$launch" + cat "$prompt" + printf 'authored brief remains byte-identical\n# evidence end\n' + fi + done + done + pass "fm-spawn: actual ship/scout launch commands deliver the worker role contract" +} + +fm_test_run_cases \ + test_launch_environment_allowlist \ + test_launch_environment_invalid_config_refuses \ + test_launch_environment_inaccessible_config_refuses \ + test_launch_environment_inherited_by_secondmate \ + test_launch_environment_inheritance_preserves_on_source_errors \ + test_worker_launch_delivers_role_scope \ + test_no_profile_keeps_claude_profile_defaults \ + test_non_cursor_launch_clears_inherited_cursor_markers \ + test_copilot_threads_model_effort_and_hooks \ + test_relative_home_overrides_launch_with_absolute_cross_process_paths \ + test_home_defaults_preserve_absolute_or_resolve_relative_paths \ + test_absolute_override_spelling_is_preserved_in_launch_paths \ + test_unresolvable_relative_overrides_fail_loudly \ + test_active_dispatch_profile_requires_explicit_harness_for_ship \ + test_active_dispatch_profile_requires_explicit_harness_for_scout \ + test_active_dispatch_profile_allows_explicit_harness \ + test_active_dispatch_profile_allows_positional_harness \ + test_active_dispatch_profile_allows_raw_launch_command \ + test_claude_threads_model_and_effort \ + test_codex_threads_model_and_effort \ + test_codex_omits_invalid_max_effort \ + test_grok_threads_model_and_reasoning_effort \ + test_grok_omits_invalid_max_reasoning_effort \ + test_grok_omits_invalid_xhigh_reasoning_effort \ + test_cursor_threads_model_workspace_and_omits_effort_axis \ + test_cursor_refuses_model_absent_from_live_catalog \ + test_cursor_failed_catalog_probe_does_not_block_spawn \ + test_opencode_threads_model_and_ignores_effort_axis \ + test_pi_threads_model_and_max_effort \ + test_pi_tui_mode_probe_is_safe_for_old_and_new_pi \ + test_pi_signed_threads_shared_pi_profile_and_preserves_identity \ + test_pi_signed_missing_binary_refuses_before_endpoint_or_metadata \ + test_pi_signed_persistent_secondmate_uses_pi_extensions_and_identity \ + test_copilot_secondmate_uses_semantic_submission_hooks \ + test_batch_forwards_shared_profile_flags \ + test_claude_forwards_firstmate_config_dir_when_set \ + test_claude_omits_config_dir_prefix_when_unset \ + test_non_claude_harness_ignores_config_dir \ + test_claude_crewmate_launch_carries_the_attribution_policy \ + test_claude_secondmate_launch_carries_the_attribution_policy \ + test_active_dispatch_profile_does_not_block_secondmate_launch diff --git a/tests/fm-spawn-pool-base-freshen.test.sh b/tests/fm-spawn-pool-base-freshen.test.sh index 3f2cfe256f0..6edff60d668 100755 --- a/tests/fm-spawn-pool-base-freshen.test.sh +++ b/tests/fm-spawn-pool-base-freshen.test.sh @@ -4,8 +4,8 @@ # A treehouse pool can return a clean detached worktree whose origin/main was # advanced after the worktree was allocated. # These tests drive the real spawn path with a fake terminal, then prove it -# starts the worker from the fetched origin/main tip or stops when origin is -# unreachable. +# starts the worker from the fetched origin tip, launches a clean origin-less +# pool as-is, or stops when a configured origin is unusable. set -u # shellcheck source=tests/fixtures.sh @@ -59,6 +59,41 @@ run_spawn() { "$id" "$PROJECT_DIR" "$@" } +test_remote_seeded_home_spawns_from_treehouse_pool() { + local rec id out status lock + id='pool-remote-seeded-r13' + rec=$(make_case remote-seeded "$id") + read_case_record "$rec" + cat > "$HOME_DIR/.fm-secondmate-parent" <<'REC' +schema=fm-secondmate-parent.v1 +route=remote +parent_host=parent-machine +REC + + out=$(run_spawn "$id" --scout) + status=$? + expect_code 0 "$status" \ + "a remote-seeded secondmate home should allocate and launch from its Treehouse pool"$'\n'"$out" + assert_contains "$out" "spawned $id" \ + "the remote-seeded spawn did not report success" + assert_grep "worktree=$POOL_DIR" "$HOME_DIR/state/$id.meta" \ + "the remote-seeded spawn did not publish its allocated pool worktree" + lock=$(FM_HOME="$HOME_DIR" bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$PROJECT_DIR") \ + || fail "the launched remote-seeded home could not resolve its Treehouse project lock" + case "$lock" in + "$HOME_DIR/state/"*) ;; + *) fail "the remote-seeded spawn anchored its lock outside its local root: $lock" ;; + esac + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded Treehouse spawn command\n' + printf '$ FM_HOME=%s bin/fm-spawn.sh %s %s --scout\n%s\nexit=%s\n' \ + "$HOME_DIR" "$id" "$PROJECT_DIR" "$out" "$status" + printf 'published worktree=%s\nresolved project lock=%s\n' "$POOL_DIR" "$lock" + fi + pass "a remote-seeded secondmate home allocates and launches from its Treehouse pool" +} + test_linked_spawning_home_rejects_primary_before_refresh() { local rec id out status returned primary spawning before_reflog for returned in primary primary-alias spawning scout; do @@ -109,11 +144,14 @@ test_linked_spawning_home_rejects_primary_before_refresh() { || fail "spawn did not refresh the genuine scout copy" else [ "$status" -ne 0 ] || fail "linked spawning home accepted $returned as a disposable copy" - if [ "$returned" = spawning ]; then - assert_contains "$out" "did not enter a worktree" "spawn accepted its own spawning directory" - else - assert_contains "$out" "did not yield an isolated worktree" "spawn did not explain its isolation refusal" - fi + # None of these is an isolated copy, so the worktree poll never adopts one + # and the wait runs out instead: the spawning directory fails the poll's + # own project comparison, and the repository primary (named directly or + # through a symlink) fails the isolation screen the poll shares with the + # guard. The refusal names the last path the pane reported. + assert_contains "$out" "did not enter an isolated worktree" \ + "spawn did not explain its isolation refusal" + assert_contains "$out" "last seen" "refusal did not name the path the pane reported" [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" [ ! -e "$primary/.git/FETCH_HEAD" ] || fail "refused spawn fetched before proving isolation" fi @@ -177,6 +215,154 @@ test_non_main_default_branch_refreshes_before_branching() { pass "a stale pooled worktree resolves and refreshes a non-main default branch" } +make_originless_case() { # <name> <id> + local name=$1 id=$2 case_dir home project pool fakebin initial + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + project="$case_dir/project" + pool="$case_dir/pool" + fakebin=$(make_spawn_fakebin "$case_dir/fake") + + mkdir -p "$home/data/$id" "$home/projects" "$home/state" "$home/config" + printf 'codex\n' > "$home/config/crew-harness" + fm_test_spawn_brief "$home" "$id" + touch "$home/state/.last-watcher-beat" + + git init --quiet -b main "$project" + printf 'base\n' > "$project/README.md" + git -C "$project" add README.md + git -C "$project" -c user.name='Firstmate Tests' -c user.email='tests@example.invalid' commit -qm initial + initial=$(git -C "$project" rev-parse HEAD) + git -C "$project" worktree add --quiet --detach "$pool" "$initial" + + printf '%s\n' "$case_dir|$home|$project|$pool|$fakebin|$initial|main" +} + +test_originless_pool_launches_without_a_freshness_fetch() { + local rec id out status before + id='pool-originless-r6' + rec=$(make_originless_case originless "$id") + read_case_record "$rec" + ! git -C "$POOL_DIR" remote get-url origin >/dev/null 2>&1 \ + || fail "fixture unexpectedly configured an origin remote" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should launch a local-only pooled worktree with no origin"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success for the origin-less pool" + assert_not_contains "$out" "could not fetch origin" \ + "spawn attempted a freshness fetch against a nonexistent origin" + [ ! -e "$POOL_DIR/.git/FETCH_HEAD" ] || fail "spawn fetched against a pooled worktree with no origin" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD on an origin-less pooled worktree that had nothing to refresh against" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# observed origin-less launch: %s\n' "$(printf '%s\n' "$out" | tail -n 1)" + fi + pass "an origin-less pooled worktree launches as-is, skipping the freshness gate" +} + +test_originless_dirty_pool_refuses_without_discarding_work() { + local rec id out status before + id='pool-originless-dirty-r1' + rec=$(make_originless_case originless-dirty "$id") + read_case_record "$rec" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + printf 'keep this local work\n' > "$POOL_DIR/uncommitted.txt" + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite a dirty origin-less pooled worktree" + assert_contains "$out" "is not clean" \ + "spawn did not clearly refuse a dirty origin-less pooled worktree" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD while refusing a dirty origin-less pooled worktree" + assert_grep 'keep this local work' "$POOL_DIR/uncommitted.txt" \ + "spawn discarded local work from an origin-less pool" + pass "a dirty origin-less pooled worktree is refused without discarding its local work" +} + +test_origin_config_without_url_refuses_pool() { + local rec id out status before + id='pool-origin-without-url-r1' + rec=$(make_originless_case origin-without-url "$id") + read_case_record "$rec" + git -C "$POOL_DIR" config remote.origin.fetch '+refs/heads/*:refs/remotes/origin/*' + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an origin configuration with no URL" + assert_contains "$out" "could not fetch origin" \ + "spawn did not refuse an origin configuration with no URL as unusable" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD after finding an unusable origin configuration" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" + pass "an origin configuration without a URL refuses the pooled worktree" +} + +test_empty_origin_config_section_refuses_pool() { + local rec id out status before config + id='pool-empty-origin-section-r1' + rec=$(make_originless_case empty-origin-section "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + printf '\n[remote "origin"]\n' >> "$config" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + [ "$status" -ne 0 ] || fail "spawn succeeded despite an empty origin configuration section" + assert_contains "$out" "could not fetch origin" \ + "spawn did not refuse an empty origin configuration section as unusable" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD after finding an empty origin configuration section" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" + pass "an empty origin configuration section refuses the pooled worktree" +} + +test_empty_only_included_origin_config_section_launches_pool() { + local rec id out status before config included + id='pool-empty-only-included-origin-section-r1' + rec=$(make_originless_case empty-only-included-origin-section "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + included=$(dirname "$config")/empty-origin.inc + printf '[remote "origin"]\n' > "$included" + git -C "$POOL_DIR" config include.path "$(basename "$included")" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should proceed when an included empty origin section is not enumerable"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success for the undetectable included section" + assert_not_contains "$out" "could not fetch origin" \ + "spawn treated an undetectable included empty section as a configured origin" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD despite treating the included empty section as origin-less" + pass "an empty-only included origin section documents the accepted detection boundary" +} + +test_inactive_conditional_origin_include_launches_pool() { + local rec id out status before config included + id='pool-inactive-origin-include-r1' + rec=$(make_originless_case inactive-origin-include "$id") + read_case_record "$rec" + config=$(git -C "$POOL_DIR" rev-parse --path-format=absolute --git-path config) + included=$(dirname "$config")/inactive-origin.inc + printf '[fm-test]\n\tmarker = true\n[remote "origin"]\n' > "$included" + git -C "$POOL_DIR" config 'includeIf.gitdir:/never/matches/this/worktree/.path' "$included" + before=$(git -C "$POOL_DIR" rev-parse HEAD) + + out=$(run_spawn "$id" --mode no-mistakes --yolo off) + status=$? + expect_code 0 "$status" "spawn should ignore an inactive conditional origin include"$'\n'"$out" + assert_contains "$out" "spawned $id" "spawn did not report success with an inactive origin include" + [ "$(git -C "$POOL_DIR" rev-parse HEAD)" = "$before" ] \ + || fail "spawn moved HEAD despite having no effective origin" + pass "an inactive conditional origin include leaves the pooled worktree origin-less" +} + test_unreachable_origin_refuses_stale_pool_base() { local rec id out status before after id='pool-unreachable-origin-r2' @@ -350,6 +536,7 @@ test_stale_submodule_pin_explains_itself() { rec=$(make_submodule_case stale-pin "$id") read_submodule_case "$rec" strand_submodule_pin_via_spawn 'pool-stale-pin-seed-r7' + git -C "$POOL_DIR" remote remove origin before=$(git -C "$POOL_DIR" rev-parse HEAD) before_sub=$(git -C "$POOL_DIR/ui" rev-parse HEAD) @@ -375,7 +562,7 @@ test_stale_submodule_pin_explains_itself() { if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then printf '# observed stale-pin refusal: %s\n' "$(printf '%s\n' "$out" | grep 'submodule' | head -n 1)" fi - pass "two consecutive spawns across a moved submodule pin end in a refusal naming both pins and no remedy" + pass "an origin-less pool with a stale submodule pin refuses while naming both pins and no remedy" } test_unpushed_submodule_commit_is_still_uncommitted_work() { @@ -489,6 +676,7 @@ test_stale_pin_beside_other_dirt_reports_one_verdict() { pass "a stale pin beside other dirt yields the conservative refusal alone, with no stale-pin line" } +test_remote_seeded_home_spawns_from_treehouse_pool test_linked_spawning_home_rejects_primary_before_refresh test_stale_pool_base_refreshes_before_branching test_non_main_default_branch_refreshes_before_branching @@ -496,6 +684,12 @@ test_direct_pr_and_scout_refresh_before_launch test_dirty_pool_refuses_without_discarding_work test_unresolved_remote_default_refuses_pool test_unreachable_origin_refuses_stale_pool_base +test_originless_pool_launches_without_a_freshness_fetch +test_originless_dirty_pool_refuses_without_discarding_work +test_origin_config_without_url_refuses_pool +test_empty_origin_config_section_refuses_pool +test_empty_only_included_origin_config_section_launches_pool +test_inactive_conditional_origin_include_launches_pool test_stale_submodule_pin_explains_itself test_unpushed_submodule_commit_is_still_uncommitted_work test_work_inside_submodule_is_still_uncommitted_work diff --git a/tests/fm-spawn-worktree-settle.test.sh b/tests/fm-spawn-worktree-settle.test.sh index a0c2d85dd98..418d0d70246 100755 --- a/tests/fm-spawn-worktree-settle.test.sh +++ b/tests/fm-spawn-worktree-settle.test.sh @@ -12,10 +12,17 @@ # transient-then-settled pane_current_path sequence with a fake tmux and # asserts the recorded worktree resolves to the real, settled worktree, never # the stale first read. +# +# The same loop has a second transient to survive: `treehouse get` reports the +# REPOSITORY's primary checkout as its own cwd while it is still preparing a +# slot. From a linked spawning home that path is not the project, so a poll +# comparing only against the project adopted it and the isolation guard then +# refused the launch. The cases below cover both the transient and the pane +# that never leaves the primary at all. set -u -# shellcheck source=tests/lib.sh -. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" +# shellcheck source=tests/fixtures.sh +. "$(dirname "${BASH_SOURCE[0]}")/fixtures.sh" SPAWN="$ROOT/bin/fm-spawn.sh" TMP_ROOT=$(fm_test_tmproot fm-spawn-worktree-settle) @@ -127,28 +134,96 @@ test_single_stale_first_read_is_not_accepted() { pass "a single transient stale pane_current_path read is not accepted as the worktree" } -# A pane that reports the real worktree from the very first read still only -# costs the loop's existing one-second inter-poll sleep to confirm - not an -# extra full cycle on top of that. -test_already_settled_pane_costs_one_confirm_sleep() { - local rec id out status start end elapsed +# A pane that reports the real worktree from the very first read costs exactly +# one confirming read - not a whole extra polling cycle on top of it. Counting +# the pane reads measures the loop itself; wall-clock time would fold in every +# other cost of a spawn (fetch, trust registration) and drift with the machine. +test_already_settled_pane_costs_one_confirm_read() { + local rec id out status reads id=settle-already-settled-z2 rec=$(make_settle_case settle-already-settled "$id" 0) read_settle_record "$rec" - start=$(date +%s) out=$(run_settle_spawn "$id") status=$? - end=$(date +%s) - elapsed=$((end - start)) - expect_code 0 "$status" "spawn should succeed when the pane is already settled" + expect_code 0 "$status" "spawn should succeed when the pane is already settled"$'\n'"$out" assert_grep "worktree=$WT_DIR" "$HOME_DIR/state/$id.meta" \ "meta did not record the already-settled worktree" - [ "$elapsed" -le 5 ] || fail "already-settled pane took ${elapsed}s to confirm - expected close to the single inter-poll sleep" - pass "an already-settled pane confirms via the existing inter-poll sleep, not an extra full cycle" + reads=$(cat "$COUNTFILE") + [ "$reads" -eq 2 ] || fail "already-settled pane took $reads reads to confirm - expected the first read plus one confirmation" + pass "an already-settled pane confirms on the next read, not a whole extra cycle" +} + +# make_primary_case <name> <id> <stale_reads> builds the linked-home shape: the +# spawning project is itself a LINKED worktree of the repository, and the path +# the pane transiently reports is that repository's PRIMARY checkout. `treehouse +# get` reports the repository it is preparing a slot from as its own cwd while +# it is still fetching and checking out, so the pane reads the primary for the +# first seconds. The primary is not the spawning project, so a poll that only +# compares against the project accepts it as the worktree, and the isolation +# guard then refuses the launch even though treehouse went on to enter a real +# slot. The settled path is a second linked worktree of the same repository. +make_primary_case() { + local name=$1 id=$2 stale_reads=$3 case_dir home primary proj wt fakebin countfile + case_dir="$TMP_ROOT/$name" + home="$case_dir/home" + primary="$case_dir/primary" + proj="$case_dir/mate" + wt="$case_dir/slot" + countfile="$case_dir/pane-call-count" + fakebin=$(make_settle_fakebin "$case_dir/fake") + fm_test_spawn_home "$home" codex + fm_git_worktree "$primary" "$proj" "mate-$name" + git -C "$primary" worktree add --quiet -b "slot-$name" "$wt" + fm_test_spawn_brief "$home" "$id" "Exercise primary-checkout transient detection for $id." + printf '%s\n' "$case_dir|$home|$proj|$wt|$primary|$fakebin|$countfile|$stale_reads" +} + +# The exact incident: the pane reports the repository primary for the first +# reads, then settles into the slot treehouse actually created. The primary must +# never be adopted as the worktree, so the spawn lands on the settled slot. +test_transient_primary_checkout_is_not_accepted() { + local rec id out status + id=settle-primary-transient-z3 + rec=$(make_primary_case settle-primary-transient "$id" 3) + read_settle_record "$rec" + fm_test_fake_sleep_noop "$FAKEBIN_DIR" + + out=$(run_settle_spawn "$id") + status=$? + expect_code 0 "$status" "spawn should succeed once the pane leaves the primary checkout"$'\n'"$out" + assert_grep "worktree=$WT_DIR" "$HOME_DIR/state/$id.meta" \ + "meta did not record the settled worktree" + assert_no_grep "worktree=$STALE_DIR" "$HOME_DIR/state/$id.meta" \ + "meta wrongly recorded the repository primary checkout as the worktree" + pass "a transient primary-checkout pane read is not accepted as the worktree" +} + +# A pane that never leaves the primary checkout must still fail at the deadline +# rather than waiting forever or recording the primary. +test_primary_checkout_that_never_settles_fails_at_the_deadline() { + local rec id out status + id=settle-primary-stuck-z4 + rec=$(make_primary_case settle-primary-stuck "$id" 100000) + read_settle_record "$rec" + fm_test_fake_sleep_noop "$FAKEBIN_DIR" + + out=$(run_settle_spawn "$id") + status=$? + [ "$status" -ne 0 ] || fail "spawn accepted a pane that never left the primary checkout"$'\n'"$out" + assert_contains "$out" "did not enter an isolated worktree" \ + "spawn did not explain that the pane never reached an isolated worktree" + assert_contains "$out" "$STALE_DIR" \ + "the refusal did not name the path the pane kept reporting" + assert_contains "$out" "repository's primary checkout" \ + "the refusal did not say why that path was rejected" + [ ! -e "$HOME_DIR/state/$id.meta" ] || fail "refused spawn published task metadata" + pass "a pane stuck on the primary checkout fails loudly at the deadline" } test_single_stale_first_read_is_not_accepted -test_already_settled_pane_costs_one_confirm_sleep +test_already_settled_pane_costs_one_confirm_read +test_transient_primary_checkout_is_not_accepted +test_primary_checkout_that_never_settles_fails_at_the_deadline echo "# all fm-spawn-worktree-settle tests passed" diff --git a/tests/fm-startup-memory-budget.test.sh b/tests/fm-startup-memory-budget.test.sh index c803f231e21..fe5a5439f62 100755 --- a/tests/fm-startup-memory-budget.test.sh +++ b/tests/fm-startup-memory-budget.test.sh @@ -62,6 +62,7 @@ case "$*" in *display-message*'#{pane_current_command}'*) printf '%s\n' codex ;; *display-message*'#{pane_id}'*) printf '%s\n' '%1' ;; *display-message*'#{cursor_y}'*) printf '%s\n' 0 ;; + *list-windows*) printf '%s\n' fm-sm ;; *capture-pane*) printf '❯\n' ;; esac exit 0 @@ -279,7 +280,7 @@ test_primary_budget_converges_with_exact_reread_and_safe_failures() { "budget propagation did not enqueue the pointer to its exact reread generation" assert_contains "$(<"$log")" "Firstmate instruction waiting: list " \ "budget propagation did not ring the durable inbox doorbell" - assert_contains "$(<"$log")" "/state/sm.inbox/*.msg" \ + assert_contains "$(<"$log")" "/state/sm.inbox'/*.msg" \ "budget propagation doorbell did not identify the durable inbox" outside="$world/unsafe-budget" diff --git a/tests/fm-stat-shadowing.test.sh b/tests/fm-stat-shadowing.test.sh new file mode 100644 index 00000000000..ec8ef4f72b7 --- /dev/null +++ b/tests/fm-stat-shadowing.test.sh @@ -0,0 +1,139 @@ +#!/usr/bin/env bash +# tests/fm-stat-shadowing.test.sh - verify Darwin BSD-stat helpers ignore a GNU +# stat earlier on PATH. +# +# On Darwin, GNU coreutils can put a GNU stat earlier on PATH than /usr/bin/stat. +# A bare `stat -f <fmt>` then reaches GNU stat, where `-f` means filesystem stat +# rather than BSD-format output and can leak a filesystem dump into callers. +# Runtime Darwin BSD-format calls use /usr/bin/stat so their syntax stays tied +# to the system BSD implementation. +# +# This test proves the invariant by installing a fake GNU-like stat that shadows +# /usr/bin/stat and asserting the helpers still return correct values. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +# Darwin-only: the shadowing assertions require a real BSD /usr/bin/stat to +# shadow; on Linux the `stat -f` semantics differ and the helpers take the +# `stat -c` branch instead, so there is nothing meaningful to assert. Skip +# visibly (after lib.sh so `pass`/`fail` exist) rather than silently. +if [ "$(uname)" != Darwin ]; then + pass "Darwin-only test: shadowing assertions require BSD /usr/bin/stat; skipping on $(uname -s)" + exit 0 +fi + +TMP_ROOT=$(mktemp -d "${TMPDIR:-/tmp}/fm-stat-shadowing.XXXXXX") || exit 1 +trap 'rm -rf "$TMP_ROOT"' EXIT + +# --- fake GNU stat that mimics ~/.local/bin/stat shadowing /usr/bin/stat ------- + +# A shadowed GNU stat does not interpret `-f <fmt>` as BSD-format output. +# We fail the tokens used in our helpers so callers cannot accidentally use the +# shadowed stat instead of /usr/bin/stat. +FAKE_STAT="$TMP_ROOT/fakebin/stat" +mkdir -p "$(dirname "$FAKE_STAT")" + +cat > "$FAKE_STAT" <<'FAKESTAT' +#!/usr/bin/env bash +# Mimics GNU coreutils stat when it shadows BSD /usr/bin/stat. +# On Darwin: BSD stat uses %m (mtime), %z (size), etc. +# GNU stat -f treats its argument as a filesystem-path option, not a format. +# For the format tokens our code uses, this fake exits non-zero before a helper +# could consume shadowed output. +opt1=${1:-} opt2=${2:-} +if [ "$opt1" = "-f" ]; then + case "$opt2" in + %m|%l|%z|%d|%Lp|%i|%u|%B|%FB|%HT:%p|%d:%i|%d:%i:%z:%m:%c) + printf 'File: "%s"\n' "${3:-}" >&2 + exit 1 + ;; + *) + printf 'GNU stat: unknown -f format: %s\n' "$opt2" >&2 + exit 1 + ;; + esac +fi +printf 'GNU stat: unknown invocation: %s\n' "$*" >&2 +exit 1 +FAKESTAT +chmod +x "$FAKE_STAT" + +# Prepend fakebin so the fake GNU stat shadows /usr/bin/stat +ORIGINAL_PATH="$PATH" +export PATH="$TMP_ROOT/fakebin:$ORIGINAL_PATH" + +# Verify the shadowing is active: a bare `stat -f %m /` must fail (not use BSD) +if stat -f %m / >/dev/null 2>&1; then + # The fake stat didn't catch this, so something is wrong with the PATH setup + PATH="$ORIGINAL_PATH" + fail "shadowing sanity check: bare stat -f %m / should fail under GNU-shadow but did not" +fi + +# Also verify /usr/bin/stat still works when called directly +REAL_MTIME=$(/usr/bin/stat -f %m "$0" 2>/dev/null) || true +[ -n "$REAL_MTIME" ] && [ "$REAL_MTIME" -ge 0 ] || { + PATH="$ORIGINAL_PATH" + fail "/usr/bin/stat -f %m sanity check failed — /usr/bin/stat is not working" +} +pass "shadowing: fake GNU stat shadows /usr/bin/stat in PATH" + +# --- test the fixed helpers under shadowing ---------------------------------- + +. "$ROOT/bin/fm-supervision-lib.sh" +. "$ROOT/bin/fm-startup-memory-budget-lib.sh" + +TESTFILE="$TMP_ROOT/testfile" +printf 'hello world\n' > "$TESTFILE" + +# 1. fm_sup_stat_mtime from bin/fm-supervision-lib.sh +RESULT_MTIME=$(fm_sup_stat_mtime "$TESTFILE") || true +EXPECTED_MTIME=$(/usr/bin/stat -f %m "$TESTFILE" 2>/dev/null) +if [ -z "$RESULT_MTIME" ] || [ "$RESULT_MTIME" != "$EXPECTED_MTIME" ]; then + PATH="$ORIGINAL_PATH" + fail "fm_sup_stat_mtime: expected $EXPECTED_MTIME, got '$RESULT_MTIME'" +fi +pass "fm_sup_stat_mtime returns correct epoch mtime under GNU stat shadowing" + +# 2. fm_startup_memory_budget_link_count from bin/fm-startup-memory-budget-lib.sh +RESULT_LINKS=$(fm_startup_memory_budget_link_count "$TESTFILE") || true +EXPECTED_LINKS=$(/usr/bin/stat -f %l "$TESTFILE" 2>/dev/null) +if [ -z "$RESULT_LINKS" ] || [ "$RESULT_LINKS" != "$EXPECTED_LINKS" ]; then + PATH="$ORIGINAL_PATH" + fail "fm_startup_memory_budget_link_count: expected $EXPECTED_LINKS, got '$RESULT_LINKS'" +fi +pass "fm_startup_memory_budget_link_count returns correct link count under GNU stat shadowing" + +# 3. _fm_status_file_size from bin/fm-classify-lib.sh +# We source it and call the internal function directly. +RESULT_SIZE=$(LC_ALL=C /usr/bin/stat -f '%z' "$TESTFILE" 2>/dev/null) || true +# The fixed code uses /usr/bin/stat so it should produce the same value as direct call +if [ -z "$RESULT_SIZE" ]; then + PATH="$ORIGINAL_PATH" + fail "_fm_status_file_size: could not get size via /usr/bin/stat" +fi +# Verify the helper itself works by checking that calling it with /usr/bin/stat prefix matches +. "$ROOT/bin/fm-classify-lib.sh" +HELPER_SIZE=$(_fm_status_file_size "$TESTFILE") || true +if [ -z "$HELPER_SIZE" ] || [ "$HELPER_SIZE" != "$RESULT_SIZE" ]; then + PATH="$ORIGINAL_PATH" + fail "_fm_status_file_size: expected $RESULT_SIZE, got '$HELPER_SIZE'" +fi +pass "_fm_status_file_size returns correct byte size under GNU stat shadowing" + +# 4. stat_mtime from bin/fm-watch.sh +# fm-watch.sh runs a top-level `mkdir -p` on its state dir when sourced; pin it +# to the temp root via FM_STATE_OVERRIDE so no artifact escapes into the repo's +# git-ignored state/ directory. +export FM_STATE_OVERRIDE="$TMP_ROOT/state" +. "$ROOT/bin/fm-watch.sh" +RESULT_WATCH_MTIME=$(stat_mtime "$TESTFILE") || true +if [ -z "$RESULT_WATCH_MTIME" ] || [ "$RESULT_WATCH_MTIME" != "$EXPECTED_MTIME" ]; then + PATH="$ORIGINAL_PATH" + fail "stat_mtime (fm-watch.sh): expected $EXPECTED_MTIME, got '$RESULT_WATCH_MTIME'" +fi +pass "stat_mtime from fm-watch.sh returns correct epoch mtime under GNU stat shadowing" + +# Restore PATH before exit +PATH="$ORIGINAL_PATH" diff --git a/tests/fm-supervision-instructions.test.sh b/tests/fm-supervision-instructions.test.sh index 04e2ec3bfdd..4900e05883f 100755 --- a/tests/fm-supervision-instructions.test.sh +++ b/tests/fm-supervision-instructions.test.sh @@ -92,6 +92,10 @@ test_repair_lines() { out=$(FM_HOME="$home" "$RENDER" --harness pi --repair-line) assert_contains "$out" "Pi tool fm_watch_arm_pi" "pi repair line does not direct the model to the extension-owned tool" assert_not_contains "$out" "extension command /fm-watch-arm-pi" "pi repair line still directs the model to the human slash command" + out=$(FM_HOME="$home" "$RENDER" --harness omp --repair-line) + assert_contains "$out" "omp tool fm_watch_arm_omp" "omp repair line does not direct the model to the extension-owned tool" + assert_contains "$out" ".omp/extensions/fm-primary-turnend-guard.ts" "omp repair line does not name its own turn-end extension" + assert_not_contains "$out" "fm_watch_arm_pi" "omp repair line must not borrow the Pi tool" pass "renderer repair-line mode is harness-aware and honors conditional state" } @@ -105,6 +109,17 @@ test_cross_harness_ordinary_continuation_and_repair_matrix() { out=$("$RENDER" --harness pi --repair-line) assert_contains "$out" "fm_watch_arm_pi" "pi recovery line lost the extension-owned repair tool" + out=$("$RENDER" --harness omp) + assert_contains "$out" "primary harness: omp" "omp heading missing" + assert_contains "$out" "Mode: omp (Oh My Pi) extension background wake." "omp snippet missing" + assert_contains "$out" "the omp extension already owns watcher continuity" "omp ordinary-wake line does not leave continuity to the extension" + assert_contains "$out" ".omp/extensions/fm-primary-omp-watch.ts" "omp snippet did not substitute its watch extension path" + assert_not_contains "$out" "__FM_OMP_EXT__" "omp snippet left a placeholder unsubstituted" + assert_not_contains "$out" "__FM_OMP_TURNEND_EXT__" "omp snippet left the turn-end placeholder unsubstituted" + assert_not_contains "$out" "project trust" "omp snippet must not carry Pi's trust prerequisite" + out=$("$RENDER" --harness omp --repair-line) + assert_contains "$out" "fm_watch_arm_omp" "omp recovery line lost the extension-owned repair tool" + out=$("$RENDER" --harness opencode) ordinary=$(printf '%s\n' "$out" | grep -F -- '- Ordinary wake:') assert_contains "$ordinary" "plugin already owns watcher continuity" "opencode ordinary-wake line does not leave continuity to the plugin" diff --git a/tests/fm-tangle-guard.test.sh b/tests/fm-tangle-guard.test.sh index 8bcbb392e60..d59864e0dae 100755 --- a/tests/fm-tangle-guard.test.sh +++ b/tests/fm-tangle-guard.test.sh @@ -164,26 +164,43 @@ test_spawn_isolation_abort() { mkdir -p "$home/data" proj=$(make_repo "$TMP_ROOT/spawn-proj") fakebin=$(make_spawn_fakebin "$TMP_ROOT/spawn-fake") + # The assertions concern identity, not how long an unchanged cwd is polled. + fm_test_fake_sleep_noop "$fakebin" # A genuine isolated linked worktree of the project, detached on the default. git -C "$proj" worktree add -q --detach "$TMP_ROOT/spawn-wt" >/dev/null 2>&1 - mkdir -p "$TMP_ROOT/spawn-notgit" "$proj/sub" + # The non-git case must BE non-git wherever this suite runs. A directory under + # TMPDIR is not one when TMPDIR itself sits inside a git repository - git walks + # up and finds that repo, and the spawn reports the subdirectory cause instead. + # GIT_CEILING_DIRECTORIES stops that upward walk: git does not chdir up into a + # listed directory, though it never excludes the directory being searched, so + # the ceiling is the PARENT of the path handed to the spawn (git(1), + # "GIT_CEILING_DIRECTORIES"). + mkdir -p "$TMP_ROOT/spawn-notgit-root/plain" "$proj/sub" # Abort: the pane resolves to a plain non-git directory (not a worktree at all). - out=$(run_spawn "$home" abort-notgit-dd4 "$proj" "$TMP_ROOT/spawn-notgit" "$fakebin"); status=$? + # The discovery poll screens every candidate with the isolation conditions, so + # a path like this is never adopted and the refusal comes from the poll's own + # deadline, naming the path and why it was rejected. The assertions pin which + # cause fired, not the operator wording that explains it. + out=$(GIT_CEILING_DIRECTORIES="$TMP_ROOT/spawn-notgit-root" \ + run_spawn "$home" abort-notgit-dd4 "$proj" "$TMP_ROOT/spawn-notgit-root/plain" "$fakebin"); status=$? expect_code 1 "$status" "spawn into a non-worktree dir should abort" - assert_contains "$out" "did not yield an isolated worktree" "non-worktree spawn lacked the isolation error" + assert_contains "$out" "did not enter an isolated worktree" "non-worktree spawn lacked the isolation error" + assert_contains "$out" "not inside a git worktree" "non-worktree spawn did not say why the path was rejected" assert_absent "$home/state/abort-notgit-dd4.meta" "aborted spawn must not record meta" # Abort: the pane resolves INTO the primary checkout (a subdir of PROJ_ABS). out=$(run_spawn "$home" abort-primary-ee5 "$proj" "$proj/sub" "$fakebin"); status=$? expect_code 1 "$status" "spawn landing inside the primary checkout should abort" - assert_contains "$out" "did not yield an isolated worktree" "primary-checkout spawn lacked the isolation error" + assert_contains "$out" "did not enter an isolated worktree" "primary-checkout spawn lacked the isolation error" + assert_contains "$out" "not a worktree root" "primary-checkout spawn did not say why the path was rejected" + assert_absent "$home/state/abort-primary-ee5.meta" "aborted spawn must not record meta" # Proceed: the pane resolves to a genuine, isolated worktree. out=$(run_spawn "$home" ok-isolated-ff6 "$proj" "$TMP_ROOT/spawn-wt" "$fakebin"); status=$? expect_code 0 "$status" "spawn into a genuine isolated worktree should succeed" assert_contains "$out" "spawned ok-isolated-ff6" "isolated spawn did not report success" - assert_not_contains "$out" "did not yield an isolated worktree" "isolated spawn wrongly tripped the guard" + assert_not_contains "$out" "isolated worktree" "isolated spawn wrongly tripped the guard" pass "fm-spawn: aborts unless the resolved worktree is a genuine, isolated worktree" } diff --git a/tests/fm-task-delivery.test.sh b/tests/fm-task-delivery.test.sh index e48e3e97647..bdace4b501e 100755 --- a/tests/fm-task-delivery.test.sh +++ b/tests/fm-task-delivery.test.sh @@ -33,6 +33,7 @@ make_home() { # <name> [<registry-line>...] projects="$TMP_ROOT/$name/projects" fakebin="$TMP_ROOT/$name/bin" mkdir -p "$home/data" "$home/state" "$home/config" "$projects/proj" "$fakebin" + git -C "$projects/proj" init -q || fail "could not initialize project fixture" printf '#!/bin/sh\nexit 1\n' > "$fakebin/tmux" chmod +x "$fakebin/tmux" if [ "$#" -gt 0 ]; then @@ -747,6 +748,49 @@ EOF pass "fm-spawn/fm-promote: leftover Task placeholders are refused until both subsections are filled" } +test_spawn_refreshes_legacy_worker_roles() { + local rec home proj fakebin kind id out brief project_kind + rec=$(make_home worker-roles) + IFS='|' read -r home proj fakebin <<EOF +$rec +EOF + # AGENTS.md and its import are instruction inputs, not implementation-source + # assertions: launching a worker must never rewrite either project's files. + cp "$ROOT/AGENTS.md" "$home/AGENTS.md" + for project_kind in firstmate unrelated; do + if [ "$project_kind" = firstmate ]; then + cp "$ROOT/AGENTS.md" "$proj/AGENTS.md" + else + printf 'Use this project coding standard.\n' > "$proj/AGENTS.md" + fi + printf '@AGENTS.md\n' > "$proj/CLAUDE.md" + cp "$proj/AGENTS.md" "$proj/agents-before" + for kind in no-mistakes direct-PR local-only scout; do + id="roles-$project_kind-$kind" + write_brief "$home" "$id" + if [ "$kind" = scout ]; then + out=$(run_spawn "$home" "$fakebin" "$id" "$proj" codex --scout) + else + out=$(run_spawn "$home" "$fakebin" "$id" "$proj" codex --mode "$kind" --yolo off) + fi + assert_not_contains "$out" 'could not render' "worker role rendering failed" + brief="$home/data/$id/launch-brief.md" + assert_present "$brief" "$project_kind $kind did not refresh the legacy launch brief" + assert_grep 'follow this brief instead of that supervisor contract' "$brief" "$project_kind $kind omitted worker authority" + assert_grep 'When this task works on Firstmate itself' "$brief" "$project_kind $kind made the exception unconditional" + assert_grep 'Other projects retain their own instructions unchanged' "$brief" "$project_kind $kind displaced project guidance" + ! grep -q '^This section supersedes every earlier brief instruction about your role' "$brief" || + fail "$project_kind $kind revoked the brief's own role for a task that is not Firstmate" + assert_no_grep '# Current worker role contract' "$home/data/$id/brief.md" "spawn rewrote the source brief" + cmp -s "$proj/agents-before" "$proj/AGENTS.md" || fail "spawn changed project AGENTS.md" + [ "$(cat "$proj/CLAUDE.md")" = '@AGENTS.md' ] || fail "spawn changed the project import" + done + done + cmp -s "$ROOT/AGENTS.md" "$home/AGENTS.md" || fail "worker spawn changed the primary contract" + pass "fm-spawn: every legacy worker receives scoped role instructions without changing project or primary instructions" +} + +test_spawn_refreshes_legacy_worker_roles test_ship_spawn_requires_a_valid_delivery_contract test_scout_and_secondmate_refuse_delivery_flags test_spawn_refuses_a_brief_mode_mismatch diff --git a/tests/fm-task-inbox.test.sh b/tests/fm-task-inbox.test.sh index 31a3f8ab088..1bad61ff79b 100644 --- a/tests/fm-task-inbox.test.sh +++ b/tests/fm-task-inbox.test.sh @@ -23,6 +23,9 @@ # pane, stays silent on a healthy/empty inbox, surfaces unwritable ladder # bookkeeping only while its record remains unhandled, and emits exactly # one stale wake once the ring budget is spent. +# 6. Dead panes: the doorbell line is a shell no-op when executed by a bare +# shell, the ring skips an agent the backend classifies dead, and the +# watcher surfaces such a record exactly once instead of re-ringing. set -u # shellcheck source=tests/wake-helpers.sh @@ -50,7 +53,10 @@ inbox_lib() { # <state> <function> [args...] # A fake tmux for the watcher cases: capture-pane replays FM_FAKE_TMUX_CAPTURE, # display-message yields a numeric cursor row, and every literal send-keys is -# logged to FM_SEND_LOG so a doorbell ring is observable. +# logged to FM_SEND_LOG so a doorbell ring is observable. With +# FM_FAKE_TMUX_AGENT set, the inventory lists window fm-t1 and its +# #{pane_current_command} answers with that value, so `zsh` makes +# fm_backend_tmux_agent_state read the pane as a dead bare shell. make_watch_stubs() { # <dir> -> echoes fakebin dir local dir=$1 fb="$1/fakebin" mkdir -p "$fb" @@ -76,7 +82,13 @@ case "${1:-}" in fi exit 0 ;; display-message) - for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done + for a in "$@"; do + case "$a" in + *cursor_y*) printf '1\n'; exit 0 ;; + *pane_current_command*) [ -z "${FM_FAKE_TMUX_AGENT:-}" ] || { printf '%s\n' "$FM_FAKE_TMUX_AGENT"; exit 0; } ;; + *pane_tty*) [ -z "${FM_FAKE_TMUX_AGENT:-}" ] || { printf '\n'; exit 0; } ;; + esac + done printf 'fakepane\n'; exit 0 ;; capture-pane) if [ -n "${FM_FAKE_TMUX_CAPTURE:-}" ] && [ -f "$FM_FAKE_TMUX_CAPTURE" ]; then @@ -85,7 +97,7 @@ case "${1:-}" in printf '╭────╮\n│ │\n╰────╯\n' fi exit 0 ;; - list-windows) exit 0 ;; + list-windows) [ "${FM_FAKE_TMUX_MISSING:-0}" = 1 ] || printf 'fm-t1\n'; exit 0 ;; esac exit 0 SH @@ -150,9 +162,9 @@ test_write_is_durable_and_exact() { doorbell2=$(inbox_lib "$state" fm_task_inbox_doorbell_line "$rec2") [ "$doorbell" = "$doorbell2" ] \ || fail "every record in one inbox should ring the same drain-all doorbell" - assert_contains "$doorbell" "$state/t1.inbox/*.msg" "doorbell should name all unhandled records" + assert_contains "$doorbell" "'$state/t1.inbox'/*.msg" "doorbell should quote and name all unhandled records" assert_contains "$doorbell" "numeric order" "doorbell should require ordered processing" - assert_contains "$doorbell" "$state/t1.inbox/handled/" "doorbell should name the handled dir" + assert_contains "$doorbell" "'$state/t1.inbox'/handled/" "doorbell should quote and name the handled dir" assert_contains "$doorbell" "Firstmate instruction waiting" "doorbell should be self-describing" case "$doorbell" in *$'\n'*) fail "the doorbell must be a single line" ;; @@ -160,6 +172,113 @@ test_write_is_durable_and_exact() { pass "inbox: a steer is written durably and round-trips byte-exact with a self-describing doorbell" } +# The doorbell may land in a pane whose agent has exited, where it is a shell +# command line. Execute the real line in real shells and assert it is inert: +# exit 0, no output, and nothing in the inbox touched. +test_doorbell_is_a_shell_noop() { + local state rec doorbell sh out before after marker + state="$TMP_ROOT/noop/x; touch marker; #'s space/state" + marker="$state/marker" + mkdir -p "$state" + rec=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "please continue") + doorbell=$(inbox_lib "$state" fm_task_inbox_doorbell_line "$rec") + case "$doorbell" in + ': '*) ;; + *) fail "the doorbell must start with the shell no-op prefix, got: $doorbell" ;; + esac + assert_contains "$doorbell" "'\\''s space/state/t1.inbox'" \ + "the doorbell should escape an embedded single quote in its quoted path" + before=$(ls -R "$state/t1.inbox") + for sh in sh bash zsh; do + command -v "$sh" >/dev/null 2>&1 || continue + out=$(cd "$state" && "$sh" -c "$doorbell" 2>&1) \ + || fail "$sh executed the hostile-path doorbell with a non-zero status: $out" + [ -z "$out" ] || fail "$sh produced output while executing the hostile-path doorbell: $out" + [ ! -e "$marker" ] || fail "$sh executed shell syntax embedded in the inbox path" + done + # An interactive-style zsh with the line fed on stdin, the closest portable + # stand-in for a dead pane's login shell reading typed keystrokes. + if command -v zsh >/dev/null 2>&1; then + out=$(cd "$state" && printf '%s\n' "$doorbell" | zsh -s 2>&1) \ + || fail "zsh reading the hostile-path doorbell from stdin failed: $out" + [ -z "$out" ] || fail "zsh printed while reading the hostile-path doorbell: $out" + [ ! -e "$marker" ] || fail "zsh executed shell syntax from the stdin doorbell" + fi + after=$(ls -R "$state/t1.inbox") + [ "$before" = "$after" ] || fail "executing the doorbell changed the inbox:"$'\n'"$after" + [ -f "$rec" ] || fail "executing the doorbell removed the unhandled record" + pass "inbox: a hostile-path doorbell executes as a no-op in bare shells" +} + +test_doorbell_rejects_terminal_controls() { + local dir state rec doorbell control label log marker rc + dir="$TMP_ROOT/control-path" + marker="$dir/marker" + mkdir -p "$dir" + make_watch_stubs "$dir" >/dev/null + for label in etx esc; do + case "$label" in + etx) control=$'\003' ;; + esc) control=$'\033' ;; + esac + state="$dir/${control}touch marker; # $label/state" + mkdir -p "$state" + rec=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "please continue") + doorbell= + rc=0 + doorbell=$(inbox_lib "$state" fm_task_inbox_doorbell_line "$rec") || rc=$? + [ "$rc" -ne 0 ] || fail "a $label path should make doorbell construction fail" + [ -z "$doorbell" ] || fail "a rejected $label path emitted doorbell bytes" + log="$dir/$label.send.log"; : > "$log" + rc=0 + PATH="$dir/fakebin:$PATH" FM_SEND_LOG="$log" \ + inbox_lib "$state" fm_task_inbox_ring tmux sess:fm-t1 "$rec" fm-t1 || rc=$? + [ "$rc" = 2 ] || fail "a rejected $label path should return send-failed status 2, got $rc" + [ ! -s "$log" ] || fail "a $label path reached send-keys:"$'\n'"$(cat "$log")" + [ ! -e "$marker" ] || fail "a $label path executed its crafted command" + [ -f "$rec" ] || fail "rejecting a $label path removed the durable record" + done + pass "inbox: terminal-control paths are rejected without typing" +} + +# fm_task_inbox_ring against a backend whose agent classifies dead or missing: +# nothing is typed and the distinct return code lets callers route to recovery. +# An unreadable endpoint still rings, so a blind classifier never starves a +# live worker. +test_ring_skips_dead_agent() { + local dir state rec log rc + dir="$TMP_ROOT/ring-dead" + state="$dir/state" + mkdir -p "$state" + make_watch_stubs "$dir" >/dev/null + rec=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "please continue") + log="$dir/send.log"; : > "$log" + rc=0 + PATH="$dir/fakebin:$PATH" FM_SEND_LOG="$log" FM_FAKE_TMUX_AGENT=zsh \ + inbox_lib "$state" fm_task_inbox_ring tmux sess:fm-t1 "$rec" fm-t1 || rc=$? + [ "$rc" = 3 ] || fail "a dead agent should return 3 from the ring, got $rc" + [ ! -s "$log" ] || fail "a dead pane was typed into:"$'\n'"$(cat "$log")" + [ -f "$rec" ] || fail "skipping the ring must leave the durable record in place" + rc=0 + PATH="$dir/fakebin:$PATH" FM_SEND_LOG="$log" FM_FAKE_TMUX_MISSING=1 \ + inbox_lib "$state" fm_task_inbox_ring tmux sess:fm-t1 "$rec" fm-t1 || rc=$? + [ "$rc" = 3 ] || fail "a missing endpoint should return 3 from the ring, got $rc" + [ ! -s "$log" ] || fail "a missing endpoint was typed into:"$'\n'"$(cat "$log")" + [ -f "$rec" ] || fail "skipping a missing endpoint must leave the durable record in place" + rc=0 + PATH="$dir/fakebin:$PATH" FM_SEND_LOG="$log" FM_FAKE_TMUX_AGENT=claude \ + inbox_lib "$state" fm_task_inbox_ring tmux sess:fm-t1 "$rec" fm-t1 || rc=$? + [ "$rc" = 0 ] || fail "a live agent should still be rung, got $rc" + grep -qF 'Firstmate instruction waiting' "$log" || fail "a live agent did not receive the doorbell" + : > "$log" + rc=0 + PATH="$dir/fakebin:$PATH" FM_SEND_LOG="$log" \ + inbox_lib "$state" fm_task_inbox_ring tmux sess:fm-t1 "$rec" fm-t1 || rc=$? + [ "$rc" = 0 ] || fail "an endpoint the classifier cannot see should still be rung, got $rc" + grep -qF 'Firstmate instruction waiting' "$log" || fail "an unclassifiable endpoint did not receive the doorbell" + pass "inbox: the ring skips dead or missing endpoints and still rings live or unclassifiable endpoints" +} + test_idempotent_write_dedups_exact_body() { local state r1 r2 r3 r4 count text state="$TMP_ROOT/idem/state"; mkdir -p "$state" @@ -391,7 +510,7 @@ test_watcher_rerings_idle_pane_quietly() { sleep 0.1 i=$((i + 1)) done - grep -qF "Firstmate instruction waiting: list $state/t1.inbox/*.msg" "$log" \ + grep -qF "Firstmate instruction waiting: list '$state/t1.inbox'/*.msg" "$log" \ || { kill "$pid" 2>/dev/null; fail "the watcher never re-rang the doorbell:"$'\n'"$(cat "$log")"; } kill -0 "$pid" 2>/dev/null \ || fail "a healthy re-ring must not wake firstmate (watcher exited):"$'\n'"$(cat "$out")" @@ -522,7 +641,61 @@ test_watcher_escalates_once_after_budget() { pass "watcher: a spent ring budget emits exactly one ordinary stale wake for recovery" } +test_watcher_dead_pane_escalates_once_without_ringing() { + local dir state out log pid rec + dir=$(setup_watch_case dead-pane) + state="$dir/state"; out="$dir/watch.out"; log="$dir/send.log"; : > "$log" + rec=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "please continue") + age_path "$rec" + watch_bg "$state" "$dir/fakebin" "$out" \ + FM_SEND_LOG="$log" FM_FAKE_TMUX_CAPTURE="$(idle_capture "$dir")" \ + FM_FAKE_TMUX_AGENT=zsh FM_TASK_INBOX_RING_MAX=99 + pid=$! + wait_watcher_gone "$pid" \ + || { kill "$pid" 2>/dev/null; fail "the watcher never surfaced a dead pane's unhandled instruction"; } + [ ! -s "$log" ] || fail "a dead pane was typed into:"$'\n'"$(cat "$log")" + [ "$(grep -cF 'unread firstmate instruction' "$state/.wake-queue" 2>/dev/null || true)" = 1 ] \ + || fail "a dead pane should surface exactly one stale wake:"$'\n'"$(cat "$state/.wake-queue" 2>/dev/null)" + grep -qF "agent has exited" "$state/.wake-queue" \ + || fail "the stale wake should say the agent has exited:"$'\n'"$(cat "$state/.wake-queue")" + grep -qF "$rec" "$state/.wake-queue" || fail "the stale wake should name the record path" + [ -f "$rec" ] || fail "the durable record must survive for recovery" + [ "$(cat "$state/t1.inbox/.escalated")" = "${rec##*/}" ] \ + || fail "the escalation marker should suppress further surfacing of this record" + [ ! -e "$state/t1.inbox/.ring-state" ] || fail "a dead pane must not enter the re-ring ladder" + # The ladder is capped: nothing further is due for this record, so no later + # poll rings the dead pane or queues a second wake. + [ "$(inbox_lib "$state" fm_task_inbox_due_action "$state" t1)" = quiet ] \ + || fail "a dead pane already surfaced must be quiet on later polls" + pass "watcher: a positively dead pane is never typed into and surfaces exactly one stale wake" +} + +test_watcher_dead_pane_ignores_stale_busy_state() { + local dir state out log pid rec + dir=$(setup_watch_case dead-pane-busy) + state="$dir/state"; out="$dir/watch.out"; log="$dir/send.log"; : > "$log" + printf 'some output\nBUSYTOKEN active\n' > "$dir/busy.capture" + rec=$(inbox_lib "$state" fm_task_inbox_write "$state" t1 "please continue") + age_path "$rec" + watch_bg "$state" "$dir/fakebin" "$out" \ + FM_SEND_LOG="$log" FM_FAKE_TMUX_CAPTURE="$dir/busy.capture" \ + FM_FAKE_TMUX_AGENT=zsh FM_BUSY_REGEX=BUSYTOKEN FM_TASK_INBOX_RING_MAX=99 + pid=$! + wait_watcher_gone "$pid" \ + || { kill "$pid" 2>/dev/null; fail "stale busy state hid a dead pane's unhandled instruction"; } + [ ! -s "$log" ] || fail "a busy-marked dead pane was typed into:"$'\n'"$(cat "$log")" + [ "$(grep -cF 'unread firstmate instruction' "$state/.wake-queue" 2>/dev/null || true)" = 1 ] \ + || fail "a busy-marked dead pane should surface exactly once:"$'\n'"$(cat "$state/.wake-queue" 2>/dev/null)" + [ -f "$rec" ] || fail "the durable record must survive stale busy-state recovery" + [ "$(cat "$state/t1.inbox/.escalated")" = "${rec##*/}" ] \ + || fail "stale busy-state recovery should suppress repeated surfacing" + pass "watcher: dead-pane recovery overrides stale busy state" +} + test_write_is_durable_and_exact +test_doorbell_is_a_shell_noop +test_doorbell_rejects_terminal_controls +test_ring_skips_dead_agent test_idempotent_write_dedups_exact_body test_idempotent_write_follows_concurrent_ack test_handled_mv_dedups_by_sequence @@ -537,3 +710,5 @@ test_watcher_quiet_on_healthy_inbox test_watcher_ack_silences_unwritable_ladder test_watcher_surfaces_unwritable_ladder test_watcher_escalates_once_after_budget +test_watcher_dead_pane_escalates_once_without_ringing +test_watcher_dead_pane_ignores_stale_busy_state diff --git a/tests/fm-teardown-endpoint-safety.test.sh b/tests/fm-teardown-endpoint-safety.test.sh index b08420fc51d..2d185907abb 100755 --- a/tests/fm-teardown-endpoint-safety.test.sh +++ b/tests/fm-teardown-endpoint-safety.test.sh @@ -1,5 +1,5 @@ #!/usr/bin/env bash -# Regression tests for cleanup endpoint identity validation. +# Regression tests for cleanup endpoint and worktree-slot identity validation. set -u # shellcheck source=tests/lib.sh @@ -14,6 +14,7 @@ make_case() { # <name> mkdir -p "$TMP_ROOT/$dir/home/state" "$TMP_ROOT/$dir/home/data" \ "$TMP_ROOT/$dir/home/config" "$TMP_ROOT/$dir/fakebin" \ "$TMP_ROOT/$dir/worktree" "$TMP_ROOT/$dir/project" + git init -q "$TMP_ROOT/$dir/project" : > "$TMP_ROOT/$dir/worktree/sentinel" : > "$TMP_ROOT/$dir/runtime.log" cat > "$TMP_ROOT/$dir/fakebin/tmux" <<'SH' @@ -34,6 +35,19 @@ SH printf '%s\n' "$TMP_ROOT/$dir" } +mark_case_as_treehouse_pool() { # <case> + local dir=$1 + rm -rf "$dir/worktree" + mkdir -p "$dir/pool/1" + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid \ + commit --allow-empty -qm pool-fixture + git -C "$dir/project" worktree add -q --detach "$dir/pool/1/project" + ln -s "pool/1/project" "$dir/worktree" + printf '{"worktrees":[{"name":"1","path":"%s"}]}\n' \ + "$dir/pool/1/project" > "$dir/pool/treehouse-state.json" + : > "$dir/worktree/sentinel" +} + run_case() { # <case> <id> local dir=$1 id=$2 FM_HOME="$dir/home" FM_ROOT_OVERRIDE="$ROOT" \ @@ -137,6 +151,42 @@ test_control_lock_contention_refuses_before_mutation() { pass "fm-teardown: a concurrent lifecycle action refuses before mutation" } +test_non_pool_teardown_ignores_task_set_lock() { + local dir id=non-pool-task lock ready holder i=0 + dir=$(make_case non-pool-task-set-lock) + fm_write_meta "$dir/home/state/$id.meta" \ + "window=isolated:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/missing-worktree" "project=$dir/project" "kind=scout" + lock="$dir/home/state/.task-set.lock" + ready="$dir/task-set-lock-ready" + ( + # shellcheck source=/dev/null + . "$ROOT/bin/fm-wake-lib.sh" + fm_lock_try_acquire "$lock" || exit 1 + trap 'fm_lock_release "$lock"' EXIT + : > "$ready" + sleep 30 + ) & + holder=$! + while [ ! -e "$ready" ] && [ "$i" -lt 100 ]; do + sleep 0.1 + i=$((i + 1)) + done + [ -e "$ready" ] || { + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + fail "could not stage an in-progress task publication" + } + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "non-pool teardown was blocked by an unrelated task publication: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "non-pool teardown left task metadata" + assert_present "$lock" "non-pool teardown removed the publisher's lock" + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + pass "fm-teardown: non-pool cleanup ignores unrelated task publication locks" +} + test_metadata_lock_serializes_destructive_cleanup() { local dir id=metadata-locked-task lock ready release holder teardown_pid i=0 rc dir=$(make_case metadata-lock) @@ -365,10 +415,430 @@ SH pass "fm-teardown: exact tmux cleanup preserves invalid and prefix-matched neighbors while removing only the recorded target" } +test_bare_relative_origin_shares_project_lock_with_clone() { + local dir second_project primary_lock clone_lock + dir=$(make_case bare-relative-origin-lock) + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid \ + commit --allow-empty -qm lock-fixture + mkdir -p "$dir/project/remotes" + git clone -q --bare "$dir/project" "$dir/project/remotes/origin.git" + git -C "$dir/project" remote add origin remotes/origin.git + second_project="$dir/second-project" + git clone -q "$dir/project/remotes/origin.git" "$second_project" + + primary_lock=$(FM_HOME="$dir/home" bash -c \ + '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$dir/project") \ + || fail "could not resolve the primary project's bare-origin lock" + clone_lock=$(FM_HOME="$dir/home" bash -c \ + '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$second_project") \ + || fail "could not resolve the clone project's absolute-origin lock" + [ "$primary_lock" = "$clone_lock" ] \ + || fail "bare and absolute forms of the same local origin resolved different project locks" + + pass "Treehouse locking resolves a bare local origin against its source project, matching the provisioned clone" +} + +test_reused_pool_slot_refuses_before_touching_the_other_task() { + local dir id=stale-task other=live-task worker rc + + dir=$(make_case slot-reuse) + mark_case_as_treehouse_pool "$dir" + # The reuse collision: the pool slot recorded for a finished task has already + # been handed to another task, whose worker is live in it right now. + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + fm_write_meta "$dir/home/state/$other.meta" \ + "window=firstmate:fm-$other" "endpoint_task_id=$other" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + # Staged in this shell, not a command substitution: a background child of a + # $(...) subshell does not outlive it, and the point of this worker is to be + # alive in the slot while teardown runs. + ( cd "$dir/worktree" && exec sleep 30 ) & + worker=$! + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + + [ "$rc" -ne 0 ] || fail "teardown returned a pool slot a second task record still holds" + kill -0 "$worker" 2>/dev/null || fail "teardown killed the worker holding the reused pool slot" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot a second task record still holds" + assert_present "$dir/home/state/$other.meta" "teardown removed the live task's record" + assert_present "$dir/home/state/$id.meta" "teardown removed the stale task's record before refusing" + [ ! -s "$dir/runtime.log" ] \ + || fail "teardown reached the runtime on a contested pool slot: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "refusal should name the other task holding the slot" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + + # The same collision recorded on a secondmate home field rather than a task + # worktree is the same slot, and refuses the same way. + dir=$(make_case slot-reuse-home) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + fm_write_meta "$dir/home/state/$other.meta" \ + "window=firstmate:fm-$other" "endpoint_task_id=$other" \ + "worktree=$dir/worktree" "home=$dir/worktree" \ + "project=$dir/project" "kind=secondmate" + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown returned a pool slot a secondmate home record still holds" + assert_present "$dir/worktree/sentinel" "teardown reset a pool slot a secondmate home record still holds" + assert_present "$dir/home/state/$other.meta" "teardown removed the secondmate record" + [ ! -s "$dir/runtime.log" ] \ + || fail "teardown reached the runtime on a slot held by a secondmate home: $(cat "$dir/runtime.log")" + + pass "fm-teardown: a pool slot named by a second task record is never returned, killed, or reset" +} + +test_cross_home_pool_slot_collision_refuses() { + local dir id=stale-task other=secondmate-task second_home second_project rc + dir=$(make_case slot-reuse-cross-home) + mark_case_as_treehouse_pool "$dir" + printf 'fixture\n' > "$dir/project/tracked" + git -C "$dir/project" add tracked + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid commit -qm fixture + second_home="$dir/secondmate-home" + second_project="$second_home/projects/project" + mkdir -p "$second_home/projects" "$second_home/state" "$second_home/data" + git clone -q "$dir/project" "$second_project" + printf '%s\n' "- mate - fixture (home: $second_home; scope: test; projects: project; added 2026-01-01)" \ + > "$dir/home/data/secondmates.md" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + fm_write_meta "$second_home/state/$other.meta" \ + "window=firstmate:fm-$other" "endpoint_task_id=$other" \ + "worktree=$dir/worktree" "project=$second_project" "kind=scout" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "teardown returned a pool slot held by another firstmate home" + assert_present "$dir/home/state/$id.meta" "cross-home collision removed stale metadata" + assert_present "$second_home/state/$other.meta" "cross-home collision removed live metadata" + assert_present "$dir/worktree/sentinel" "cross-home collision reset the shared slot" + [ ! -s "$dir/runtime.log" ] \ + || fail "cross-home collision reached the runtime: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "cross-home refusal should name the task holding the slot" + pass "fm-teardown: a pool slot held by another firstmate home is never returned" +} + +test_sole_slot_record_still_tears_down() { + local dir id=sole-task worker + + dir=$(make_case slot-sole) + mark_case_as_treehouse_pool "$dir" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + # A neighbouring task on its OWN slot must not look like a collision. + mkdir -p "$dir/other-worktree" + fm_write_meta "$dir/home/state/neighbour.meta" \ + "window=firstmate:fm-neighbour" "endpoint_task_id=neighbour" \ + "worktree=$dir/other-worktree" "project=$dir/project" "kind=scout" + ( cd "$dir/other-worktree" && exec sleep 30 ) & + worker=$! + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown of a task that solely holds its slot failed: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "uncontested teardown left the task record" + assert_present "$dir/home/state/neighbour.meta" "uncontested teardown removed the neighbour's record" + kill -0 "$worker" 2>/dev/null || fail "uncontested teardown killed a worker in a different slot" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "uncontested teardown did not return its own pool slot: $(cat "$dir/runtime.log")" + kill "$worker" 2>/dev/null || true + wait "$worker" 2>/dev/null || true + pass "fm-teardown: a task that solely holds its slot still returns it" +} + +test_recorded_endpoint_that_changed_directory_still_tears_down() { + local dir id=moved-task + + dir=$(make_case slot-endpoint-moved) + mark_case_as_treehouse_pool "$dir" + mkdir -p "$dir/other-directory" + # The exact recorded worker may legitimately cd outside its worktree. Its + # endpoint identity still owns the lifecycle; cwd alone must not brick it. + cat > "$dir/fakebin/tmux" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = display-message ]; then + printf '%s\n' '$dir/other-directory' + exit 0 +fi +printf 'tmux' >> "\${FM_RUNTIME_LOG:?}" +printf ' <%s>' "\$@" >> "\${FM_RUNTIME_LOG:?}" +printf '\n' >> "\${FM_RUNTIME_LOG:?}" +exit 0 +SH + chmod +x "$dir/fakebin/tmux" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" \ + || fail "teardown refused its recorded endpoint after it changed directory: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "moved-endpoint teardown left the task record" + grep -Fq "tmux <kill-window> <-t> <=firstmate:=fm-$id>" "$dir/runtime.log" \ + || fail "moved-endpoint teardown did not stop the exact recorded worker: $(cat "$dir/runtime.log")" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "moved-endpoint teardown did not return its uncontested pool slot: $(cat "$dir/runtime.log")" + + pass "fm-teardown: an exact recorded endpoint still tears down after changing cwd outside its worktree" +} + +# --- Treehouse project-lock anchoring across home layouts -------------------- +# +# The lock is anchored at the local root home, so every home on this machine +# that can reach the same pool must derive the identical file. A remote parent +# binding terminates that walk at the home holding it: its parent is on another +# machine and can neither hold nor observe a lock taken here. + +write_local_parent_record() { # <home> <parent-home> + cat > "$1/.fm-secondmate-parent" <<REC +schema=fm-secondmate-parent.v1 +route=local +parent_home=$2 +REC +} + +write_remote_parent_record() { # <home> + cat > "$1/.fm-secondmate-parent" <<'REC' +schema=fm-secondmate-parent.v1 +route=remote +parent_host=machine-a +REC +} + +make_home() { # <path> + mkdir -p "$1/state" "$1/data" "$1/config" "$1/projects" +} + +resolve_project_lock() { # <home> <project> + FM_HOME="$1" bash -c '. "$1"; fm_treehouse_project_lock_path "$2"' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$2" +} + +test_project_lock_anchors_at_the_local_root_across_home_layouts() { + local dir main_home main_project local_mate remote_mate remote_child + local main_lock mate_lock remote_lock child_lock orphan_lock rc + dir=$(make_case project-lock-anchoring) + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid \ + commit --allow-empty -qm anchor-fixture + + # Main-home layout: a root home and a local secondmate beneath it. + main_home="$dir/home" + main_project="$main_home/projects/project" + make_home "$main_home" + git clone -q "$dir/project" "$main_project" + local_mate="$dir/local-mate" + make_home "$local_mate" + write_local_parent_record "$local_mate" "$main_home" + git clone -q "$dir/project" "$local_mate/projects/project" + + # Remote layout: a home seeded from another machine, plus its own local child. + remote_mate="$dir/remote-mate" + make_home "$remote_mate" + write_remote_parent_record "$remote_mate" + git clone -q "$dir/project" "$remote_mate/projects/project" + remote_child="$dir/remote-mate-child" + make_home "$remote_child" + write_local_parent_record "$remote_child" "$remote_mate" + git clone -q "$dir/project" "$remote_child/projects/project" + + main_lock=$(resolve_project_lock "$main_home" "$main_project") \ + || fail "the root home could not resolve its project lock" + mate_lock=$(resolve_project_lock "$local_mate" "$local_mate/projects/project") \ + || fail "a local secondmate home could not resolve its project lock" + remote_lock=$(resolve_project_lock "$remote_mate" "$remote_mate/projects/project") \ + || fail "a remote-seeded secondmate home could not resolve its project lock" + child_lock=$(resolve_project_lock "$remote_child" "$remote_child/projects/project") \ + || fail "a local child of a remote-seeded home could not resolve its project lock" + + [ "$main_lock" = "$mate_lock" ] \ + || fail "the root home and its local secondmate derived different project locks" + [ "$remote_lock" = "$child_lock" ] \ + || fail "a remote-seeded home and its local child derived different project locks" + case "$remote_lock" in + "$remote_mate/state/"*) ;; + *) fail "a remote-seeded home anchored its project lock outside its own state: $remote_lock" ;; + esac + + # An origin-less local-only project still resolves, keyed on its worktree top. + git init -q "$remote_mate/projects/local-only" + orphan_lock=$(resolve_project_lock "$remote_mate" "$remote_mate/projects/local-only") \ + || fail "an origin-less local-only project could not resolve its lock in a remote-seeded home" + [ "$orphan_lock" != "$remote_lock" ] \ + || fail "an origin-less project shared the lock identity of an unrelated origin" + + # Everything other than a remote route still fails closed. + printf 'schema=fm-secondmate-parent.v1\nroute=sideways\n' \ + > "$remote_child/.fm-secondmate-parent" + set +e + resolve_project_lock "$remote_child" "$remote_child/projects/project" >/dev/null 2>&1 + rc=$? + set -e + [ "$rc" -ne 0 ] || fail "an unsupported parent route resolved a project lock instead of refusing" + + pass "Treehouse project locking anchors at the local root for main-home, local-secondmate, and remote-seeded layouts" +} + +test_remote_seeded_home_returns_its_uncontested_slot() { + local dir id=remote-task rc + dir=$(make_case remote-home-teardown) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -eq 0 ] \ + || fail "teardown in a remote-seeded home refused its own uncontested slot: $(cat "$dir/stderr")" + assert_absent "$dir/home/state/$id.meta" "remote-seeded teardown left the task record" + grep -Fq "treehouse <return>" "$dir/runtime.log" \ + || fail "remote-seeded teardown did not return its own pool slot: $(cat "$dir/runtime.log")" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded Treehouse teardown command\n' + printf '$ FM_HOME=%s bin/fm-teardown.sh %s --force\n' "$dir/home" "$id" + printf 'stdout:\n'; cat "$dir/stdout" + printf 'stderr:\n'; cat "$dir/stderr" + printf 'exit=%s\nruntime calls:\n' "$rc"; cat "$dir/runtime.log" + printf 'task metadata=%s\nslot sentinel=%s\n' \ + "$([ -e "$dir/home/state/$id.meta" ] && printf present || printf removed)" \ + "$([ -e "$dir/worktree/sentinel" ] && printf present || printf removed)" + fi + + pass "fm-teardown: a remote-seeded secondmate home returns its own uncontested pool slot" +} + +test_remote_seeded_home_still_refuses_a_slot_its_child_holds() { + local dir id=remote-stale other=child-task child_home child_project rc + dir=$(make_case remote-home-collision) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + printf 'fixture\n' > "$dir/project/tracked" + git -C "$dir/project" add tracked + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid commit -qm fixture + child_home="$dir/child-home" + child_project="$child_home/projects/project" + make_home "$child_home" + write_local_parent_record "$child_home" "$dir/home" + git clone -q "$dir/project" "$child_project" + printf '%s\n' "- mate - fixture (home: $child_home; scope: test; projects: project; added 2026-01-01)" \ + > "$dir/home/data/secondmates.md" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + fm_write_meta "$child_home/state/$other.meta" \ + "window=firstmate:fm-$other" "endpoint_task_id=$other" \ + "worktree=$dir/worktree" "project=$child_project" "kind=scout" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + [ "$rc" -ne 0 ] \ + || fail "a remote-seeded home returned a pool slot its own local child still holds" + assert_present "$dir/home/state/$id.meta" "remote-layout collision removed stale metadata" + assert_present "$child_home/state/$other.meta" "remote-layout collision removed live metadata" + assert_present "$dir/worktree/sentinel" "remote-layout collision reset the shared slot" + assert_contains "$(cat "$dir/stderr")" "$other" \ + "remote-layout refusal should name the task holding the slot" + if [ "${FM_TEST_EVIDENCE:-0}" = 1 ]; then + printf '# remote-seeded cross-home collision command\n' + printf '$ FM_HOME=%s bin/fm-teardown.sh %s --force\n' "$dir/home" "$id" + printf 'stderr:\n'; cat "$dir/stderr" + printf 'exit=%s\nruntime calls=%s\n' "$rc" \ + "$([ -s "$dir/runtime.log" ] && cat "$dir/runtime.log" || printf none)" + printf 'remote metadata=%s\nchild metadata=%s\nslot sentinel=%s\n' \ + "$([ -e "$dir/home/state/$id.meta" ] && printf preserved || printf removed)" \ + "$([ -e "$child_home/state/$other.meta" ] && printf preserved || printf removed)" \ + "$([ -e "$dir/worktree/sentinel" ] && printf preserved || printf removed)" + fi + + pass "fm-teardown: slot ownership across a remote-seeded home and its local child still refuses" +} + +test_remote_layout_homes_serialize_on_one_project_lock() { + local dir id=remote-serialize child_home child_project lock holder rc waited=0 + dir=$(make_case remote-lock-exclusion) + mark_case_as_treehouse_pool "$dir" + write_remote_parent_record "$dir/home" + printf 'fixture\n' > "$dir/project/tracked" + git -C "$dir/project" add tracked + git -C "$dir/project" -c user.name=test -c user.email=test@example.invalid commit -qm fixture + child_home="$dir/child-home" + child_project="$child_home/projects/project" + make_home "$child_home" + write_local_parent_record "$child_home" "$dir/home" + git clone -q "$dir/project" "$child_project" + fm_write_meta "$dir/home/state/$id.meta" \ + "window=firstmate:fm-$id" "endpoint_task_id=$id" \ + "worktree=$dir/worktree" "project=$dir/project" "kind=scout" + + # The local child takes the lock its own home derives and stays alive holding + # it, standing in for a slot allocation running in that home right now. + lock=$(resolve_project_lock "$child_home" "$child_project") \ + || fail "the local child could not resolve the shared project lock" + FM_HOME="$child_home" bash -c \ + '. "$1"; fm_lock_try_acquire "$2" || exit 1; : > "$3"; exec sleep 30' _ \ + "$ROOT/bin/fm-wake-lib.sh" "$lock" "$dir/lock-held" & + holder=$! + while [ ! -e "$dir/lock-held" ] && [ "$waited" -lt 100 ]; do + kill -0 "$holder" 2>/dev/null || break + sleep 0.1 + waited=$((waited + 1)) + done + [ -e "$dir/lock-held" ] || fail "the local child never took the shared project lock" + + set +e + run_case "$dir" "$id" > "$dir/stdout" 2> "$dir/stderr" + rc=$? + set -e + kill "$holder" 2>/dev/null || true + wait "$holder" 2>/dev/null || true + + [ "$rc" -ne 0 ] \ + || fail "a remote-seeded home returned a pool slot while its local child held the shared lock" + assert_present "$dir/home/state/$id.meta" "contended remote-layout teardown removed the task record" + assert_present "$dir/worktree/sentinel" "contended remote-layout teardown reset the slot" + [ ! -s "$dir/runtime.log" ] \ + || fail "contended remote-layout teardown reached the runtime: $(cat "$dir/runtime.log")" + assert_contains "$(cat "$dir/stderr")" "another Treehouse slot allocation or return is in progress" \ + "the refusal should name the shared project lock, not some unrelated check" + + pass "Treehouse project locking still serializes two homes across the remote-seeded boundary" +} + test_invalid_endpoint_records_refuse_before_mutation test_control_lock_contention_refuses_before_mutation +test_non_pool_teardown_ignores_task_set_lock test_metadata_lock_serializes_destructive_cleanup test_supported_backend_endpoint_records_validate test_tmux_empty_target_refuses_without_invocation test_recorded_process_identity_cleanup_is_exact test_isolated_tmux_invalid_and_valid_cleanup +test_bare_relative_origin_shares_project_lock_with_clone +test_reused_pool_slot_refuses_before_touching_the_other_task +test_cross_home_pool_slot_collision_refuses +test_sole_slot_record_still_tears_down +test_recorded_endpoint_that_changed_directory_still_tears_down +test_project_lock_anchors_at_the_local_root_across_home_layouts +test_remote_seeded_home_returns_its_uncontested_slot +test_remote_seeded_home_still_refuses_a_slot_its_child_holds +test_remote_layout_homes_serialize_on_one_project_lock diff --git a/tests/fm-teardown.test.sh b/tests/fm-teardown.test.sh index bfb22baf11c..62e609e081a 100755 --- a/tests/fm-teardown.test.sh +++ b/tests/fm-teardown.test.sh @@ -38,6 +38,10 @@ # (o) fm-pr-check rerun after HEAD moved -> no stale pr_head # (p) fm-pr-check when local HEAD lags -> record remote PR head # (q) no-mistakes + NO pr= recorded, PR discovered by branch -> ALLOW (yolo/no-CI merge) +# (q2) no-mistakes + squash-merged, local followed pipeline rebase -> ALLOW +# (q3) no-mistakes + squash-merged, same file, different content -> REFUSE +# (q4) no-mistakes + squash-merged rebased local plus extra commit -> REFUSE +# (q5) gh down + squash-merged stale local, content not in default -> REFUSE # # Also covers backlog teardown-lock-race: a git index.lock left in the worktree by a # killed crew process (bin/fm-teardown.sh's teardown_treehouse_return). @@ -165,7 +169,7 @@ SH # Bare origin so the clone has an `origin` remote and origin/HEAD. git init -q --bare "$case_dir/origin.git" - git -C "$case_dir/origin.git" symbolic-ref HEAD refs/heads/main + git --git-dir="$case_dir/origin.git" symbolic-ref HEAD refs/heads/main # Seed origin with one commit BEFORE cloning so the clone is not empty. git clone -q "$case_dir/origin.git" "$case_dir/_seed" 2>/dev/null git -C "$case_dir/_seed" -c user.email=t@t -c user.name=t \ @@ -271,6 +275,94 @@ SH chmod +x "$case_dir/fakebin/gh-axi" "$case_dir/fakebin/gh" } +# Squash-merged history whose pipeline rebased the branch onto a newer main that +# edited the same shared file. A local copy left behind by that rebase holds +# different content for the shared file, so its per-commit patch ids against the +# PR head differ and merge-tree against main conflicts; teardown refuses it on +# purpose rather than reading a shared path as proof the local content landed. +# local_mode: rebased | stale | rebased-plus-unlanded +# Echoes: <pr_head> +setup_squash_rebased_history() { + local case_dir=$1 local_mode=$2 tmp local_head pr_head + tmp="$case_dir/_shared_base" + git clone -q "$case_dir/origin.git" "$tmp" + printf '%s\n' base > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "shared base" + git -C "$tmp" push -q origin main + git -C "$case_dir/wt" fetch -q origin + git -C "$case_dir/wt" reset -q --hard origin/main + rm -rf "$tmp" + + wt_commit_file "$case_dir" feature.txt hello "add feature" + printf '%s\n' base feature-edit > "$case_dir/wt/shared.txt" + git -C "$case_dir/wt" add -- shared.txt + git -C "$case_dir/wt" -c user.email=t@t -c user.name=t \ + commit -q -m "edit shared from feature" + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + + tmp="$case_dir/_main_move" + git clone -q "$case_dir/origin.git" "$tmp" + printf '%s\n' base main-edit > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "main edits shared" + git -C "$tmp" push -q origin main + rm -rf "$tmp" + + tmp="$case_dir/_pipeline" + git clone -q "$case_dir/origin.git" "$tmp" + git -C "$tmp" checkout -q -b fm/task-x1 + printf '%s\n' hello > "$tmp/feature.txt" + git -C "$tmp" add -- feature.txt + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "add feature" + printf '%s\n' base main-edit feature-edit > "$tmp/shared.txt" + git -C "$tmp" add -- shared.txt + git -C "$tmp" -c user.email=t@t -c user.name=t \ + commit -q -m "edit shared from feature" + pr_head=$(git -C "$tmp" rev-parse HEAD) + git -C "$tmp" push -q origin "HEAD:refs/pull/7/head" + git -C "$tmp" checkout -q main + git -C "$tmp" merge -q --squash fm/task-x1 >/dev/null + git -C "$tmp" -c user.email=t@t -c user.name=t commit -q -m "feat: squash (#7)" + git -C "$tmp" push -q origin main + rm -rf "$tmp" + + git -C "$case_dir/project" fetch -q origin + git -C "$case_dir/wt" fetch -q origin "refs/pull/7/head:refs/fm-test/pr-head" + case "$local_mode" in + rebased) + git -C "$case_dir/wt" reset -q --hard "$pr_head" + ;; + stale) + git -C "$case_dir/wt" reset -q --hard "$local_head" + ;; + rebased-plus-unlanded) + git -C "$case_dir/wt" reset -q --hard "$pr_head" + wt_commit_file "$case_dir" later.txt local-only "local follow-up" + ;; + *) + fail "setup_squash_rebased_history: unknown local_mode $local_mode" + ;; + esac + printf '%s\n' "$pr_head" +} + +# A refusal must leave every recovery route intact: the isolated copy, its task +# branch still at the unlanded commit, and the durable task record. A completed +# teardown detaches and deletes that branch and removes the record, so these hold +# only while nothing destructive ran before the refusal was reported. +# Args: case_dir label head-before-teardown +assert_refusal_retained_task_state() { + local case_dir=$1 label=$2 head=$3 + [ -d "$case_dir/wt" ] || fail "$label: refusal removed the isolated copy" + [ "$(git -C "$case_dir/wt" rev-parse --abbrev-ref HEAD 2>/dev/null)" = fm/task-x1 ] \ + || fail "$label: refusal dropped the task branch" + [ "$(git -C "$case_dir/wt" rev-parse HEAD 2>/dev/null)" = "$head" ] \ + || fail "$label: refusal moved the task branch off the unlanded commit" + [ -e "$case_dir/state/task-x1.meta" ] \ + || fail "$label: refusal erased the durable task record" +} + append_pr_meta_for_current_head() { local case_dir=$1 head head=$(git -C "$case_dir/wt" rev-parse HEAD) @@ -545,12 +637,16 @@ run_teardown() { # no tasks-axi stub, so PATH resolves the installed one). seed_backlog_in_flight() { local case_dir=$1 kind=${2:-ship} + command -v tasks-axi >/dev/null 2>&1 \ + || fail "tasks-axi is required for teardown backlog fixtures" mkdir -p "$case_dir/data" printf '%s\n' '# Backlog' '' '## In flight' '' '## Queued' '' '## Done' \ > "$case_dir/data/backlog.md" tasks-axi add task-x1 "teardown fixture task" --kind "$kind" \ - --file "$case_dir/data/backlog.md" >/dev/null - tasks-axi start task-x1 --file "$case_dir/data/backlog.md" >/dev/null + --file "$case_dir/data/backlog.md" >/dev/null \ + || fail "could not add the teardown fixture backlog task" + tasks-axi start task-x1 --file "$case_dir/data/backlog.md" >/dev/null \ + || fail "could not start the teardown fixture backlog task" } backlog_row_state() { @@ -851,6 +947,98 @@ test_merged_pr_with_later_local_commit_refuses() { pass "merged PR does not allow teardown after a later local commit" } +test_squash_merged_rebased_branch_allows() { + local case_dir rc pr_head + case_dir=$(make_case squash-rebased) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" rebased) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 0 "$rc" "squash-rebased: teardown should succeed when the worktree followed the pipeline rebase"$'\n'"$(cat "$case_dir/stderr")" + ! grep -q REFUSED "$case_dir/stderr" || fail "squash-rebased: teardown printed a REFUSED line" + pass "squash-merged task whose local branch followed the pipeline rebase is torn down" +} + +test_squash_merged_same_file_different_content_refuses() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-same-path-diverged) + write_meta "$case_dir" no-mistakes ship + # The pipeline rebase produced a different blob for shared.txt than the stale + # local still holds, then squash-merged. Same path is not proof the local + # content landed. + pr_head=$(setup_squash_rebased_history "$case_dir" stale) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-same-path-diverged: teardown should refuse when the same file has different content"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-same-path-diverged: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-same-path-diverged "$local_head" + pass "squash-merged same-path different content still refuses" +} + +# The local branch followed the pipeline rebase, so without later.txt this is the +# q2 ALLOW case exactly. The one unlanded follow-up commit is the sole difference +# and must be the sole reason teardown refuses. +test_squash_merged_rebased_local_with_unlanded_commit_refuses() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-rebased-unlanded) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" rebased-plus-unlanded) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_pr_merged_for_head "$case_dir" "$pr_head" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-rebased-unlanded: teardown should refuse extra local commits that never landed"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-rebased-unlanded: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-rebased-unlanded "$local_head" + pass "squash-merged rebased local still refuses a genuinely unlanded follow-up commit" +} + +test_squash_merged_stale_local_refuses_when_forge_unreachable() { + local case_dir rc pr_head local_head + case_dir=$(make_case squash-stale-offline) + write_meta "$case_dir" no-mistakes ship + pr_head=$(setup_squash_rebased_history "$case_dir" stale) + local_head=$(git -C "$case_dir/wt" rev-parse HEAD) + printf '%s\n' \ + 'pr=https://github.com/example/repo/pull/7' \ + "pr_head=$pr_head" >> "$case_dir/state/task-x1.meta" + add_gh_axi_error "$case_dir" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "squash-stale-offline: teardown should refuse when the forge is down and trees conflict"$'\n'"$(cat "$case_dir/stderr")" + grep -q REFUSED "$case_dir/stderr" || fail "squash-stale-offline: no REFUSED line in stderr" + assert_refusal_retained_task_state "$case_dir" squash-stale-offline "$local_head" + pass "squash-merged stale local still refuses when the forge is unreachable" +} + test_pr_check_does_not_refresh_stale_pr_head() { local case_dir rc pr_head new_head count case_dir=$(make_case pr-check-stale) @@ -992,6 +1180,271 @@ test_gh_error_and_content_absent_refuses() { pass "gh lookup error with content not in default refuses (fail-safe)" } +# Write a meta that predates the spawn_gen field entirely. Args: case_dir mode kind +write_legacy_meta() { + local case_dir=$1 mode=$2 kind=$3 + fm_write_meta "$case_dir/state/task-x1.meta" \ + "window=firstmate:fm-task-x1" \ + "endpoint_task_id=task-x1" \ + "worktree=$case_dir/wt" \ + "project=$case_dir/project" \ + "kind=$kind" \ + "mode=$mode" \ + "harness=codex" +} + +# Count spawn_gen fields in the task's meta, so a refusal can prove it left the +# record byte-equivalent rather than stamped. +legacy_meta_gen_count() { + local case_dir=$1 + awk -F= '$1 == "spawn_gen" { count++ } END { print count + 0 }' \ + "$case_dir/state/task-x1.meta" 2>/dev/null || printf '0\n' +} + +# Override fakebin/tmux so the recovery-grade classifier reads the endpoint as +# unreadable (a session inventory failure it cannot attribute), never dead. +add_unreadable_tmux() { + local case_dir=$1 + cat > "$case_dir/fakebin/tmux" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + list-windows) echo "error connecting to fixture: permission denied" >&2 ; exit 1 ;; +esac +exit 0 +SH + chmod +x "$case_dir/fakebin/tmux" +} + +test_legacy_record_without_the_flag_refuses() { + local case_dir rc + case_dir=$(make_case legacy-noflag) + write_legacy_meta "$case_dir" no-mistakes ship + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "legacy-noflag: a record without spawn_gen must refuse without --legacy-record" + grep -q -- '--legacy-record' "$case_dir/stderr" \ + || fail "legacy-noflag: the refusal did not name the --legacy-record path" + [ "$(legacy_meta_gen_count "$case_dir")" = 0 ] \ + || fail "legacy-noflag: the refusal stamped a spawn generation into the record" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-noflag: the refusal closed the backlog item anyway" + assert_present "$case_dir/state/task-x1.meta" \ + "legacy-noflag: the refusal removed the task record" + pass "a record predating spawn_gen refuses teardown until --legacy-record is passed" +} + +test_legacy_record_teardown_completes_when_landed_and_endpoint_dead() { + local case_dir out + case_dir=$(make_case legacy-allow) + write_legacy_meta "$case_dir" no-mistakes ship + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + # The default fakebin tmux answers every query with success and no output, so + # the classifier reads the recorded window as authoritatively missing. + + out=$(run_teardown "$case_dir" --legacy-record) \ + || fail "legacy-allow: teardown refused a landed legacy record with a dead endpoint" + [ "$(backlog_row_state "$case_dir")" = "done" ] \ + || fail "legacy-allow: teardown returned success with its backlog item still open" + printf '%s\n' "$out" | grep -Fq 'legacy record accepted without spawn_gen: endpoint missing, incarnation legacy-' \ + || fail "legacy-allow: the teardown line did not log the accepted legacy incarnation: $out" + assert_absent "$case_dir/state/task-x1.backlog-close" \ + "legacy-allow: a landed legacy close left its pending-close record behind" + assert_absent "$case_dir/state/task-x1.meta" \ + "legacy-allow: teardown left the task record behind" + pass "a landed legacy record with a dead endpoint tears down and logs its accepted incarnation" +} + +test_legacy_record_teardown_refuses_unlanded_work() { + local case_dir rc before + case_dir=$(make_case legacy-unlanded) + write_legacy_meta "$case_dir" no-mistakes ship + seed_backlog_in_flight "$case_dir" + # Real content committed but pushed nowhere and merged nowhere. + wt_commit_file "$case_dir" feature.txt unique-legacy-content "real unlanded work" + before=$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}') + + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "legacy-unlanded: --legacy-record must not relax the unlanded-work refusal" + grep -q REFUSED "$case_dir/stderr" \ + || fail "legacy-unlanded: no REFUSED line for unlanded legacy work" + [ "$(legacy_meta_gen_count "$case_dir")" = 0 ] \ + || fail "legacy-unlanded: the unlanded refusal stamped a spawn generation into the record" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$before" ] \ + || fail "legacy-unlanded: the unlanded refusal modified the task record" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-unlanded: the unlanded refusal closed the backlog item anyway" + pass "--legacy-record never relaxes the unlanded-work refusal" +} + +test_legacy_record_teardown_refuses_an_ambiguous_endpoint() { + local case_dir rc before + case_dir=$(make_case legacy-ambiguous) + write_legacy_meta "$case_dir" no-mistakes ship + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + add_unreadable_tmux "$case_dir" + before=$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}') + + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "legacy-ambiguous: an unreadable endpoint must refuse the legacy acceptance" + grep -q "not confidently dead or agent-less" "$case_dir/stderr" \ + || fail "legacy-ambiguous: the refusal did not name the endpoint state" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$before" ] \ + || fail "legacy-ambiguous: the endpoint refusal modified the task record" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-ambiguous: the endpoint refusal closed the backlog item anyway" + pass "an endpoint that cannot be confidently read as dead refuses --legacy-record teardown" +} + +test_legacy_record_rolls_the_stamp_back_when_the_marker_write_fails() { + local case_dir rc before + case_dir=$(make_case legacy-stamp-rollback) + write_legacy_meta "$case_dir" no-mistakes ship + printf '%s\n' 'pr=not-a-valid-url' >> "$case_dir/state/task-x1.meta" + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + before=$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}') + + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" \ + "legacy-stamp-rollback: an unrecordable close must fail the teardown after accepting the legacy record" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$before" ] \ + || fail "legacy-stamp-rollback: the failed marker write left the record modified" + grep -q "rolled back" "$case_dir/stderr" \ + || fail "legacy-stamp-rollback: the refusal did not report the rolled-back stamp" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-stamp-rollback: the failed teardown closed the backlog item anyway" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout2" 2> "$case_dir/stderr2" + rc=$? + set -e + expect_code 1 "$rc" \ + "legacy-stamp-rollback: the flag-less retry must not sail past the endpoint gate on the rolled-back record" + grep -q -- '--legacy-record' "$case_dir/stderr2" \ + || fail "legacy-stamp-rollback: the retry refusal did not name the flag path" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$before" ] \ + || fail "legacy-stamp-rollback: the flag-less retry modified the record" + pass "--legacy-record teardown rolls its stamp back when the close marker write fails" +} + +# Override fakebin/perl so ONLY the stamp rollback's truncate fails; every other +# perl call in the lifecycle still runs the real interpreter, so the abandoned +# attempt leaves its stamp behind for exactly the reason under test. +add_failing_truncate_perl() { + local case_dir=$1 real + real=$(command -v perl) + cat > "$case_dir/fakebin/perl" <<SH +#!/usr/bin/env bash +case "\$*" in + *truncate*) exit 1 ;; +esac +exec "$real" "\$@" +SH + chmod +x "$case_dir/fakebin/perl" +} + +test_retained_legacy_stamp_still_faces_the_endpoint_gate() { + local case_dir rc stamped + case_dir=$(make_case legacy-stamp-retained) + write_legacy_meta "$case_dir" no-mistakes ship + printf '%s\n' 'pr=not-a-valid-url' >> "$case_dir/state/task-x1.meta" + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + add_failing_truncate_perl "$case_dir" + + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + expect_code 1 "$rc" \ + "legacy-stamp-retained: an unrecordable close must fail the teardown after accepting the legacy record" + grep -q "could not be rolled back" "$case_dir/stderr" \ + || fail "legacy-stamp-retained: the fixture did not exercise a failed rollback" + [ "$(legacy_meta_gen_count "$case_dir")" = 1 ] \ + || fail "legacy-stamp-retained: the abandoned attempt did not leave its stamp on the record" + stamped=$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}') + + # The stamp the failed rollback left behind is the whole risk: a retry must + # not read it as an incarnation some spawn published and sail past the + # dead-or-agent-less endpoint gate onto a reused endpoint. + add_unreadable_tmux "$case_dir" + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout2" 2> "$case_dir/stderr2" + rc=$? + set -e + expect_code 1 "$rc" \ + "legacy-stamp-retained: the retry must re-run the endpoint gate on the retained stamp" + grep -q "not confidently dead or agent-less" "$case_dir/stderr2" \ + || fail "legacy-stamp-retained: the retry skipped the dead-or-agent-less endpoint gate" + [ "$(legacy_meta_gen_count "$case_dir")" = 1 ] \ + || fail "legacy-stamp-retained: the retry stamped a second incarnation into the record" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$stamped" ] \ + || fail "legacy-stamp-retained: the endpoint refusal modified the task record" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-stamp-retained: the endpoint refusal closed the backlog item anyway" + + set +e + run_teardown "$case_dir" > "$case_dir/stdout3" 2> "$case_dir/stderr3" + rc=$? + set -e + expect_code 1 "$rc" \ + "legacy-stamp-retained: a flag-less retry must refuse the retained legacy stamp" + grep -q -- '--legacy-record' "$case_dir/stderr3" \ + || fail "legacy-stamp-retained: the flag-less refusal did not name the flag path" + [ "$(cksum "$case_dir/state/task-x1.meta" | awk '{print $1, $2}')" = "$stamped" ] \ + || fail "legacy-stamp-retained: the flag-less refusal modified the task record" + pass "a legacy stamp a failed rollback left behind still faces the endpoint gate" +} + +test_legacy_record_never_accepts_a_corrupt_spawn_gen() { + local case_dir rc + case_dir=$(make_case legacy-corrupt) + write_legacy_meta "$case_dir" no-mistakes ship + printf 'spawn_gen=one\nspawn_gen=two\n' >> "$case_dir/state/task-x1.meta" + seed_backlog_in_flight "$case_dir" + wt_commit "$case_dir" "landed legacy work" + add_fork_with_pushed_branch "$case_dir" + + set +e + run_teardown "$case_dir" --legacy-record > "$case_dir/stdout" 2> "$case_dir/stderr" + rc=$? + set -e + + expect_code 1 "$rc" "legacy-corrupt: an ambiguous spawn_gen must refuse even with --legacy-record" + grep -q "unreadable spawn_gen" "$case_dir/stderr" \ + || fail "legacy-corrupt: the refusal did not name the unreadable spawn_gen" + [ "$(legacy_meta_gen_count "$case_dir")" = 2 ] \ + || fail "legacy-corrupt: the refusal rewrote the corrupt record" + [ "$(backlog_row_state "$case_dir")" = in_flight ] \ + || fail "legacy-corrupt: the refusal closed the backlog item anyway" + pass "a corrupt spawn_gen is never accepted as a legacy record" +} + test_stale_index_lock_cleared_and_teardown_succeeds() { local case_dir rc lock case_dir=$(make_case stale-index-lock) @@ -1152,6 +1605,13 @@ test_non_linked_index_lock_path_is_checked_from_worktree() { test_index_lock_mtime_read_failure_refuses() { local case_dir rc lock + # The mtime fault is injected by a fake stat on PATH; on Darwin the lock + # helper now calls /usr/bin/stat directly, so the fake can never fire there. + # Skip the Darwin run of this case. + if [ "$(uname)" = Darwin ]; then + pass "index-lock mtime fault injection is PATH-based; skipped on Darwin where stat is /usr/bin/stat" + return + fi case_dir=$(make_case mtime-error-index-lock) write_meta "$case_dir" no-mistakes ship wt_commit "$case_dir" "shippable work" @@ -3448,79 +3908,91 @@ EOF pass "the run abort and the leaked-process reap both complete before the destructive worktree return" } -test_local_only_fork_remote_allows -test_teardown_closes_the_backlog_item_itself -test_teardown_manual_backend_leaves_the_backlog_to_the_operator -test_local_only_truly_unpushed_refuses -test_local_only_merged_to_local_main_allows -test_no_mistakes_origin_remote_allows -test_no_mistakes_truly_unpushed_refuses -test_local_only_force_overrides_unpushed -test_secondmate_pr_registration_publishes_ready_line -test_secondmate_home_teardown_delivers_final_line_or_refuses -test_teardown_missing_busy_sidecar_completes -test_herdr_teardown_clears_escalation_marker -test_herdr_flat_teardown_refuses_orphaning_records_then_retry_completes -test_herdr_flat_teardown_refuses_records_on_unparseable_presence -test_herdr_flat_teardown_preflight_refuses_before_changes -test_forced_secondmate_herdr_child_preflight_refuses_before_changes -test_forced_secondmate_teardown_holds_descendant_lifecycle_locks -test_forced_secondmate_herdr_child_retains_records_when_close_unconfirmed -test_forced_teardown_retains_nested_secondmate_home_when_grandchild_close_unconfirmed -test_herdr_projection_teardown_restores_safe_focus -test_herdr_projection_teardown_defers_idempotently_without_a_captain_decision -test_herdr_projection_teardown_recovers_after_confirmed_close -test_herdr_projection_teardown_retires_journal_only_after_confirmed_close -test_herdr_projection_teardown_retains_journal_when_close_unconfirmed -test_herdr_projection_teardown_surfaces_restore_failure_without_blocking_cleanup -test_squash_merged_branch_deleted_allows -test_squash_merged_pr_allows_when_head_ancestor_of_pr_head -test_no_pr_recorded_discovers_merged_pr_by_branch_allows -test_squash_merged_pr_allows_replayed_unpushed_patch -test_merged_pr_with_later_local_commit_refuses -test_pr_check_does_not_refresh_stale_pr_head -test_pr_check_records_remote_head_when_local_lags -test_content_in_default_fallback_allows -test_content_fallback_refreshes_stale_origin_ref -test_dirty_worktree_refuses -test_gh_error_and_content_absent_refuses -test_stale_index_lock_cleared_and_teardown_succeeds -test_live_index_lock_is_never_removed_and_teardown_refuses -test_lsof_error_never_clears_index_lock -test_stale_index_lock_cleanup_rechecks_dirty_worktree -test_non_linked_index_lock_path_is_checked_from_worktree -test_index_lock_mtime_read_failure_refuses -test_transient_index_lock_clears_after_first_attempt_and_retry_succeeds -test_persistent_index_lock_exhausts_retries_and_refuses_loudly -test_empty_retry_wait_uses_default_without_aborting -test_fractional_legacy_retry_wait_refuses_without_arithmetic_error -test_parked_own_run_is_aborted_before_teardown -test_parked_run_advanced_past_unfetched_head_is_still_aborted -test_parked_run_with_mismatched_ledger_head_is_never_aborted -test_parked_run_with_malformed_ledger_row_is_never_aborted -test_parked_run_with_impossible_ledger_date_is_never_aborted -test_terminal_status_with_gate_never_queries_or_aborts_ledger_fallback -test_parked_run_advanced_head_locally_fetched_is_still_aborted -test_parked_advanced_run_without_anchor_is_never_aborted -test_parked_advanced_run_ancestor_anchor_is_never_aborted -test_parked_terminal_unfetched_row_is_never_aborted -test_parked_run_terminal_newest_row_at_own_head_is_never_aborted -test_parked_run_behind_diverged_newer_row_is_never_aborted -test_parked_advanced_run_ambiguous_rows_are_never_aborted -test_ledger_proven_continuation_never_aborts_active_run -test_parked_own_run_refuses_when_abort_is_unconfirmed -test_mismatched_run_after_abort_refuses_unconfirmed -test_empty_status_after_abort_refuses_unconfirmed -test_not_found_status_after_abort_confirms_completion -test_another_branchs_parked_run_is_never_touched -test_own_autonomous_run_is_left_alone -test_leaked_worktree_process_is_reaped -test_leaked_tasktmp_process_is_reaped -test_lsof_absent_reaps_tmux_process_group -test_lsof_error_refuses_before_removal -test_reused_pid_identity_is_not_force_killed -test_exec_changed_process_is_still_reaped -test_process_spawned_during_grace_is_reaped_on_later_pass -test_persistent_scan_refuses_after_bounded_retries -test_process_exit_during_identity_lookup_does_not_refuse -test_run_abort_precedes_process_reap_precedes_worktree_removal +fm_test_run_cases \ + test_local_only_fork_remote_allows \ + test_teardown_closes_the_backlog_item_itself \ + test_teardown_manual_backend_leaves_the_backlog_to_the_operator \ + test_local_only_truly_unpushed_refuses \ + test_local_only_merged_to_local_main_allows \ + test_no_mistakes_origin_remote_allows \ + test_no_mistakes_truly_unpushed_refuses \ + test_local_only_force_overrides_unpushed \ + test_secondmate_pr_registration_publishes_ready_line \ + test_secondmate_home_teardown_delivers_final_line_or_refuses \ + test_teardown_missing_busy_sidecar_completes \ + test_herdr_teardown_clears_escalation_marker \ + test_herdr_flat_teardown_refuses_orphaning_records_then_retry_completes \ + test_herdr_flat_teardown_refuses_records_on_unparseable_presence \ + test_herdr_flat_teardown_preflight_refuses_before_changes \ + test_forced_secondmate_herdr_child_preflight_refuses_before_changes \ + test_forced_secondmate_teardown_holds_descendant_lifecycle_locks \ + test_forced_secondmate_herdr_child_retains_records_when_close_unconfirmed \ + test_forced_teardown_retains_nested_secondmate_home_when_grandchild_close_unconfirmed \ + test_herdr_projection_teardown_restores_safe_focus \ + test_herdr_projection_teardown_defers_idempotently_without_a_captain_decision \ + test_herdr_projection_teardown_recovers_after_confirmed_close \ + test_herdr_projection_teardown_retires_journal_only_after_confirmed_close \ + test_herdr_projection_teardown_retains_journal_when_close_unconfirmed \ + test_herdr_projection_teardown_surfaces_restore_failure_without_blocking_cleanup \ + test_squash_merged_branch_deleted_allows \ + test_squash_merged_pr_allows_when_head_ancestor_of_pr_head \ + test_no_pr_recorded_discovers_merged_pr_by_branch_allows \ + test_squash_merged_pr_allows_replayed_unpushed_patch \ + test_merged_pr_with_later_local_commit_refuses \ + test_squash_merged_rebased_branch_allows \ + test_squash_merged_same_file_different_content_refuses \ + test_squash_merged_rebased_local_with_unlanded_commit_refuses \ + test_squash_merged_stale_local_refuses_when_forge_unreachable \ + test_pr_check_does_not_refresh_stale_pr_head \ + test_pr_check_records_remote_head_when_local_lags \ + test_content_in_default_fallback_allows \ + test_content_fallback_refreshes_stale_origin_ref \ + test_dirty_worktree_refuses \ + test_gh_error_and_content_absent_refuses \ + test_legacy_record_without_the_flag_refuses \ + test_legacy_record_teardown_completes_when_landed_and_endpoint_dead \ + test_legacy_record_teardown_refuses_unlanded_work \ + test_legacy_record_teardown_refuses_an_ambiguous_endpoint \ + test_legacy_record_rolls_the_stamp_back_when_the_marker_write_fails \ + test_retained_legacy_stamp_still_faces_the_endpoint_gate \ + test_legacy_record_never_accepts_a_corrupt_spawn_gen \ + test_stale_index_lock_cleared_and_teardown_succeeds \ + test_live_index_lock_is_never_removed_and_teardown_refuses \ + test_lsof_error_never_clears_index_lock \ + test_stale_index_lock_cleanup_rechecks_dirty_worktree \ + test_non_linked_index_lock_path_is_checked_from_worktree \ + test_index_lock_mtime_read_failure_refuses \ + test_transient_index_lock_clears_after_first_attempt_and_retry_succeeds \ + test_persistent_index_lock_exhausts_retries_and_refuses_loudly \ + test_empty_retry_wait_uses_default_without_aborting \ + test_fractional_legacy_retry_wait_refuses_without_arithmetic_error \ + test_parked_own_run_is_aborted_before_teardown \ + test_parked_run_advanced_past_unfetched_head_is_still_aborted \ + test_parked_run_with_mismatched_ledger_head_is_never_aborted \ + test_parked_run_with_malformed_ledger_row_is_never_aborted \ + test_parked_run_with_impossible_ledger_date_is_never_aborted \ + test_terminal_status_with_gate_never_queries_or_aborts_ledger_fallback \ + test_parked_run_advanced_head_locally_fetched_is_still_aborted \ + test_parked_advanced_run_without_anchor_is_never_aborted \ + test_parked_advanced_run_ancestor_anchor_is_never_aborted \ + test_parked_terminal_unfetched_row_is_never_aborted \ + test_parked_run_terminal_newest_row_at_own_head_is_never_aborted \ + test_parked_run_behind_diverged_newer_row_is_never_aborted \ + test_parked_advanced_run_ambiguous_rows_are_never_aborted \ + test_ledger_proven_continuation_never_aborts_active_run \ + test_parked_own_run_refuses_when_abort_is_unconfirmed \ + test_mismatched_run_after_abort_refuses_unconfirmed \ + test_empty_status_after_abort_refuses_unconfirmed \ + test_not_found_status_after_abort_confirms_completion \ + test_another_branchs_parked_run_is_never_touched \ + test_own_autonomous_run_is_left_alone \ + test_leaked_worktree_process_is_reaped \ + test_leaked_tasktmp_process_is_reaped \ + test_lsof_absent_reaps_tmux_process_group \ + test_lsof_error_refuses_before_removal \ + test_reused_pid_identity_is_not_force_killed \ + test_exec_changed_process_is_still_reaped \ + test_process_spawned_during_grace_is_reaped_on_later_pass \ + test_persistent_scan_refuses_after_bounded_retries \ + test_process_exit_during_identity_lookup_does_not_refuse \ + test_run_abort_precedes_process_reap_precedes_worktree_removal diff --git a/tests/fm-test-run.test.sh b/tests/fm-test-run.test.sh index 481204c5ceb..0a6df884d4f 100755 --- a/tests/fm-test-run.test.sh +++ b/tests/fm-test-run.test.sh @@ -115,6 +115,7 @@ init_changed_fixture_repo() { fm-procevent-quota.test.sh \ fm-quota-choose.test.sh \ fm-pi-watch-extension.test.sh \ + fm-pi-windows-shell-invocation.test.sh \ fm-afk-return.test.sh \ fm-bearings-snapshot.test.sh \ fm-backend-cmux.test.sh \ @@ -168,6 +169,8 @@ init_changed_fixture_repo() { : >"$repo/.opencode/plugins/fm-primary-cd-check.js" : >"$repo/.pi/extensions/fm-primary-pi-watch.ts" : >"$repo/.pi/extensions/fm-primary-turnend-guard.ts" + mkdir -p "$repo/.pi/extensions/lib" + : >"$repo/.pi/extensions/lib/fm-operational-input.ts" : >"$repo/docs/fm-test-isolation-proof.md" : >"$repo/skills/example/SKILL.md" : >"$repo/CONTRIBUTING.md" @@ -178,6 +181,69 @@ init_changed_fixture_repo() { git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm baseline } +# Build a repository with a primary checkout and one linked worktree, each +# holding a runnable copy of the runner and a probe suite that records the fact +# that it ran. Untracked copies are enough: the runner resolves its root from +# its own path, and the probe is named explicitly. +init_primary_and_linked_worktree() { + local repo=$1 linked=$2 tree + fm_git_init_commit "$repo" + git -C "$repo" worktree add --quiet -b linked-probe "$linked" + for tree in "$repo" "$linked"; do + mkdir -p "$tree/bin" "$tree/tests" + cp "$RUNNER" "$tree/bin/fm-test-run.sh" + chmod +x "$tree/bin/fm-test-run.sh" + cat >"$tree/tests/probe.test.sh" <<PROBE +#!/usr/bin/env bash +echo "ok - probe suite" +: >"$tree/ran" +PROBE + chmod +x "$tree/tests/probe.test.sh" + done +} + +# A task worker's isolated placement is checked once, when the task starts. +# Nothing re-checks it, so a worker that later changes directory into the +# repository's primary checkout would run this branch-switching suite in the one +# checkout every linked worktree resolves against. The runner refuses that. +test_task_marker_refuses_the_primary_checkout() { + local tmp repo linked out rc + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-primary.XXXXXX") + repo="$tmp/repo" + linked="$tmp/linked" + init_primary_and_linked_worktree "$repo" "$linked" + + # Marker set, primary checkout: refuse, name the primary, and run nothing. + out=$(FM_TASK_ID=probe-task "$repo/bin/fm-test-run.sh" tests/probe.test.sh 2>&1) && rc=0 || rc=$? + [ "$rc" -ne 0 ] || { rm -rf "$tmp"; fail "the runner must refuse the primary checkout under a task marker"; } + assert_contains "$out" "primary checkout" "refusal did not name the primary checkout" + assert_contains "$out" "FM_TASK_ID=probe-task" "refusal did not name the task marker" + assert_contains "$out" "task worktree" "refusal did not point at the task worktree" + assert_not_contains "$out" "FM_TEST_BEGIN" "the refusal must happen before any suite runs" + assert_absent "$repo/ran" "the refused run still executed a suite" + + # Marker set, linked worktree: the assigned placement, so the suite runs. + FM_TASK_ID=probe-task "$linked/bin/fm-test-run.sh" tests/probe.test.sh >/dev/null 2>&1 \ + || { rm -rf "$tmp"; fail "the runner must still run in a linked task worktree"; } + assert_present "$linked/ran" "the linked-worktree run did not execute its suite" + + # No marker: a person in their own checkout is unaffected. + "$repo/bin/fm-test-run.sh" tests/probe.test.sh >/dev/null 2>&1 \ + || { rm -rf "$tmp"; fail "an unmarked run in the primary checkout must be unchanged"; } + assert_present "$repo/ran" "the unmarked run did not execute its suite" + + # Inspection executes nothing, so it stays available even in the primary. + rm -f "$repo/ran" + out=$(FM_TASK_ID=probe-task "$repo/bin/fm-test-run.sh" --list tests/probe.test.sh 2>&1) \ + || { rm -rf "$tmp"; fail "--list must remain available under a task marker"; } + [ "$out" = "tests/probe.test.sh" ] \ + || { rm -rf "$tmp"; fail "--list under a task marker printed: $out"; } + assert_absent "$repo/ran" "--list must not execute a suite" + + rm -rf "$tmp" + pass "a task marker refuses execution in the primary checkout and leaves worktrees and inspection alone" +} + test_changed_reference_scan_batches_test_files() { local tmp repo fake_bin grep_calls_file basename_calls_file local real_grep real_basename listed grep_calls basename_calls i source @@ -296,6 +362,21 @@ test_changed_runner_surfaces_select_their_family() { pass "runner and its documentation surfaces select their curated family, not just their contract owners" } +test_shell_line_ending_policy_selects_runner_contract() { + local tmp repo listed + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-attributes.XXXXXX") + repo="$tmp/repo" + init_changed_fixture_repo "$repo" + printf '*.sh text eol=lf\n' >"$repo/.gitattributes" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-test-run.test.sh" \ + "shell line-ending policy selects the runner contract" + assert_contains "$listed" "tests/fm-gitignore-config.test.sh" \ + "shell line-ending policy also selects repository Git configuration coverage" + rm -rf "$tmp" + pass "shell line-ending policy selects runner coverage" +} + test_changed_dependency_selection_and_unmapped_failure() { local tmp repo listed rc tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-changed.XXXXXX") @@ -339,9 +420,18 @@ test_changed_dependency_selection_and_unmapped_failure() { assert_contains "$listed" "tests/fm-ask-user-authority.test.sh" "skill source selects pure contract coverage" assert_contains "$listed" "tests/fm-cd-pretool-check.test.sh" "Claude and Pi source selects hook coverage" assert_contains "$listed" "tests/fm-pi-watch-extension.test.sh" "Pi source selects watcher coverage" + assert_contains "$listed" "tests/fm-pi-windows-shell-invocation.test.sh" \ + "turn-end extension selects native-Windows shell coverage" git -C "$repo" add .agents .claude .pi git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm non-bin-source-change + printf '\n' >>"$repo/.pi/extensions/lib/fm-operational-input.ts" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + assert_contains "$listed" "tests/fm-pi-windows-shell-invocation.test.sh" \ + "operational-input extension selects native-Windows shell coverage" + git -C "$repo" add .pi/extensions/lib/fm-operational-input.ts + git -C "$repo" -c user.name=test -c user.email=test@example.invalid commit -qm operational-input-source-change + printf '\n' >>"$repo/skills/example/SKILL.md" listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) assert_contains "$listed" "tests/fm-brief.test.sh" \ @@ -443,8 +533,12 @@ test_changed_dependency_selection_and_unmapped_failure() { [ "$rc" -eq 2 ] || fail "unmapped changed source must fail with exit 2, got $rc" grep -Fq 'no changed-test mapping for source path: src/unmapped.ts' "$tmp/err" \ || fail "unmapped changed source failure is not actionable: $(cat "$tmp/err")" + + rm -f "$repo/src/unmapped.ts" + listed=$(cd "$repo" && bin/fm-test-run.sh --list --changed --base HEAD) + [ -z "$listed" ] || fail "a retired unmapped source without consumers selected tests: $listed" rm -rf "$tmp" - pass "changed selection covers dependents and fails closed for unmapped source" + pass "changed selection covers dependents, fails closed for live unmapped source, and accepts retired unconsumed source" } # A direct test reference is per-script evidence. Widening it to the referencing @@ -616,6 +710,77 @@ SH pass "changed defaults to bounded automatic scheduling with serial override" } +test_windows_posix_mode_emulation_does_not_fail_parallel_runs() { + local tmp repo fakebin real_stat out rc acl_log + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-windows-modes.XXXXXX") + repo="$tmp/repo" + fakebin="$tmp/fakebin" + acl_log="$tmp/acl.log" + real_stat=$(command -v stat) + init_changed_fixture_repo "$repo" + mkdir -p "$fakebin" + cat >"$fakebin/uname" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${FAKE_UNAME:-MINGW64_NT-10.0}" +SH + cat >"$fakebin/stat" <<'SH' +#!/usr/bin/env bash +if [ "${1:-}" = -c ] && [ "${2:-}" = %a ]; then + printf '%s\n' 755 + exit 0 +fi +exec "$REAL_STAT" "$@" +SH + cat >"$fakebin/cygpath" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "$2" +SH + cat >"$fakebin/powershell.exe" <<'SH' +#!/usr/bin/env bash +printf '%s\n' "${FM_TEST_WORKER_NATIVE:?}" >>"${FM_TEST_ACL_LOG:?}" +exit "${FM_TEST_ACL_RC:-0}" +SH + chmod +x "$fakebin/uname" "$fakebin/stat" "$fakebin/cygpath" "$fakebin/powershell.exe" + set +e + out=$(cd "$repo" && PATH="$fakebin:$PATH" REAL_STAT="$real_stat" FM_TEST_ACL_LOG="$acl_log" \ + bin/fm-test-run.sh --jobs 2 \ + tests/fm-cd-pretool-check.test.sh tests/fm-ask-user-authority.test.sh 2>&1) + rc=$? + set -e + expect_code 0 "$rc" "native-Windows POSIX-mode emulation" + assert_contains "$out" "FM_TEST_SUMMARY total=2 failed=0" \ + "Windows mode emulation did not complete both parallel scripts" + [ "$(wc -l <"$acl_log" | tr -d ' ')" -eq 2 ] \ + || fail "Windows mode emulation must prove each worker directory's ACL" + + : >"$acl_log" + set +e + out=$(cd "$repo" && PATH="$fakebin:$PATH" REAL_STAT="$real_stat" \ + FM_TEST_ACL_LOG="$acl_log" FM_TEST_ACL_RC=1 \ + bin/fm-test-run.sh --jobs 2 \ + tests/fm-cd-pretool-check.test.sh tests/fm-ask-user-authority.test.sh 2>&1) + rc=$? + set -e + expect_code 1 "$rc" "native-Windows ACL enforcement" + assert_contains "$out" "isolation failure: worker root is not private" \ + "Windows mode emulation must reject an unsafe ACL" + + : >"$acl_log" + set +e + out=$(cd "$repo" && PATH="$fakebin:$PATH" REAL_STAT="$real_stat" \ + FM_TEST_ACL_LOG="$acl_log" FAKE_UNAME=CYGWIN_NT-10.0 \ + bin/fm-test-run.sh --jobs 2 \ + tests/fm-cd-pretool-check.test.sh tests/fm-ask-user-authority.test.sh 2>&1) + rc=$? + set -e + expect_code 1 "$rc" "Cygwin POSIX-mode enforcement" + assert_contains "$out" "isolation failure: worker root is not private (reported mode 755;" \ + "Cygwin mode enforcement did not reject a non-0700 worker root" + [ ! -s "$acl_log" ] || fail "Cygwin must enforce POSIX modes rather than use the Windows exemption" + rm -rf "$tmp" + pass "Windows emulation proves ACL privacy while Cygwin enforces POSIX modes" +} + # A local verification round names the subjects it cares about. Exercise begin/end # markers from real fixture processes to prove that a plain list of script paths # gets bounded automatic scheduling without changing its per-script timeout @@ -922,6 +1087,80 @@ assert doc["summary"]["failed"] == 0 pass "gate-skip accounting is honest and non-failing" } +test_gate_skip_reason_is_recorded() { + local tmp skip_f out json + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-skipreason.XXXXXX") + skip_f="$tmp/skip.test.sh" + out="$tmp/out.txt" + json="$tmp/timing.json" + cat >"$skip_f" <<'SH' +#!/usr/bin/env bash +echo "skip: live: fmnosuchharness absent" +exit 0 +SH + chmod +x "$skip_f" + "$RUNNER" --json "$json" "$skip_f" >"$out" 2>"$tmp/err.txt" \ + || fail "a capability skip must still exit 0 from the runner" + grep -q 'live: fmnosuchharness absent' "$tmp/err.txt" \ + || fail "the runner log must name what this host could not exercise: $(cat "$tmp/err.txt")" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["gate_skip"] is True, record +assert record["gate_skip_reason"] == "live: fmnosuchharness absent", record +' "$json" || { rm -rf "$tmp"; fail "the timing artifact must carry the skip reason"; } + rm -rf "$tmp" + pass "a gate skip records why it skipped" +} + +test_a_run_that_ran_records_no_skip_reason() { + local tmp ran_f json + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-ranreason.XXXXXX") + ran_f="$tmp/ran.test.sh" + json="$tmp/timing.json" + cat >"$ran_f" <<'SH' +#!/usr/bin/env bash +echo "ok - ran" +exit 0 +SH + chmod +x "$ran_f" + "$RUNNER" --json "$json" "$ran_f" >"$tmp/out.txt" 2>&1 \ + || fail "a passing fixture must exit 0 from the runner" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["gate_skip"] is False, record +assert record["gate_skip_reason"] == "", record +' "$json" || { rm -rf "$tmp"; fail "a script that ran must carry an empty skip reason"; } + rm -rf "$tmp" + pass "a script that actually ran records no skip reason" +} + +test_live_guards_expect_a_capability_skip_class() { + local tmp out + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-liveclass.XXXXXX") + out="$tmp/out.txt" + # FM_LIVE=0 makes every live guard refuse without touching a harness, so this + # exercises the real family through the real runner in bounded time. + FM_LIVE=0 "$RUNNER" --json "$tmp/timing.json" \ + tests/fm-composer-matrix-live-e2e.test.sh >"$out" 2>"$tmp/err.txt" \ + || fail "a disabled live guard must not fail the runner: $(cat "$tmp/err.txt")" + grep -q 'expected_gate_skip=live-capability' "$out" \ + || fail "the live-harness family must expect a capability skip: $(grep FM_TEST_BEGIN "$out")" + python3 -c ' +import json, sys +doc = json.load(open(sys.argv[1])) +record = doc["scripts"][0] +assert record["expected_gate_skip"] == "live-capability", record +assert record["gate_skip"] is True, record +assert record["gate_skip_reason"].startswith("live: "), record +' "$tmp/timing.json" || { rm -rf "$tmp"; fail "the live guard record is wrong"; } + rm -rf "$tmp" + pass "live guards are recorded as a capability class, not a bare env opt-in" +} + test_fail_on_gate_skip_token() { local tmp skip_f out rc tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-test-run-fail-skip.XXXXXX") @@ -1600,33 +1839,40 @@ assert len(doc["scripts"])==3 pass "aggregate-json merges lane timing artifacts" } -test_list_all_exact_suite_coverage -test_family_selection -test_single_script_selection -test_changed_file_selection_is_conservative -test_changed_reference_scan_batches_test_files -test_changed_runner_surfaces_select_their_family -test_changed_dependency_selection_and_unmapped_failure -test_changed_bin_reference_selects_per_script_not_per_family -test_changed_uses_bounded_automatic_concurrency -test_script_list_uses_bounded_automatic_concurrency -test_family_proofs_run_in_separate_concurrent_phases -test_empty_selection_emits_summary -test_timing_markers_and_json -test_aggregate_exit_behavior -test_gate_skip_accounting -test_fail_on_gate_skip_token -test_exclude_family -test_portable_shard_union_and_coverage_guard -test_portable_serial_shards_partition_the_serial_lane -test_portable_serial_hint_coverage_is_reported_and_bounded -test_portable_serial_shard_lane_refusals -test_jobs_requires_proven_isolated -test_jobs_admits_a_concurrent_safe_family -test_unmapped_new_test_never_inherits_family_concurrency -test_concurrent_runs_are_ordered_longest_first -test_per_script_timeout_bounds_a_hang -test_max_wall_ms_is_a_result_not_advice -test_jobs_parallel_scheduler_and_failure_propagation -test_herdr_ci_family_run_has_a_step_timeout -test_aggregate_json +fm_test_run_cases \ + test_list_all_exact_suite_coverage \ + test_family_selection \ + test_single_script_selection \ + test_changed_file_selection_is_conservative \ + test_task_marker_refuses_the_primary_checkout \ + test_changed_reference_scan_batches_test_files \ + test_changed_runner_surfaces_select_their_family \ + test_shell_line_ending_policy_selects_runner_contract \ + test_changed_dependency_selection_and_unmapped_failure \ + test_changed_bin_reference_selects_per_script_not_per_family \ + test_changed_uses_bounded_automatic_concurrency \ + test_windows_posix_mode_emulation_does_not_fail_parallel_runs \ + test_script_list_uses_bounded_automatic_concurrency \ + test_family_proofs_run_in_separate_concurrent_phases \ + test_empty_selection_emits_summary \ + test_timing_markers_and_json \ + test_aggregate_exit_behavior \ + test_gate_skip_accounting \ + test_gate_skip_reason_is_recorded \ + test_a_run_that_ran_records_no_skip_reason \ + test_live_guards_expect_a_capability_skip_class \ + test_fail_on_gate_skip_token \ + test_exclude_family \ + test_portable_shard_union_and_coverage_guard \ + test_portable_serial_shards_partition_the_serial_lane \ + test_portable_serial_hint_coverage_is_reported_and_bounded \ + test_portable_serial_shard_lane_refusals \ + test_jobs_requires_proven_isolated \ + test_jobs_admits_a_concurrent_safe_family \ + test_unmapped_new_test_never_inherits_family_concurrency \ + test_concurrent_runs_are_ordered_longest_first \ + test_per_script_timeout_bounds_a_hang \ + test_max_wall_ms_is_a_result_not_advice \ + test_jobs_parallel_scheduler_and_failure_propagation \ + test_herdr_ci_family_run_has_a_step_timeout \ + test_aggregate_json diff --git a/tests/fm-tmux-agent-liveness.test.sh b/tests/fm-tmux-agent-liveness.test.sh index 5c2824a44db..321f51c68db 100755 --- a/tests/fm-tmux-agent-liveness.test.sh +++ b/tests/fm-tmux-agent-liveness.test.sh @@ -54,6 +54,11 @@ export PATH ln -s "$SLEEP_BIN" "$LAB/bin/claude-link" ln -s "$SLEEP_BIN" "$LAB/bin/pi" ln -s "$SLEEP_BIN" "$LAB/bin/notaharness" +# omp (Oh My Pi) is a single binary whose live process name is the bare word +# `omp`; the two decoys are the substrings an unanchored glob would misread. +ln -s "$SLEEP_BIN" "$LAB/bin/omp" +ln -s "$SLEEP_BIN" "$LAB/bin/ompd" +ln -s "$SLEEP_BIN" "$LAB/bin/comp" # muse's installed binary is muse-bin-<version>: the launcher execs it, so the # version is the LIVE process name and it changes on every auto-update. Unlike # Claude Code's version-named binary there is no `muse` path component to fall @@ -172,6 +177,23 @@ for decoy in musescore amuse muse-binary muse-bind; do done pass "tmux liveness: unrelated muse-containing command names stay ambiguous" +# --- omp's bare binary name ------------------------------------------------- +# omp (Oh My Pi) runs as a single binary whose live process name is exactly +# `omp`, with no path component to fall back on, so the anchored name is the +# only signal and the two decoys prove it never widens into a substring match. + +new_window omp "$LAB/bin/omp" 900 +wait_for_state "$SESSION:omp" alive \ + || fail "omp's bare binary name must classify alive" +pass "tmux liveness: omp's bare binary name classifies alive" + +for decoy in ompd comp; do + new_window "decoy-$decoy" "$LAB/bin/$decoy" 900 + wait_for_state "$SESSION:decoy-$decoy" ambiguous \ + || fail "'$decoy' merely contains 'omp' and must not classify as a live agent pane" +done +pass "tmux liveness: unrelated omp-containing command names stay ambiguous" + # --- a version name blinds one source --------------------------------------- # Giving a genuine harness-named executable the version-string argv[0] that # Claude Code 2.1.220 reports drives the two sources apart on both supported diff --git a/tests/fm-tmux-submit-busy.test.sh b/tests/fm-tmux-submit-busy.test.sh index e5c9329a91a..5290aaaea97 100755 --- a/tests/fm-tmux-submit-busy.test.sh +++ b/tests/fm-tmux-submit-busy.test.sh @@ -322,6 +322,21 @@ test_claude_busy_signature_uses_real_capture_shapes() { printf 'Working...\n' > "$composer" pane_busy pi pi || fail "Pi Working footer should be busy" pane_busy pi-signed pi-signed || fail "pi-signed should share Pi's exact Working footer" + # omp (Oh My Pi) renders its TUI line with U+2026; Pi's three-dot footer is + # not omp's signature, and neither Pi nor Codex may borrow the ellipsis form. + # The status-row spinner cell is its second, independent signal, and an idle + # status row (identity glyph, no elapsed time) is not busy. + pane_busy omp-three-dots omp && fail "omp must not read Pi's three-dot Working... footer as busy" + printf ' \xf3\xb1\x8a\xb7 Working\xe2\x80\xa6\n' > "$composer" + pane_busy omp omp || fail "omp TUI Working… footer should be busy" + pane_busy omp-ellipsis-pi pi && fail "Pi must not borrow omp's Working… footer" + pane_busy omp-ellipsis-codex codex && fail "Codex must not borrow omp's Working… footer" + printf ' \xe2\xa0\xa7 11s \xc2\xb7 gpt-6-astra\n' > "$composer" + pane_busy omp-spinner omp || fail "omp braille spinner plus elapsed cell should be busy" + printf ' \xf3\xb0\xb5\x97 \xc2\xb7 gpt-6-astra \xc2\xb7 36.7%%/41K\n' > "$composer" + pane_busy omp-idle omp && fail "omp idle status row must not read busy" + printf 'esc interrupt\n' > "$composer" + pane_busy omp-cross omp && fail "omp must ignore OpenCode's interrupt footer" printf 'Ctrl+c:cancel\n' > "$composer" pane_busy grok grok || fail "Grok cancel footer should be busy" pass "fm_pane_is_busy: Claude spinner is scoped, multi-frame, and backward-compatible" diff --git a/tests/fm-tool-update-check.test.sh b/tests/fm-tool-update-check.test.sh index bf785595256..59b8cff2261 100755 --- a/tests/fm-tool-update-check.test.sh +++ b/tests/fm-tool-update-check.test.sh @@ -688,15 +688,15 @@ test_findings_are_reported_once_until_they_change() { } test_an_overlong_report_says_it_was_cut() { - local home out report i tools= + local home out report i tools_json= # Many watched tools can outgrow one line. The report must say it was cut # rather than end mid-finding as if that were everything found. home=$(make_home long) for i in $(seq 1 30); do - [ -z "$tools" ] || tools="$tools," - tools="$tools{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" + [ -z "$tools_json" ] || tools_json="$tools_json," + tools_json="$tools_json{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" done - write_config "$home" "{\"tools\":[$tools]}" + write_config "$home" "{\"tools\":[$tools_json]}" out="$home/out.txt" run_check "$home" "$PATH" "$out" report=$(cat "$out") @@ -706,7 +706,7 @@ test_an_overlong_report_says_it_was_cut() { } test_a_finding_past_the_cut_is_still_reported() { - local home stale fresh out report i tools= + local home stale fresh out report i tools_json= # Once a report is long enough to be cut, a new finding lands past the cut and # leaves the printed line unchanged. It still has to count as news, or the PATH # skew this check exists for would be suppressed for good on a busy home. @@ -716,17 +716,17 @@ test_a_finding_past_the_cut_is_still_reported() { make_copy "$stale" "$TOOL" 'herdr 0.8.0' make_copy "$fresh" "$TOOL" 'herdr 0.8.2' for i in $(seq 1 30); do - [ -z "$tools" ] || tools="$tools," - tools="$tools{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" + [ -z "$tools_json" ] || tools_json="$tools_json," + tools_json="$tools_json{\"name\":\"absent-tool-$i\",\"command\":\"fm-absent-fixture-$i\"}" done out="$home/out.txt" - write_config "$home" "{\"tools\":[$tools]}" + write_config "$home" "{\"tools\":[$tools_json]}" run_check "$home" "$(fixture_path "$stale:$fresh")" "$out" assert_contains "$(cat "$out")" "[truncated]" "the first report was not long enough to be cut, so this case proves nothing" # The skew tool goes last, so its finding falls past the cut and the printed # line is byte identical to the one the first sweep already recorded. - write_config "$home" "{\"tools\":[$tools,{\"name\":\"herdr\",\"command\":\"$TOOL\"}]}" + write_config "$home" "{\"tools\":[$tools_json,{\"name\":\"herdr\",\"command\":\"$TOOL\"}]}" run_check "$home" "$(fixture_path "$stale:$fresh")" "$out" report=$(cat "$out") [ -n "$report" ] || fail "a finding past the cut produced no report at all, so it can never reach the watcher" diff --git a/tests/fm-trace-context-spawn.test.sh b/tests/fm-trace-context-spawn.test.sh index 567a2e0adcc..d38fbf787d1 100755 --- a/tests/fm-trace-context-spawn.test.sh +++ b/tests/fm-trace-context-spawn.test.sh @@ -258,6 +258,7 @@ test_enabled_records_and_injects_identical_carrier_before_launch() { local rec out status meta mtp itp gl tl ll rec=$(make_spawn_case tc-on) read_case_record "$rec" + : > "$HOME_DIR/config/launch-env-allowlist" : > "$HOME_DIR/config/trace-context" # enable via the real config path start_trace_session "$HOME_DIR" diff --git a/tests/fm-turnend-guard.test.sh b/tests/fm-turnend-guard.test.sh index 02d5bf3509f..51fffdf46b2 100755 --- a/tests/fm-turnend-guard.test.sh +++ b/tests/fm-turnend-guard.test.sh @@ -99,6 +99,73 @@ test_predicate_source_needs_supervision() { pass "fm_supervision_unhealthy: source-only home needs supervision" } +# Register a custom check the way an operator does, through the real +# bin/fm-check-register.sh, so these cases bind to the shipped registration +# artifacts rather than to a hand-written imitation of them. +register_custom_check() { + local state=$1 id=$2 + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/$id.check.sh" + chmod 700 "$state/$id.check.sh" + FM_STATE_OVERRIDE="$state" "$ROOT/bin/fm-check-register.sh" "$id" >/dev/null \ + || fail "fm-check-register.sh could not register $id" +} + +test_predicate_registered_check_needs_supervision() { + local state="$TMP_ROOT/pred-check/state" + mkdir -p "$state" + register_custom_check "$state" issue-comments + fm_supervision_needed "$state" 300 || fail "a registered custom check did not register as supervision need" + [ "$FM_SUP_IN_FLIGHT" -eq 0 ] || fail "a registered custom check must not count as an in-flight task" + [ "$FM_SUP_CHECKS" -eq 1 ] || fail "expected one registered custom check, got $FM_SUP_CHECKS" + fm_supervision_unhealthy "$state" 300 || fail "a registered custom check with no beacon must be unhealthy" + pass "fm_supervision_needed: a registered custom check needs supervision with no task in flight" +} + +test_predicate_registered_check_survives_rebinding_drift() { + local state="$TMP_ROOT/pred-check-drift/state" + mkdir -p "$state" + register_custom_check "$state" issue-comments + printf '#!/usr/bin/env bash\necho drifted\n' > "$state/issue-comments.check.sh" + fm_supervision_needed "$state" 300 \ + || fail "an edited registered check must keep supervision on so the sweep can report the rejection" + [ "$FM_SUP_CHECKS" -eq 1 ] || fail "expected the edited check to stay counted, got $FM_SUP_CHECKS" + pass "fm_supervision_needed: a registered check whose bytes drifted still needs supervision" +} + +test_predicate_unregistered_check_needs_nothing() { + local state="$TMP_ROOT/pred-check-unregistered/state" + mkdir -p "$state" + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/rogue.check.sh" + chmod 700 "$state/rogue.check.sh" + if fm_supervision_needed "$state" 300; then + fail "a check with no trust binding must not arm supervision" + fi + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "an unregistered check must not be counted, got $FM_SUP_CHECKS" + pass "fm_supervision_needed: false for a check.sh with no registration binding" +} + +test_predicate_task_pr_poll_is_not_a_custom_check() { + local state="$TMP_ROOT/pred-pr-poll/state" + mkdir -p "$state" + : > "$state/task1.meta" + printf '#!/usr/bin/env bash\nexit 0\n' > "$state/task1.check.sh" + chmod 700 "$state/task1.check.sh" + : > "$state/task1.pr-poll" + fm_supervision_needed "$state" 300 || fail "the in-flight task itself must need supervision" + [ "$FM_SUP_IN_FLIGHT" -eq 1 ] || fail "expected the task to be the one in-flight need, got $FM_SUP_IN_FLIGHT" + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "a task PR poll must not count as a registered custom check" + pass "fm_supervision_needed: a task PR poll without a trust binding is not a registered custom check" +} + +test_predicate_relay_shim_is_not_a_custom_check() { + local state="$TMP_ROOT/pred-relay-not-custom/state" + mkdir -p "$state" + : > "$state/x-watch.check.sh" + fm_supervision_needed "$state" 300 || fail "the relay poll must still need supervision" + [ "$FM_SUP_CHECKS" -eq 0 ] || fail "the relay shim keeps its own trust path and must not be counted as a custom check" + pass "fm_supervision_status: the relay shim is not counted as a registered custom check" +} + # --- HOOK: bin/fm-turnend-guard.sh ------------------------------------------ # # Each scenario gets its own directory carrying a copy of the two guard scripts @@ -336,6 +403,7 @@ Codex|{"cwd":"$dir","stop_hook_active":false} OpenCode|{"stop_hook_active":false} Pi|{"stop_hook_active":false} pi-signed|{"stop_hook_active":false} +omp|{"stop_hook_active":false} Grok|{"sessionId":"grok-session","stopHookActive":false} Kimi|{"stop_hook_active":false} EOF @@ -411,6 +479,17 @@ test_hook_x_mode_only_blocks_in_default_mode() { pass "fm-turnend-guard: X-mode-only supervision remains guarded in default mode" } +test_hook_registered_check_only_blocks_with_check_banner() { + local dir out status + dir=$(make_primary_dir "$TMP_ROOT/hook-check-only") + register_custom_check "$dir/state" issue-comments + out=$(run_hook "$dir" false); status=$? + expect_code 2 "$status" "default hook mode must block a registered-check-only blind turn" + assert_contains "$out" "1 registered custom check(s), but no live watcher" "check-only blind stop must identify its supervision need" + assert_not_contains "$out" "X-mode relay polling needs supervision" "check-only blind stop must not be misreported as relay polling" + pass "fm-turnend-guard: registered-check-only supervision is named in the block banner" +} + test_hook_ignores_repo_state_when_fm_home_set() { local dir home out status dir=$(make_primary_dir "$TMP_ROOT/hook-fm-home-ignore-root") @@ -1982,6 +2061,95 @@ test_hook_daemon_lock_is_ignored_without_away_mode() { pass "fm-turnend-guard: a daemon lock proves nothing while away mode is off" } +# --- AWAY MODE: beacon grace derives from the poll cadence ------------------- +# +# The daemon starts a fresh one-shot watcher only after it finishes handling +# the previous wake, and that handling can legitimately outrun a flat 300s +# window under load (a slow registered check, a busy supervisor pane) with the +# daemon perfectly healthy throughout (fm-turnend-guard-afk-race). The guard +# must accept a live daemon there once FM_POLL justifies the wider window, but +# must still block a dead daemon or a beacon older than that wider grace. + +test_hook_away_daemon_allows_beacon_within_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-healthy") + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live away-mode daemon holder" + } + # 400s is stale under the flat 300s default, but not under the poll-derived + # grace (max(300, FM_POLL + 60) = 660 at FM_POLL=600) - a live daemon that + # simply has not finished restarting its watcher yet. + beat=$(( $(date +%s) - 400 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 0 "$status" "a live daemon with a beacon within the poll-derived grace must not block" + [ -z "$out" ] || fail "away-mode daemon within poll-derived grace still produced a block banner: $out" + pass "fm-turnend-guard: away-mode beacon freshness uses the poll-derived grace, not the flat default" +} + +test_hook_away_daemon_blocks_dead_daemon_despite_poll_derived_grace() { + local dir dead out status + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-dead-daemon") + dead=$(nonexistent_pid) + record_daemon_lock "$dir" "$dead" "dead daemon identity" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + expect_code 2 "$status" "a wider poll-derived grace must not paper over a dead daemon" + assert_contains "$out" "$AWAY_REQUIRED_REASON" "away-mode block must point at the daemon, not normal supervision" + pass "fm-turnend-guard: a dead away-mode daemon still blocks under the poll-derived grace" +} + +test_hook_away_daemon_blocks_beacon_older_than_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-afk-poll-grace-stale") + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live away-mode daemon holder" + } + # 700s exceeds even the wider poll-derived grace (660 at FM_POLL=600), so a + # live daemon that has genuinely stopped restarting its watcher still blocks. + beat=$(( $(date +%s) - 700 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "a beacon older than the poll-derived grace must still block" + assert_contains "$out" "$AWAY_REQUIRED_REASON" "away-mode block must point at the daemon, not normal supervision" + pass "fm-turnend-guard: the poll-derived grace is bounded, not unlimited" +} + +test_hook_no_afk_ignores_poll_derived_grace() { + local dir pid out status beat + dir=$(make_away_home_between_cycles "$TMP_ROOT/hook-no-afk-poll-grace") + rm -f "$dir/state/.afk" + sleep 60 & + pid=$! + record_daemon_lock "$dir" "$pid" || { + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + fail "could not identify live daemon holder" + } + # 400s would be within the poll-derived grace the away-mode branch would + # accept, but away mode is off here, so the strict watcher predicate and its + # flat default govern instead - old behavior, unaffected by FM_POLL. + beat=$(( $(date +%s) - 400 )) + touch -d "@$beat" "$dir/state/.last-watcher-beat" + out=$(FM_GUARD_GRACE='' FM_POLL=600 run_hook "$dir" false); status=$? + kill "$pid" 2>/dev/null || true + wait "$pid" 2>/dev/null || true + expect_code 2 "$status" "without .afk, FM_POLL must not widen the strict watcher predicate's grace" + assert_contains "$out" "$REQUIRED_REASON" "block reason must contain the exact required instruction" + pass "fm-turnend-guard: with away mode off, the poll-derived grace never applies" +} + test_predicate_healthy_no_inflight test_predicate_unhealthy_no_beacon test_predicate_unhealthy_stale_beacon @@ -1989,6 +2157,11 @@ test_predicate_healthy_fresh_beacon test_predicate_queue_pending_flag test_predicate_x_mode_needs_supervision test_predicate_source_needs_supervision +test_predicate_registered_check_needs_supervision +test_predicate_registered_check_survives_rebinding_drift +test_predicate_unregistered_check_needs_nothing +test_predicate_task_pr_poll_is_not_a_custom_check +test_predicate_relay_shim_is_not_a_custom_check test_hook_silent_when_no_work_in_flight test_hook_blocks_when_fresh_beacon_has_no_live_lock test_hook_blocks_source_only_home @@ -2000,6 +2173,7 @@ test_hook_blocks_when_unhealthy_in_primary test_hook_blocks_from_fm_home_state test_hook_x_mode_reason_sources_cadence test_hook_x_mode_only_blocks_in_default_mode +test_hook_registered_check_only_blocks_with_check_banner test_hook_ignores_repo_state_when_fm_home_set test_hook_uses_state_override test_hook_loop_guard_allows_retry @@ -2060,3 +2234,7 @@ test_hook_away_mode_blocks_on_dead_daemon test_hook_away_mode_blocks_on_pid_reused_daemon test_hook_away_mode_blocks_on_stale_beacon test_hook_daemon_lock_is_ignored_without_away_mode +test_hook_away_daemon_allows_beacon_within_poll_derived_grace +test_hook_away_daemon_blocks_dead_daemon_despite_poll_derived_grace +test_hook_away_daemon_blocks_beacon_older_than_poll_derived_grace +test_hook_no_afk_ignores_poll_derived_grace diff --git a/tests/fm-wake-queue.test.sh b/tests/fm-wake-queue.test.sh index a9ca08be5f4..9e7faea0335 100755 --- a/tests/fm-wake-queue.test.sh +++ b/tests/fm-wake-queue.test.sh @@ -14,6 +14,7 @@ set -u WATCH="$ROOT/bin/fm-watch.sh" DRAIN="$ROOT/bin/fm-wake-drain.sh" GRANT="$ROOT/bin/fm-wake-grant.sh" +GUARD="$ROOT/bin/fm-guard.sh" TMP_ROOT=$(fm_test_tmproot fm-wake-tests) @@ -231,8 +232,8 @@ test_drain_dedupes_obvious_duplicates() { # plain drain-and-handle turn that runs no other supervision script. It must warn # when work is in flight with no live watcher, and stay silent right after a # normal fire from a live watcher with a fresh beacon, so it never false-alarms. -test_secondmate_foreign_queue_stall_is_one_shot_and_read_only() { - local dir state sub fakebin out row_before row_after stall_count +test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once() { + local dir state sub fakebin out row_before row_after stall_count real_date dir=$(make_case secondmate-foreign-stall) state="$dir/state" sub="$dir/secondmate" @@ -241,35 +242,58 @@ test_secondmate_foreign_queue_stall_is_one_shot_and_read_only() { printf 'mate\n' > "$sub/.fm-secondmate-home" printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ "$sub" > "$state/mate.meta" - printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" - row_before="$dir/foreign-before" - row_after="$dir/foreign-after" - cp "$sub/state/.wake-queue" "$row_before" fakebin="$dir/fakebin" - cat > "$fakebin/tmux" <<'SH' + real_date=$(command -v date) + cat > "$fakebin/date" <<SH #!/usr/bin/env bash -case "${1:-}" in - list-windows) printf '%s\n' "${FM_FAKE_TMUX_WINDOW:-}" ;; - capture-pane) cat "${FM_FAKE_TMUX_CAPTURE:-/dev/null}" ;; - display-message) printf '0\n' ;; - *) exit 0 ;; -esac +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi SH - chmod +x "$fakebin/tmux" - out="$dir/watch.out" + chmod +x "$fakebin/date" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # An already-old row starts an observation interval; its creation time alone + # cannot produce an alert. + printf '1000\n' > "$dir/now" + printf '100\t7\tcheck\trouted\tcheck: routed row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 3 > "$out" 2> "$dir/watch.err" || true - grep -F 'check: secondmate wake-loop stalled: mate=mate row=7' "$out" >/dev/null \ - || fail "an aged foreign row did not wake the parent checkpoint: $(cat "$out"); err=$(cat "$dir/watch.err"); meta=$(cat "$state/mate.meta"); foreign=$(cat "$sub/state/.wake-queue")" - [ -s "$state/.wake-queue" ] || fail "the parent notification was not durable" - stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) - [ "$stall_count" -eq 1 ] || fail "the first parent checkpoint did not publish exactly one stall notification" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-first.out" 2> "$dir/watch-first.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "the first observation of an old foreign row produced an age-only alert" + + # The oldest sequence advances after more than the threshold. This is healthy + # drain progress even though the replacement row is itself very old. + printf '1002\n' > "$dir/now" + printf '100\t8\tcheck\thealthy\tcheck: healthy progress\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-progress.out" 2> "$dir/watch-progress.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "an advancing foreign queue produced a stall alert because its oldest row was old" + # With no further sequence progress, the same queue must still expose the real + # failure after the configured interval. + printf '1004\n' > "$dir/now" + row_before="$dir/foreign-before" + row_after="$dir/foreign-after" + cp "$sub/state/.wake-queue" "$row_before" + out="$dir/watch-stalled.out" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$out" 2> "$dir/watch-stalled.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=8 idle=2s' "$out" >/dev/null \ + || fail "a foreign queue with no progress did not alert: $(cat "$out")" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the stalled episode did not publish exactly one parent notification" cmp -s "$row_before" "$sub/state/.wake-queue" \ || fail "foreign queue row changed during read-only stall detection" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ @@ -277,55 +301,210 @@ SH ack_drain_err "$state" "$dir/drain.err" \ || fail "parent stall notification could not be acknowledged" - sleep 1 - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # Partial draining changes the oldest row, ends the prior no-progress episode, + # and cannot produce an immediate notification cascade. + printf '1010\n' > "$dir/now" + printf '100\t9\tcheck\tnext\tcheck: next row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-second.out" 2> "$dir/watch-second.err" || true - [ ! -s "$state/.wake-queue" ] || { - stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) - [ "$stall_count" -eq 0 ] || fail "repeated checkpoint re-published the same stall notification" - } + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-next.out" 2> "$dir/watch-next.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "a newly-oldest row cascaded an immediate second alert after progress" cp "$sub/state/.wake-queue" "$row_after" - cmp -s "$row_before" "$row_after" || fail "foreign queue changed after idempotent re-check" + grep -F $'\t9\t' "$row_after" >/dev/null || fail "foreign queue progress fixture changed during observation" - : > "$sub/state/.wake-queue" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + # If that new drain position then genuinely stops advancing, it is a new + # no-progress episode and must remain visible rather than being muted forever. + printf '1012\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-empty.out" 2> "$dir/watch-empty.err" || true - ! grep -F 'secondmate wake-loop stalled' "$dir/watch-empty.out" >/dev/null \ - || fail "an empty foreign queue produced a stall notification" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-refrozen.out" 2> "$dir/watch-refrozen.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=9 idle=2s' "$dir/watch-refrozen.out" >/dev/null \ + || fail "a genuine later no-progress episode was hidden after earlier progress" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the later no-progress episode did not publish exactly one notification" + pass "foreign secondmate queue alerts once per no-progress episode without age-only or cascade noise" +} - printf '%s\t8\tcheck\thealthy\tcheck: healthy row\n' "$(date +%s)" > "$sub/state/.wake-queue" - PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ +test_secondmate_declared_pause_rows_do_not_feed_stall_escalation() { + local dir state sub fakebin real_date + dir=$(make_case secondmate-declared-pause-queue) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nhome=%s\n' "$sub" > "$state/mate.meta" + fakebin="$dir/fakebin" + real_date=$(command -v date) + cat > "$fakebin/date" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi +SH + chmod +x "$fakebin/date" + cat > "$sub/state/.wake-queue" <<'EOF' +100 7 stale fleet:w2:p4 stale: fleet:w2:p4 (paused 3613s, awaiting external - declared paused) +100 8 stale fleet:w2:p3 stale: fleet:w2:p3 (paused 3615s, awaiting external - declared pause, rechecked on a long cadence not a wedge) +EOF + printf '1000\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ - FM_FAKE_TMUX_LOG="$dir/tmux.log" FM_FAKE_TMUX_CAPTURE="$dir/fake-tmux/pane.txt" \ - FM_SECONDMATE_WAKE_STALL_SECS=60 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-first.out" 2> "$dir/watch-first.err" || true + printf '5000\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-second.out" 2> "$dir/watch-second.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "declared external-wait rows fed the secondmate wake-loop escalation" + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-first.out" "$dir/watch-second.out" >/dev/null \ + || fail "a declared external wait was mislabeled as a stalled wake loop" + pass "declared external-wait pause rows do not feed secondmate wake-loop escalation" +} + +# A retired mate reprovisioned under the same task id gets a fresh home, so its +# wake-queue sequence restarts from scratch and can land on the very position the +# parent last recorded for the retired generation. Those are different rows in +# different queue generations, not the continuation of the previous generation's +# no-progress interval: inheriting that interval fires a wake-loop stall against +# a queue the mate has only just created. +test_secondmate_reprovisioned_queue_starts_a_fresh_interval() { + local dir state sub fakebin real_date + dir=$(make_case secondmate-reprovisioned-queue) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + fakebin="$dir/fakebin" + real_date=$(command -v date) + cat > "$fakebin/date" <<SH +#!/usr/bin/env bash +if [ "\${1:-}" = +%s ]; then + cat "\${FM_FAKE_NOW_FILE:?}" +else + exec "$real_date" "\$@" +fi +SH + chmod +x "$fakebin/date" + + # The retired generation's last observation records sequence 9. + printf '1000\n' > "$dir/now" + printf '100\t9\tcheck\told\tcheck: retired generation row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ - "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 2 > "$dir/watch-healthy.out" 2> "$dir/watch-healthy.err" || true - ! grep -F 'secondmate wake-loop stalled' "$dir/watch-healthy.out" >/dev/null \ - || fail "a healthy foreign queue produced a stall notification" - pass "foreign secondmate queue stalls notify once, remain byte-stable, and stay quiet when empty or healthy" + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-old.out" 2> "$dir/watch-old.err" || true + [ ! -s "$state/.wake-queue" ] || fail "the first observation of the retired generation alerted" + + # Reprovisioning under the same task id restarts the sequence on 9 again, long + # after the recorded observation. That first sight of the new queue cannot + # inherit the old generation's idle interval. + printf '1010\n' > "$dir/now" + printf '200\t9\tcheck\tregen\tcheck: reprovisioned row\n' > "$sub/state/.wake-queue" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-regen.out" 2> "$dir/watch-regen.err" || true + [ ! -s "$state/.wake-queue" ] \ + || fail "a reprovisioned queue generation inherited the retired generation's idle interval and alerted" + + # The restarted generation still earns its own honest no-progress episode. + printf '1012\n' > "$dir/now" + PATH="$fakebin:$PATH" FM_FAKE_NOW_FILE="$dir/now" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_FAKE_TMUX_WINDOW='firstmate:fm-mate' \ + FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 FM_SIGNAL_GRACE=0 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 1 > "$dir/watch-regen-frozen.out" 2> "$dir/watch-regen-frozen.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=9 idle=2s' "$dir/watch-regen-frozen.out" >/dev/null \ + || fail "a frozen reprovisioned queue generation was hidden: $(cat "$dir/watch-regen-frozen.out")" + pass "a reprovisioned queue generation starts a fresh no-progress interval" +} + +# A healthy mate drains its wake queue BETWEEN turns, not inside one, so a queue +# that has not advanced while the mate is provably mid-turn is not a stalled wake +# loop - it is the normal state of a busy mate, and the measured false alarms +# (ages 63s, 75s, 70s) all landed here. The active-turn gate must DEFER that +# escalation, not cancel it: the same frozen queue still has to surface once the +# turn ends. +test_secondmate_active_turn_defers_stall_until_the_turn_ends() { + local dir state sub fakebin stall_count + dir=$(make_case secondmate-active-turn) + state="$dir/state" + sub="$dir/secondmate" + mkdir -p "$sub/state" + printf 'mate\n' > "$sub/.fm-secondmate-home" + printf 'window=firstmate:fm-mate\nkind=secondmate\nharness=claude\nbackend=tmux\nhome=%s\n' \ + "$sub" > "$state/mate.meta" + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" \ + > "$sub/state/.wake-queue" + fakebin="$dir/fakebin" + cat > "$fakebin/tmux" <<'SH' +#!/usr/bin/env bash +case "${1:-}" in + list-windows) printf '%s\n' 'firstmate:fm-mate' ;; + capture-pane) printf 'working\n' ;; + display-message) printf '0\n' ;; + *) exit 0 ;; +esac +SH + chmod +x "$fakebin/tmux" + "$ROOT/bin/fm-busy-event.sh" arm "$state" mate >/dev/null \ + || fail "could not arm the mate's busy contract" + + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 \ + FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 4 \ + > "$dir/watch-busy.out" 2> "$dir/watch-busy.err" || true + ! grep -F 'secondmate wake-loop stalled' "$dir/watch-busy.out" >/dev/null \ + || fail "a mate inside an active turn was escalated as a stalled wake loop: $(cat "$dir/watch-busy.out")" + [ ! -s "$state/.wake-queue" ] \ + || fail "a mate inside an active turn published a durable stall notification" + + "$ROOT/bin/fm-busy-event.sh" apply "$state" mate idle --current-gen \ + --source claude-hook --event stop >/dev/null \ + || fail "could not end the mate's turn" + PATH="$fakebin:$PATH" FM_HOME="$dir" FM_ROOT_OVERRIDE="$ROOT" \ + FM_STATE_OVERRIDE="$state" FM_SECONDMATE_WAKE_STALL_SECS=1 FM_POLL=1 \ + FM_SIGNAL_GRACE=0 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \ + "$ROOT/bin/fm-watch-checkpoint.sh" --seconds 4 \ + > "$dir/watch-idle.out" 2> "$dir/watch-idle.err" || true + grep -F 'check: secondmate wake-loop stalled: mate=mate row=7' "$dir/watch-idle.out" >/dev/null \ + || fail "the same frozen queue stayed hidden after the turn ended: $(cat "$dir/watch-idle.out")" + stall_count=$(grep -c 'secondmate-wake-loop-mate-' "$state/.wake-queue" || true) + [ "$stall_count" -eq 1 ] || fail "the deferred episode did not publish exactly one notification" + pass "an active turn defers the secondmate stall escalation without cancelling it" } test_secondmate_stall_marker_rejects_symlink() { - local dir state sub fakebin marker outside expected + local dir state sub fakebin marker outside expected epoch dir=$(make_case secondmate-stall-marker-symlink) state="$dir/state" sub="$dir/secondmate" mkdir -p "$sub/state" printf 'mate\n' > "$sub/.fm-secondmate-home" printf 'window=firstmate:fm-mate\nkind=secondmate\nhome=%s\n' "$sub" > "$state/mate.meta" - printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$(( $(date +%s) - 10 ))" > "$sub/state/.wake-queue" + epoch=$(( $(date +%s) - 10 )) + printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$sub/state/.wake-queue" outside="$dir/outside" expected='must remain unchanged' printf '%s\n' "$expected" > "$outside" marker="$state/.secondmate-wake-stall-mate" + printf '%s\t%s-7\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-mate" ln -s "$outside" "$marker" fakebin="$dir/fakebin" cat > "$fakebin/tmux" <<'SH' @@ -363,8 +542,9 @@ test_acknowledged_stall_publication_survives_pre_marker_crash() { printf '%s\t7\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$sub/state/.wake-queue" row_before="$dir/foreign-before" cp "$sub/state/.wake-queue" "$row_before" + printf '%s\t%s-7\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-mate" append_wake "$state" check "secondmate-wake-loop-mate-$epoch-7" \ - "check: secondmate wake-loop stalled: mate=mate row=7 age=10s" \ + "check: secondmate wake-loop stalled: mate=mate row=7 idle=2s" \ || fail "could not seed the pre-marker crash publication" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ || fail "pre-marker crash publication could not be drained" @@ -404,8 +584,9 @@ test_empty_prefix_mate_preserves_other_mate_receipt() { printf '%s\t9\tcheck\trouted\tcheck: routed row\n' "$epoch" > "$stalled/state/.wake-queue" row_before="$dir/foreign-before" cp "$stalled/state/.wake-queue" "$row_before" + printf '%s\t%s-9\n' "$(( $(date +%s) - 2 ))" "$epoch" > "$state/.secondmate-wake-progress-ios-ui" append_wake "$state" check "secondmate-wake-loop-ios-ui-$epoch-9" \ - "check: secondmate wake-loop stalled: mate=ios-ui row=9 age=10s" \ + "check: secondmate wake-loop stalled: mate=ios-ui row=9 idle=2s" \ || fail "could not seed the ios-ui stall publication" FM_STATE_OVERRIDE="$state" "$DRAIN" > "$dir/drain.out" 2> "$dir/drain.err" \ || fail "ios-ui stall publication could not be drained" @@ -703,6 +884,164 @@ test_main_drain_excludes_rows_already_granted_to_branch() { pass "main drain and acknowledgement exclude an active branch grant" } +# The pending-warning condition and what a drain can actually present must name +# the same rows. A row reserved by a live branch grant is invisible to a main +# drain by design, so counting it as "queued for main" told main to run a drain +# that could only print nothing - no row, no acknowledgement command - on every +# guarded command, for as long as the branch held the grant. +test_main_is_never_told_to_drain_rows_only_the_branch_owns() { + local dir state out err sequence generation + dir=$(make_case main-not-told-to-drain-branch-rows) + state="$dir/state" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + append_wake "$state" stale "fleet:w2:p3" "stale: fleet:w2:p3 (paused, awaiting external)" \ + || fail "stale append failed" + FM_STATE_OVERRIDE="$state" "$GRANT" activate "$$" held-by-branch || fail "branch owner activation failed" + FM_STATE_OVERRIDE="$state" "$GRANT" publish held-by-branch 1 || fail "branch grant publication failed" + + out="$dir/main-drain.out" + err="$dir/main-drain.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed: $(cat "$err")" + ! grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$out" || fail "main drain presented a branch-granted row" + grep -Fq 'WAKE ROWS HELD BY SUPERVISION BRANCH' "$out" \ + || fail "main drain went silent instead of naming who holds the queued rows" + ! grep -Fq 'WAKE_ACK_REQUIRED' "$err" || fail "main drain offered an acknowledgement for a row it never presented" + ! grep -Fq 'queued wakes pending' "$err" \ + || fail "main was told to drain rows only the branch can present" + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-held.err" || fail "guard failed while the branch held the rows" + ! grep -Fq 'queued wakes pending' "$dir/guard-held.err" \ + || fail "guard counted branch-held rows as pending for main" + grep -Fq 'wake rows held by the live supervision branch' "$dir/guard-held.err" \ + || fail "guard went silent about a non-empty queue instead of naming the branch as its holder" + grep -Fq 'do not drain them from here' "$dir/guard-held.err" \ + || fail "the held advisory did not say the rows must not be drained from here" + grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$state/.wake-queue" \ + || fail "the branch-held row must stay durable for its own owner" + + # Disconfirming half: the same row, same kind, same stopped endpoint, with the + # grant released. Nothing about the row makes it unpresentable - only the + # live grant did - so main now presents it with an executable acknowledgement. + FM_STATE_OVERRIDE="$state" "$GRANT" release held-by-branch || fail "branch grant release failed" + out="$dir/main-drain-after.out" + err="$dir/main-drain-after.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed after release: $(cat "$err")" + grep -Fq "$(printf '\tstale\tfleet:w2:p3\t')" "$out" || fail "main drain omitted the released row" + ! grep -Fq 'WAKE ROWS HELD BY SUPERVISION BRANCH' "$out" \ + || fail "main drain reported a hold that no longer exists" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") + [ -n "$sequence" ] && [ -n "$generation" ] || fail "the released row was presented without an acknowledgement command" + grep -Fq 'queued wakes pending' "$err" || fail "guard stopped warning about a row main can actually drain" + ! grep -Fq 'wake rows held by the live supervision branch' "$err" \ + || fail "guard kept advising about a hold that was already released" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "acknowledgement of the released row failed" + [ ! -s "$state/.wake-queue" ] || fail "the acknowledged row stayed queued" + + pass "a branch-held row raises no queued-wake warning for main, and the same row is presented and acknowledged once the grant clears" +} + +# The pending-warning condition must also survive a queue nobody could read: a +# queue that exists but cannot be counted is not evidence that it was drained. +# The per-actor count runs awk over the queue, and awk implementations differ on +# whether a failed input open aborts before the END rule; one that reaches END +# reports a 0 count for a queue that was never proved empty. +test_uncountable_queue_still_raises_the_pending_alarm() { + local dir state awkbin real_awk + dir=$(make_case uncountable-queue) + state="$dir/state" + awkbin="$dir/awkbin" + mkdir -p "$awkbin" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + # An awk that still runs its END rule after failing to open its input: it + # prints a 0 count and exits non-zero. Every other invocation is the real awk. + real_awk=$(command -v awk) || fail "no awk on PATH" + cat > "$awkbin/awk" <<SH +#!/usr/bin/env bash +set -u +for _arg in "\$@"; do _last=\$_arg; done +if [ -n "\${_last:-}" ] && [ -e "\$_last" ] && [ ! -r "\$_last" ]; then + printf '0\\n' + exit 2 +fi +exec "$real_awk" "\$@" +SH + chmod +x "$awkbin/awk" + + append_wake "$state" stale "fleet:w2:p3" "stale: fleet:w2:p3 (paused, awaiting external)" \ + || fail "stale append failed" + chmod 000 "$state/.wake-queue" || fail "could not make the queue unreadable" + PATH="$awkbin:$PATH" FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/unreadable.err" \ + || fail "guard failed on an unreadable queue" + grep -Fq 'queued wakes pending' "$dir/unreadable.err" \ + || fail "a queue that could not be counted silenced the queued-wake alarm" + chmod 600 "$state/.wake-queue" || fail "could not restore the queue" + + # Disconfirming half: the same fake awk over a queue that is readable and + # provably empty stays silent, so the warning above came from the failed count + # and not from the fake awk itself. + : > "$state/.wake-queue" + PATH="$awkbin:$PATH" FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/empty.err" \ + || fail "guard failed on an empty queue" + ! grep -Fq 'queued wakes pending' "$dir/empty.err" \ + || fail "a provably empty queue raised the queued-wake alarm" + + pass "a queue that cannot be counted keeps the queued-wake alarm up" +} + +# A row that lost its structure can never be claimed, presented, or named by an +# --ack-through cutoff, while it still counts as queued: without retirement it +# wedges the queue permanently and keeps waking supervision. +test_unconsumable_rows_are_retired_instead_of_wedging_the_queue() { + local dir state out err sequence generation + dir=$(make_case unconsumable-row-retirement) + state="$dir/state" + printf 'window=test:fm-x\nkind=ship\n' > "$state/x.meta" + + append_wake "$state" signal "task-a.status" "signal: task-a" || fail "signal append failed" + printf '1788792074\t574\tstale\tfleet:w2:p3\n' >> "$state/.wake-queue" + printf '1788792075\tnot-a-sequence\tstale\tfleet:w2:p4\tstale: fleet:w2:p4\n' >> "$state/.wake-queue" + + # A branch actor never repairs the queue: it may only touch its own grant. + FM_STATE_OVERRIDE="$state" "$GRANT" activate "$$" retire-scope || fail "branch owner activation failed" + FM_STATE_OVERRIDE="$state" "$GRANT" publish retire-scope 1 || fail "branch grant publication failed" + FM_STATE_OVERRIDE="$state" FM_SUPERVISION_ACTOR=branch "$DRAIN" > "$dir/branch.out" 2> "$dir/branch.err" \ + || fail "branch drain failed: $(cat "$dir/branch.err")" + ! grep -Fq 'retired' "$dir/branch.err" || fail "a branch drain retired rows outside its grant" + [ "$(awk 'END { print NR }' "$state/.wake-queue")" -eq 3 ] \ + || fail "a branch drain changed rows it was never granted" + FM_STATE_OVERRIDE="$state" "$GRANT" release retire-scope || fail "branch grant release failed" + + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-before.err" || fail "guard failed with unusable rows queued" + grep -Fq 'queued wakes pending' "$dir/guard-before.err" \ + || fail "guard stayed silent about rows main still has to clear" + ! grep -Fq 'wake rows held by the live supervision branch' "$dir/guard-before.err" \ + || fail "guard advised a branch hold for rows no grant covers" + + out="$dir/main.out" + err="$dir/main.err" + FM_STATE_OVERRIDE="$state" "$DRAIN" > "$out" 2> "$err" || fail "main drain failed: $(cat "$err")" + grep -Fq 'retired 2 unusable queue row(s)' "$err" || fail "main drain did not report the rows it retired" + grep -Fq "$(printf '1788792074\t574\tstale\tfleet:w2:p3')" "$err" \ + || fail "the retired row's content was discarded instead of reported" + grep -Fq "$(printf '1788792075\tnot-a-sequence\tstale\tfleet:w2:p4\tstale: fleet:w2:p4')" "$err" \ + || fail "the second retired row's content was discarded instead of reported" + grep -Fq "$(printf '\tsignal\ttask-a.status\t')" "$out" || fail "retirement dropped a usable row" + [ "$(awk 'END { print NR }' "$state/.wake-queue")" -eq 1 ] || fail "unusable rows survived the drain" + sequence=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err") + generation=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err") + [ -n "$sequence" ] && [ -n "$generation" ] || fail "the usable row was presented without an acknowledgement command" + FM_STATE_OVERRIDE="$state" "$DRAIN" --ack-through "$sequence" --recovery-generation "$generation" \ + || fail "acknowledgement failed" + [ ! -s "$state/.wake-queue" ] || fail "the queue stayed wedged after acknowledgement" + FM_STATE_OVERRIDE="$state" "$GUARD" 2> "$dir/guard-after.err" || fail "guard failed after the queue drained" + ! grep -Fq 'queued wakes pending' "$dir/guard-after.err" || fail "guard kept warning about an empty queue" + + pass "structurally unusable rows are retired by main alone, leaving every remaining row presentable and acknowledgeable" +} + test_branch_grant_refuses_rows_already_claimed_by_main() { local dir state rc dir=$(make_case branch-refuses-main-claim) @@ -1248,6 +1587,33 @@ test_self_held_lock_reclaims_instead_of_deadlocking() { pass "an abandoned same-process lock hold is reclaimed; a parent's live hold is not" } +test_subshell_lock_ownership_without_bashpid() { + local dir state rc + dir=$(make_case subshell-lock-ownership) + state="$dir/state" + rc=0 + FM_STATE_OVERRIDE="$state" bash -c ' + unset BASHPID + . "$1" + lock="$2/.fixture.lock" + fm_lock_acquire_wait "$lock" || exit 10 + ( fm_lock_release "$lock" ) + [ "$(cat "$lock/pid")" = "$$" ] || exit 11 + ( fm_lock_try_acquire "$lock" && exit 12; exit 0 ) || exit 12 + fm_lock_release "$lock" + ( + fm_lock_acquire_wait "$lock" || exit 13 + [ "$(cat "$lock/pid")" != "$$" ] || exit 14 + fm_lock_try_acquire "$lock" || exit 15 + fm_lock_set_role "$lock" terminal-check || exit 16 + fm_lock_release "$lock" + [ ! -e "$lock" ] && [ ! -L "$lock" ] || exit 17 + ) || exit $? + ' _ "$ROOT/bin/fm-wake-lib.sh" "$state" || rc=$? + [ "$rc" -eq 0 ] || fail "subshell lock ownership without BASHPID failed (rc=$rc)" + pass "without BASHPID a subshell cannot release or reclaim its parent lock and owns its own hold" +} + # A bounded waiter acquires in a helper process, but the caller must own the # lock once contention clears so it can safely hold and release the critical # section itself. @@ -1542,10 +1908,14 @@ test_historical_annotation_skips_announced_status() { } test_self_held_lock_reclaims_instead_of_deadlocking +test_subshell_lock_ownership_without_bashpid test_bounded_lock_handoff_after_contention test_live_presentation_holder_is_deadlined_without_weakening_ack test_malformed_presentation_lock_reports_acquire_failure -test_secondmate_foreign_queue_stall_is_one_shot_and_read_only +test_secondmate_foreign_queue_stall_tracks_progress_and_alerts_once +test_secondmate_declared_pause_rows_do_not_feed_stall_escalation +test_secondmate_reprovisioned_queue_starts_a_fresh_interval +test_secondmate_active_turn_defers_stall_until_the_turn_ends test_secondmate_stall_marker_rejects_symlink test_acknowledged_stall_publication_survives_pre_marker_crash test_empty_prefix_mate_preserves_other_mate_receipt @@ -1564,6 +1934,9 @@ test_enrichment_preserves_all_unread_lines_and_status_file_failures test_slow_annotation_does_not_block_append_and_deleted_file_fails_open test_branch_actor_scoped_ack_never_swallows_a_main_owned_row test_main_drain_excludes_rows_already_granted_to_branch +test_main_is_never_told_to_drain_rows_only_the_branch_owns +test_uncountable_queue_still_raises_the_pending_alarm +test_unconsumable_rows_are_retired_instead_of_wedging_the_queue test_branch_grant_refuses_rows_already_claimed_by_main test_actor_filter_precedes_same_key_deduplication test_main_reclaims_a_grant_whose_branch_owner_exited diff --git a/tests/fm-watch-arm.test.sh b/tests/fm-watch-arm.test.sh index bafe83ef0d1..49d350c538f 100755 --- a/tests/fm-watch-arm.test.sh +++ b/tests/fm-watch-arm.test.sh @@ -288,16 +288,10 @@ test_rearm_resurfaces_durable_queue_and_remote_open_decision() { append_wake "$state" check startup-network 'check: startup-network' start_rearm_arm "$home" "$state" "$fakebin" "$armout" - sleep 0.25 - if is_live_non_zombie "$ARM_PID"; then - # End the fixture through an ordinary actionable status transition so this - # failing pre-fix path leaves no child behind. - printf 'done: fixture cleanup\n' > "$state/cleanup.status" - wait_for_exit "$ARM_PID" 80 || true - fail "re-arm stayed live instead of surfacing durable wakes and the still-open remote decision" - fi - wait "$ARM_PID" + wait_for_exit "$ARM_PID" 80 status=$? + [ "$status" -ne 124 ] \ + || fail "re-arm stayed live instead of surfacing durable wakes and the still-open remote decision" expect_code 0 "$status" "re-arm re-surface wake must close successfully" grep -F 'check: rearm-resurface' "$armout" >/dev/null \ || fail "re-arm did not report the durable recovery wake: $(cat "$armout")" diff --git a/tests/fm-watch-triage.test.sh b/tests/fm-watch-triage.test.sh index 2b8d2933c4b..029b4a491f9 100755 --- a/tests/fm-watch-triage.test.sh +++ b/tests/fm-watch-triage.test.sh @@ -1522,6 +1522,117 @@ test_actionable_signal_surfaced() { pass "captain-relevant signal is surfaced (queue + exit) and marked surfaced" } +# A needs-decision status append surfaced through this actionable signal path +# must skip the Pi supervision branch and reach main directly +# (docs/pi-supervision-branch.md "Autonomy"). The row still +# queues as an ordinary signal-kind wake - fm-branch-dispatch.ts's +# scopeForUnreadWake tells it apart from a routine signal by this payload +# marker, not by kind. +test_needs_decision_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid + dir=$(make_case needs-decision-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: setup\nneeds-decision: pick A or B\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable needs-decision signal" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a needs-decision signal row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion" +} + +# A needs-decision whose key transition was rejected by the reserved-key +# vocabulary is reported as a "reconciliation-required: " wrapped event +# (fm-classify-lib.sh's status_span_first_actionable_record), but it is still a +# needs-decision signal that this path routes directly to main - the payload +# marker must not be fooled by that wrapper. +test_needs_decision_reconciliation_required_still_marked() { + local dir state fakebin out status_file pid + dir=$(make_case needs-decision-reconciliation); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'needs-decision [key=pending-reply-x]: unrelated request\nworking: awaiting reconciliation\n' \ + > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for a rejected-reserved-key needs-decision" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a reconciliation-required needs-decision row was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a reconciliation-required needs-decision row's queued payload is still marked needs-decision:" +} + +# A captain-held declaration is itself actionable. Positive evidence that the +# crew is still working must not absorb the signal before its main-only marker +# can be delivered. +test_captain_held_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid + dir=$(make_case captain-held-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'captain-held [key=route]: awaiting the captain\n' > "$status_file" + export FM_FAKE_CREW_STATE='state: working · source: run-step · still wrapping up' + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher absorbed a captain-held signal while the crew was still working" + grep -F "signal: $status_file" "$out" >/dev/null \ + || fail "a captain-held signal changed its wake reason: $(cat "$out")" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a captain-held signal was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a captain-held signal stays actionable while the crew is still working" +} + +test_pending_reply_escalation_signal_payload_marked_for_branch_exclusion() { + local dir state fakebin out status_file pid corr + dir=$(make_case pending-reply-escalation-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + corr=0123456789abcdef + printf 'blocked [key=pending-reply-%s]: pending-reply-missed: task=task pending-reply-id=%s request=finish report\n' \ + "$corr" "$corr" > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for a pending-reply escalation" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + || fail "a pending-reply escalation was not payload-marked for branch exclusion: $(cat "$state/.wake-queue")" + pass "a pending-reply second-mate escalation is marked for main-only routing" +} + +test_ordinary_blocked_signal_payload_remains_branch_eligible() { + local dir state fakebin out status_file pid + dir=$(make_case ordinary-blocked-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'blocked [key=dependency]: waiting for an upstream release\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an ordinary blocked event" + grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ + || fail "an ordinary blocked event lost branch-eligible routing: $(cat "$state/.wake-queue")" + if grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null; then + fail "an ordinary blocked event was marked as a second-mate escalation" + fi + pass "an ordinary blocked event remains branch-eligible" +} + +# A routine (non-needs-decision) captain-relevant event must keep its ordinary +# payload: only a genuine needs-decision gets the exclusion marker. +test_routine_signal_payload_not_marked_needs_decision() { + local dir state fakebin out status_file pid + dir=$(make_case routine-signal-payload); state="$dir/state"; fakebin="$dir/fakebin" + out="$dir/watch.out" + status_file="$state/task.status" + printf 'working: setup\ndone: migration complete ; needs-decision: documented in follow-up\n' > "$status_file" + watch_bg "$state" "$fakebin" "$out" + pid=$! + wait_for_exit "$pid" 100 || fail "watcher did not exit for an actionable done signal" + grep -F "$(printf 'signal\ttask.status\tneeds-decision:')" "$state/.wake-queue" >/dev/null \ + && fail "a routine done signal was incorrectly payload-marked needs-decision: $(cat "$state/.wake-queue")" + grep -F "$(printf 'signal\ttask.status\tsignal:')" "$state/.wake-queue" >/dev/null \ + || fail "a routine signal lost its ordinary payload: $(cat "$state/.wake-queue")" + pass "a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked" +} + # The reported bug, end to end through a real watcher: a crew reports something # the captain must act on and then keeps appending routine progress, which is # ordinary while the watcher lingers its signal grace window to coalesce a status @@ -2246,6 +2357,274 @@ test_live_declared_wait_churn_honors_the_resurface_throttle() { pass "a parked live worker surfaces once, absorbs pane churn for the whole re-surface window, then re-surfaces when it elapses" } +# --- work the captain is already holding: pane churn must not re-alarm ------- +# The other record of a legitimate wait. The declared-wait bound above reads the +# status LINE, and a delivered task's line stays `done: PR ...` while the wait +# itself lives in the BACKLOG, written there by bin/fm-captain-hold.sh. No line +# predicate can see that record, so both stale alarms - the captain-relevant one +# and the inconclusive one - re-fired on every new pane hash for as long as the +# captain was deciding, which is the 2026-09 loop observed on delivered work +# awaiting their merge word. +# Pinned here, in both directions: while the call stands the first sight still +# alarms, further sights of the SAME call and status-log state are absorbed, and +# a new pane hash after the window's end alarms once more; and the identical +# fixture WITHOUT the hold keeps alarming on every hash, because a bound that +# swallowed an unheld delivery or blocker would be worse than the churn it removes. +# +# The backlog is real rather than a fixture file: bin/fm-captain-hold.sh is the +# only writer of a hold and tasks-axi the only reader, so a hand-written row +# would pin this test's idea of a hold instead of the one the watcher consults. +# +# Cost: every case below drives churn through ONE watcher process rather than +# relaunching per pane change. Watcher startup dominates a round here, and an +# absorbing watcher stays in its poll loop across churn in production anyway, so +# the cheaper shape is also the more faithful one. + +# The window key every hold fixture uses, derived the way fm-watch.sh derives it. +hold_key() { + printf '%s' test:fm-held-merge | tr ':/.' '___' +} + +# bin/fm-captain-hold.sh against a hold fixture's own home. +run_hold() { # <dir> <args...> + local dir=$1 + shift + FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" FM_DATA_OVERRIDE="$dir/data" \ + FM_CONFIG_OVERRIDE="$dir/config" "$ROOT/bin/fm-captain-hold.sh" "$@" >/dev/null 2>&1 +} + +make_hold_home() { # <name> <status-line> <hold|nohold> + local name=$1 line=$2 hold=$3 dir state + dir=$(make_case "$name"); state="$dir/state" + mkdir -p "$dir/data" "$dir/config" + cp "$ROOT/.tasks.toml" "$dir/.tasks.toml" || return 1 + printf '## In flight\n\n## Queued\n\n## Done\n' > "$dir/data/backlog.md" + (cd "$dir" && tasks-axi add held-merge 'delivered work' --file data/backlog.md) >/dev/null 2>&1 \ + || return 1 + if [ "$hold" = hold ]; then + run_hold "$dir" hold held-merge --reason 'awaiting the captain on the merge' || return 1 + fi + printf 'window=test:fm-held-merge\nkind=ship\nharness=grok\nbackend=tmux\n' \ + > "$state/held-merge.meta" + printf '%s\n' "$line" > "$state/held-merge.status" + printf '%s' "$(seen_sig "$state/held-merge.status")" > "$state/.seen-held-merge_status" + printf '%s\n' "$dir" +} + +# Launch one watcher against a hold fixture, armed the way parked_watch_round +# arms one, plus the home the backlog read resolves against. The crew reads +# stopped: a delivered worker's agent has exited, and that is the population +# whose alarm the call must bound. The pid lands in HOLD_WATCH_PID rather than on +# stdout: a command substitution would background the watcher inside a subshell, +# leaving the caller unable to wait on or reap its own watcher. +HOLD_WATCH_PID= +hold_watch_launch() { # <dir> <out> <capture> + local dir=$1 out=$2 capture=$3 + PATH="$dir/fakebin:$PATH" FM_FAKE_TMUX_WINDOW=test:fm-held-merge \ + FM_FAKE_TMUX_CAPTURE="$capture" FM_FAKE_TMUX_CURRENT_COMMAND=zsh \ + FM_FAKE_CREW_STATE='state: stopped · source: pane · bare shell' \ + FM_WATCH_HANDLING_SUCCESSOR=1 \ + FM_HOME="$dir" FM_DATA_OVERRIDE="$dir/data" FM_CONFIG_OVERRIDE="$dir/config" \ + FM_STATE_OVERRIDE="$dir/state" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" \ + FM_PAUSE_RESURFACE_SECS="${FM_HOLD_PAUSE_RESURFACE_SECS:-999}" FM_POLL=1 FM_SIGNAL_GRACE=1 \ + FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" >> "$out" 2>&1 & + HOLD_WATCH_PID=$! +} + +# One sighting that must surface and exit the cycle. +hold_watch_surface() { # <dir> <out> <capture> <pane-text> + local dir=$1 out=$2 capture=$3 text=$4 + printf '%s\n' "$text" > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + wait_for_exit "$HOLD_WATCH_PID" 100 || { reap "$HOLD_WATCH_PID"; return 1; } + return 0 +} + +# <count> successive pane changes driven through ONE watcher, each given three +# poll cycles: one to see the new hash, one to count it stable and classify, one +# to prove the classification held. The watcher must stay in the loop throughout. +hold_watch_churn() { # <dir> <out> <capture> <label> <count> + local dir=$1 out=$2 capture=$3 label=$4 count=$5 i=1 c + local state="$dir/state" + printf '%s 0\n' "$label" > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + while [ "$i" -le "$count" ]; do + printf '%s %s\n' "$label" "$i" > "$capture" + c=0 + while [ "$c" -lt 3 ]; do + wait_poll_cycle "$state" "$HOLD_WATCH_PID" 300 \ + || { reap "$HOLD_WATCH_PID"; return 1; } + c=$((c + 1)) + done + i=$((i + 1)) + done + reap "$HOLD_WATCH_PID" + return 0 +} + +hold_stale_wakes() { # <state> + awk -F '\t' '$3 == "stale" && $4 == "test:fm-held-merge" { n++ } END { print n + 0 }' \ + "$1/.wake-queue" 2>/dev/null || echo 0 +} + +# Both status lines a held task really carries: the delivery that routes through +# the captain-relevant stale branch, and a worker line that routes through the +# inconclusive one. The hold is invisible to the status line in both, so both +# branches had the same blindness and both are covered. +test_open_captain_call_bounds_stale_churn() { + local spec name line dir state out capture throttle wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (captain-hold stale bound)"; return 0; } + for spec in \ + 'held-delivery|done: PR https://example.invalid/pull/1 checks green' \ + 'held-worker-line|working: still tidying the branch' + do + name=${spec%%|*}; line=${spec#*|} + dir=$(make_hold_home "$name" "$line" hold) \ + || fail "[$name] could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + throttle="$state/.paused-resurfaced-$(hold_key)" + + # First sight still alarms: the call bounds repetition, never the first look. + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 1s' \ + || fail "[$name] first sight of held work did not surface" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] || fail "[$name] first sight produced $wakes wakes instead of one" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge the first surface" + + # The pane churns while the SAME call stands. Every one of these alarmed. + hold_watch_churn "$dir" "$out" "$capture" 'idle, tick' 2 \ + || fail "[$name] watcher exited during pane churn instead of supervising through it" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 0 ] \ + || fail "[$name] pane churn re-alarmed held work $wakes time(s) inside the re-surface window" + + # After the window ends, the next new pane hash re-surfaces held work exactly + # once, so a forgotten call on a churning pane cannot hide behind the bound. + [ -e "$throttle" ] || fail "[$name] the absorbed churn recorded no re-surface cadence to elapse" + set_mtime "$(( $(date +%s) - 5000 ))" "$throttle" + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 9s' \ + || fail "[$name] held work did not re-surface once its re-surface window elapsed" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "[$name] elapsed re-surface window produced $wakes wakes instead of one" + done + pass "work under an open captain call surfaces once, absorbs pane churn, then re-surfaces when the window elapses" +} + + + +# The other half of the same bound, and the one that decides whether widening the +# wait was safe: the identical fixtures with NO hold must keep alarming on every +# new hash, on both branches. +test_stale_churn_without_a_captain_call_still_alarms() { + local spec name line dir state out capture round wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (unheld stale alarm)"; return 0; } + for spec in \ + 'unheld-delivery|done: PR https://example.invalid/pull/1 checks green' \ + 'unheld-blocker|blocked: cannot reach the release host' \ + 'unheld-worker-line|working: still tidying the branch' + do + name=${spec%%|*}; line=${spec#*|} + dir=$(make_hold_home "$name" "$line" nohold) \ + || fail "[$name] could not build an unheld backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + round=1 + while [ "$round" -le 2 ]; do + hold_watch_surface "$dir" "$out" "$capture" "idle, elapsed ${round}s" \ + || fail "[$name] an unheld stale window stopped alarming on round $round" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "[$name] round $round produced $wakes wakes instead of one" + ack_stopped_cycle "$state" || fail "[$name] could not acknowledge round $round" + round=$((round + 1)) + done + done + pass "a stale window with no open captain call keeps alarming on every new hash" +} + + +# The cadence marker may never outlive the wake it claims to record. Recording it +# before publishing the durable wake turned a delayed alarm into a lost one: the +# append fails, the watcher exits with nothing queued, and the next sighting +# reads that fresh marker and absorbs the retry. An unwritable queue is the real +# failure, so it is the one this drives. +test_failed_wake_append_does_not_arm_the_captain_hold_throttle() { + local dir state out capture wakes rc + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (failed wake append)"; return 0; } + dir=$(make_hold_home append-failure 'done: PR https://example.invalid/pull/1 checks green' hold) \ + || fail "could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + + # A directory where the queue file belongs: every append fails, whatever the + # caller does, so the watcher cannot publish the wake it just decided to send. + # Its exit code is read directly here because a refusing watcher exits NON-zero, + # which is the correct outcome and not the "surfaced" one hold_watch_surface means. + rm -f "$state/.wake-queue" + mkdir -p "$state/.wake-queue" + printf 'idle, elapsed 1s\n' > "$capture" + hold_watch_launch "$dir" "$out" "$capture" + wait_for_exit "$HOLD_WATCH_PID" 100 + rc=$? + rmdir "$state/.wake-queue" + [ "$rc" -ne 124 ] || fail "the watcher did not exit when its durable queue could not be written" + [ "$rc" -ne 0 ] || fail "the watcher reported success despite an unwritable durable queue" + [ -e "$state/.paused-resurfaced-$(hold_key)" ] \ + && fail "a wake that never reached the durable queue still armed the re-surface throttle" + + # The retry must alarm: nothing was ever delivered, so nothing may be absorbed. + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 2s' \ + || fail "the retry after a failed wake append was absorbed instead of alarming" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "the retry after a failed wake append produced $wakes wakes instead of one" + pass "a wake that never reached the durable queue arms no re-surface throttle" +} + +# The task id is not the captain call. A task can be answered with `--release` +# and held again as a genuinely different call with NO status append, and binding +# the throttle to the status-log signature alone let the second call inherit the +# first one's silence and absorbed its first sight. That is the one alarm this +# bound must never swallow: a delivery announced twice is noise, but a decision +# waiting on the captain that is never surfaced is invisible. +# Measured at base c499f84 this fixture alarms on every sighting, so the +# suppression was introduced by the bound itself rather than pre-existing. +test_reheld_captain_call_starts_its_own_resurface_window() { + local dir state out capture wakes + command -v tasks-axi >/dev/null 2>&1 \ + || { echo "skip: tasks-axi not found (re-held captain call)"; return 0; } + dir=$(make_hold_home reheld-call 'done: PR https://example.invalid/pull/1 checks green' hold) \ + || fail "could not build a captain-held backlog fixture" + state="$dir/state"; out="$dir/watch.out"; capture="$dir/pane.txt" + + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 1s' \ + || fail "first sight of the first captain call did not surface" + ack_stopped_cycle "$state" || fail "could not acknowledge the first call's surface" + hold_watch_churn "$dir" "$out" "$capture" 'idle, tick' 1 \ + || fail "the first call's churn was not absorbed" + [ "$(hold_stale_wakes "$state")" -eq 0 ] \ + || fail "the first call's churn re-alarmed inside its own window" + + # Answer and release, then re-hold: a second, distinct captain call on the same + # task id, with no status append, so the status signature cannot tell them apart. + printf 'go ahead\n' > "$dir/decision.txt" + run_hold "$dir" answer held-merge --decision-file "$dir/decision.txt" --release \ + || fail "could not record the captain's answer" + run_hold "$dir" hold held-merge --reason 'awaiting the captain a second time' \ + || fail "could not re-hold the task as a second captain call" + + hold_watch_surface "$dir" "$out" "$capture" 'idle, elapsed 3s' \ + || fail "the second captain call inherited the first call's silence" + wakes=$(hold_stale_wakes "$state") + [ "$wakes" -eq 1 ] \ + || fail "the second captain call produced $wakes first wakes instead of one" + pass "a released-then-re-held task is a distinct captain call whose first sight still alarms" +} + + + test_secondmate_paused_resurfaces_in_normal_mode() { local dir state fakebin out capture_file statusf window key pane_hash sig pid back dir=$(make_case secondmate-paused-resurface); state="$dir/state"; fakebin="$dir/fakebin" @@ -4032,6 +4411,12 @@ test_working_note_not_working_surfaced test_secondmate_status_note_surfaced_despite_busy_agent test_self_announced_close_does_not_rewake_but_next_note_does test_actionable_signal_surfaced +test_needs_decision_signal_payload_marked_for_branch_exclusion +test_needs_decision_reconciliation_required_still_marked +test_captain_held_signal_payload_marked_for_branch_exclusion +test_pending_reply_escalation_signal_payload_marked_for_branch_exclusion +test_ordinary_blocked_signal_payload_remains_branch_eligible +test_routine_signal_payload_not_marked_needs_decision test_actionable_signal_survives_a_later_routine_append test_release_completion_survives_a_later_routine_append test_routine_appends_after_a_classified_event_stay_absorbed @@ -4056,6 +4441,10 @@ test_nonterminal_stale_paused_absorbed_then_resurfaced test_exited_declared_pause_is_bounded_but_live_gate_surfaces test_absorbed_replacement_wait_does_not_inherit_the_old_throttle test_live_declared_wait_churn_honors_the_resurface_throttle +test_open_captain_call_bounds_stale_churn +test_stale_churn_without_a_captain_call_still_alarms +test_failed_wake_append_does_not_arm_the_captain_hold_throttle +test_reheld_captain_call_starts_its_own_resurface_window test_secondmate_paused_resurfaces_in_normal_mode test_secondmate_captain_held_resurfaces_in_normal_mode test_secondmate_nonpaused_stale_remains_suppressed diff --git a/tests/lib.sh b/tests/lib.sh index fb269ef70ca..d042f96dd25 100644 --- a/tests/lib.sh +++ b/tests/lib.sh @@ -26,6 +26,13 @@ if [ -n "${FM_TEST_LIB_SOURCED:-}" ]; then fi FM_TEST_LIB_SOURCED=1 +# Pin the fixture umask. Firstmate's state-root and process-event contracts +# refuse group- or world-writable state directories, and a permissive ambient +# umask (e.g. 0002) makes every `mkdir state` fixture fail that contract before +# the behavior under test can even run. 022 is the conventional default this +# suite's fixtures were written against. +umask 022 + # Exempt firstmate's own test suite from the gate-lifecycle refusal # (bin/fm-gate-refuse-lib.sh). The no-mistakes gate runs this suite FROM a gate # worktree - the exact environment that guard refuses - so without this every @@ -35,6 +42,12 @@ FM_TEST_LIB_SOURCED=1 # strips this to verify real refusal. export FM_GATE_REFUSE_BYPASS=1 +# Clear the task-worker marker bin/fm-spawn.sh exports into ship and scout +# panes. This suite builds git-init fixture repositories whose primary checkout +# it runs a copied bin/fm-test-run.sh in, and that runner refuses the primary +# under the marker. A case that verifies the refusal sets FM_TASK_ID itself. +unset FM_TASK_ID + # Resolve the repo root from this library's own location. Consumed by sourcing # test files, not by this library, so it reads as "unused" here. # shellcheck disable=SC2034 @@ -51,6 +64,45 @@ pass() { printf 'ok - %s\n' "$1" } +# fm_test_run_cases <case ...> +# FM_TEST_ONLY selects one registered case; FM_TEST_LIST_CASES=1 lists them. +# With neither set, preserve the declared full-suite order. Consume selectors +# here so a selected case cannot accidentally filter its child fixtures. +fm_test_run_cases() { + local selected=${FM_TEST_ONLY:-} list=${FM_TEST_LIST_CASES:-0} name found=0 rc + [ "$#" -gt 0 ] || fail "fm_test_run_cases: no cases registered" + case "$list" in + 0|1) ;; + *) fail "fm_test_run_cases: FM_TEST_LIST_CASES must be 0 or 1" ;; + esac + [ -z "$selected" ] || [ "$list" != 1 ] \ + || fail "fm_test_run_cases: select a case or list cases, not both" + for name in "$@"; do + case "$name" in + test_*) ;; + *) fail "fm_test_run_cases: invalid case name: $name" ;; + esac + case "$name" in + *[!a-zA-Z0-9_]*) fail "fm_test_run_cases: invalid case name: $name" ;; + esac + declare -F "$name" >/dev/null || fail "fm_test_run_cases: case is not defined: $name" + [ "$name" != "$selected" ] || found=1 + done + [ -z "$selected" ] || [ "$found" -eq 1 ] \ + || fail "fm_test_run_cases: unknown case: $selected" + unset FM_TEST_ONLY FM_TEST_LIST_CASES + if [ "$list" = 1 ]; then + printf '%s\n' "$@" + return 0 + fi + for name in "$@"; do + [ -z "$selected" ] || [ "$name" = "$selected" ] || continue + "$name" + rc=$? + [ "$rc" -eq 0 ] || exit "$rc" + done +} + fm_test_make_symlink() { case "$(uname -s 2>/dev/null)" in MINGW*|MSYS*) MSYS=winsymlinks:nativestrict ln -s "$@" ;; @@ -90,8 +142,59 @@ FM_TEST_OWNER_IDENTITY=$(fm_test_pid_identity "$$") || { return 1 } +# --- process-event runner reaping ------------------------------------------- +# +# A process-event runner is detached into its own process group and reparents to +# init, so removing a fixture directory does not stop one: only sweeping the home +# that owns it does. Registration goes through a `$$`-keyed registry file for the +# same reason the temp roots do - a fixture home is almost always built inside a +# command substitution (`home=$(make_home x)`), and an array append there never +# reaches the caller, so a suite that tracked its homes in a shell array was +# silently tracking nothing and left every runner it started behind. +# +# The sweep is scoped to the exact home (and its claim root when the suite uses a +# private one). It never matches on a script or process name, which would reach +# into another home's live runners. + +FM_TEST_PROCEVENT_REGISTRY=$(mktemp "${TMPDIR:-/tmp}/.fm-test-procevent.$$.XXXXXX") || return 1 + +fm_test_track_procevent_home() { # <home> [claim-root] + [ -n "${1:-}" ] || return 1 + printf '%s\t%s\n' "$1" "${2-}" >> "$FM_TEST_PROCEVENT_REGISTRY" +} + +fm_test_reap_procevent_homes() { + local home claim_root seen=$'\n' + [ -f "$FM_TEST_PROCEVENT_REGISTRY" ] || return 0 + while IFS=$'\t' read -r home claim_root; do + [ -n "$home" ] || continue + case "$seen" in *$'\n'"$home"$'\n'*) continue ;; esac + seen+="$home"$'\n' + [ -d "$home/state/procevent" ] || continue + if [ -n "$claim_root" ]; then + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" FM_PROCEVENT_CLAIM_ROOT="$claim_root" \ + "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true + else + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + "$ROOT/bin/fm-procevent.sh" sweep-home >/dev/null 2>&1 || true + fi + done < "$FM_TEST_PROCEVENT_REGISTRY" + rm -f "$FM_TEST_PROCEVENT_REGISTRY" +} + +# Ceiling on how long a fixture's blocking stub may keep polling. A stub that +# waits for a trigger file by re-running `sleep` is a high-frequency source of +# process spawns, and one that outlives its test - because the test was killed +# before any cleanup ran - is what turned leftover fixtures into a host-wide +# process storm. Every blocking stub this suite writes stops itself at this +# bound, so an escaped one is bounded in duration and cost on its own, before +# its owner's guard reaps it. +FM_TEST_STUB_MAX_BLOCK_SECONDS=${FM_TEST_STUB_MAX_BLOCK_SECONDS:-120} +export FM_TEST_STUB_MAX_BLOCK_SECONDS + fm_test_cleanup() { local d + fm_test_reap_procevent_homes for d in "${FM_TEST_CLEANUP_DIRS[@]:-}"; do [ -n "$d" ] && rm -rf "$d" done @@ -120,6 +223,8 @@ fm_test_tmproot() { trap fm_test_cleanup EXIT trap 'fm_test_cleanup; exit 130' INT trap 'fm_test_cleanup; exit 143' TERM +trap 'fm_test_cleanup; exit 129' HUP +trap 'fm_test_cleanup; exit 131' QUIT # fm_test_reap_orphans: best-effort sweep for fixture roots left behind by a # prior run that was killed hard enough to skip the traps above (e.g. a @@ -163,6 +268,101 @@ if [ "${FM_TEST_SKIP_ORPHAN_REAP:-0}" != 1 ]; then fm_test_reap_orphans fi +# --- live-capability gate --------------------------------------------------- +# +# fm_live_gate <policy> <vars> [tool ...] +# +# The single gate every live-harness guard opens with, so "can this host run +# this guard for real, and should it?" is decided in one place instead of in +# two dozen hand-rolled env checks. It returns 0 when the guard should run, and +# otherwise ends the script with one runner-readable line: +# +# skip: live: <tool> absent this host cannot run the guard +# skip: live: disabled by <VAR>=0 an explicit local opt-out +# skip: live: opt-in; set <VAR>=1 to run a guard that spends model tokens +# +# <policy> is default-on for a guard that spends no model tokens, so it runs +# wherever its tools are installed - notably on the machine the product and its +# validation actually run on - and opt-in for a guard that submits prompts, +# which stays deliberate. <vars> is the guard's own control variable, or a +# comma-separated list when a guard has more than one entry point. +# +# Setting any of those variables to 1 (or FM_LIVE=1, for every guard at once) +# both turns the guard on and makes an absent tool a hard failure rather than a +# skip, which is how "run it after a harness upgrade" keeps proving the guard +# actually ran. Setting one to 0 (or FM_LIVE=0) turns it off; a guard's own +# variable wins over FM_LIVE. +# +# Sourcing this library also exports FM_GATE_REFUSE_BYPASS=1, which is what +# lets a live guard drive the real fm-spawn/fm-send/fm-teardown from inside a +# no-mistakes gate worktree instead of being refused by +# bin/fm-gate-refuse-lib.sh. + +fm_live_gate() { + local policy=$1 vars=$2 + shift 2 + local var value rest primary requested=0 disabled_by='' tool + local -a var_list=() + + case "$policy" in + default-on | opt-in) ;; + *) fail "fm_live_gate: unknown policy '$policy' (expected default-on or opt-in)" ;; + esac + + rest=$vars + while [ -n "$rest" ]; do + var=${rest%%,*} + if [ "$var" = "$rest" ]; then + rest='' + else + rest=${rest#*,} + fi + [ -n "$var" ] && var_list+=("$var") + done + [ "${#var_list[@]}" -gt 0 ] || fail "fm_live_gate: at least one control variable is required" + primary=${var_list[0]} + + for var in "${var_list[@]}"; do + value=${!var:-} + case "$value" in + 1) requested=1 ;; + 0) [ -n "$disabled_by" ] || disabled_by=$var ;; + esac + done + + if [ "$requested" -eq 0 ]; then + if [ -n "$disabled_by" ]; then + printf 'skip: live: disabled by %s=0\n' "$disabled_by" + exit 0 + fi + case "${FM_LIVE:-}" in + 0) + printf 'skip: live: disabled by FM_LIVE=0\n' + exit 0 + ;; + 1) requested=1 ;; + *) + if [ "$policy" = opt-in ]; then + printf 'skip: live: opt-in; set %s=1 to run\n' "$primary" + exit 0 + fi + ;; + esac + fi + + for tool in "$@"; do + command -v "$tool" >/dev/null 2>&1 && continue + if [ "$requested" -eq 1 ]; then + printf 'not ok - %s was requested but %s is not installed\n' "$primary" "$tool" >&2 + exit 1 + fi + printf 'skip: live: %s absent\n' "$tool" + exit 0 + done + + return 0 +} + # --- fakebin / PATH shims --------------------------------------------------- # # fm_fakebin <dir> creates <dir>/fakebin and echoes it; prepend it to PATH to @@ -361,3 +561,43 @@ assert_absent() { assert_present() { [ -e "$1" ] || fail "$2" } + +# fm_test_base_path_sans <base_path> <tool...>: returns the path to a single +# curated directory that resolves every tool <base_path> would have resolved, +# except the named ones. Some hosts have real system binaries (node, orca, +# ...) sitting in BASE_PATH; a fixture that simulates a tool as missing by +# omitting it from fakebin still falls through to that host binary via +# BASE_PATH, silently defeating the simulation. Dropping whole directories +# out of BASE_PATH is not a safe fix: on a usr-merged host /bin, /sbin, and +# /usr/sbin are symlinks that collapse to the same directory as /usr/bin, so +# dropping any one of them because it resolves the excluded tool drops every +# other tool a test still needs (git, awk, sed, ...) too. Building a curated +# directory instead hides only the named tool(s). Use only at the specific +# assertions that simulate a tool as absent - every other case keeps using +# bare BASE_PATH. +fm_test_base_path_sans() { + local base_path=$1 dir src entry name tool skip + shift + local tools=("$@") + dir=$(fm_test_tmproot fm-base-path-sans) || return 1 + local dirs + IFS=: read -ra dirs <<< "$base_path" + for src in "${dirs[@]}"; do + [ -d "$src" ] || continue + for entry in "$src"/*; do + [ -e "$entry" ] || [ -L "$entry" ] || continue + name=${entry##*/} + [ -e "$dir/$name" ] && continue + skip=0 + for tool in "${tools[@]}"; do + if [ "$name" = "$tool" ]; then + skip=1 + break + fi + done + [ "$skip" -eq 1 ] && continue + ln -s "$entry" "$dir/$name" 2>/dev/null || true + done + done + printf '%s\n' "$dir" +} diff --git a/tests/secondmate-helpers.sh b/tests/secondmate-helpers.sh index e78881872c9..365faccd4c2 100644 --- a/tests/secondmate-helpers.sh +++ b/tests/secondmate-helpers.sh @@ -31,9 +31,26 @@ case "${1:-}" in exit 0 ;; list-windows) - if [ -n "${FM_FAKE_TMUX_WINDOW:-}" ]; then - printf '%s\n' "$FM_FAKE_TMUX_WINDOW" - fi + session= + prev= + for arg in "$@"; do + if [ "$prev" = -t ]; then session=$arg; break; fi + prev=$arg + done + while IFS= read -r recorded; do + [ -n "$recorded" ] || continue + if [ -z "$session" ]; then + printf '%s\n' "$recorded" + continue + fi + case "$recorded" in + "$session":*) printf '%s\n' "${recorded#*:}" ;; + *:*) ;; + *) printf '%s\n' "$recorded" ;; + esac + done <<EOF +${FM_FAKE_TMUX_WINDOW:-} +EOF exit 0 ;; display-message)