diff --git a/.agents/skills/afk/SKILL.md b/.agents/skills/afk/SKILL.md index aba6e3fb00c..8df771825fa 100644 --- a/.agents/skills/afk/SKILL.md +++ b/.agents/skills/afk/SKILL.md @@ -160,6 +160,8 @@ Classify each wake this way: Other signals with no captain-relevant status -> self-handle. - `signal` or `stale` for a declared `paused:` external wait -> self-handle and track the pause rather than a wedge. If it remains declared and idle past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one awaiting-external recheck and resets the pause window. + The recheck exists to re-ask a wait that can change without the captain, so it is skipped for a wait the backlog records as captain-gated (`hold_kind: captain`): that one clears only when the captain acts, and the captain acting is already the away-mode exit, which runs the full return catch-up. + Suppression is a cadence decision only - the wait is still tracked, still reset each window, and still as visible as before in the backlog digest, the fleet view, and the return catch-up - and any kind that cannot be established is rechecked normally rather than dropped. - `check` -> always escalate. Check scripts print only when firstmate should wake. - `stale` with a terminal status or bare legacy captain-relevant line -> escalate. Nonterminal progress remains transient even when its prose contains a legacy free-text token or its seen-status marker already matches, so record a marker and self-handle. diff --git a/.agents/skills/bootstrap-diagnostics/SKILL.md b/.agents/skills/bootstrap-diagnostics/SKILL.md index 477980b8df1..7b001179f34 100644 --- a/.agents/skills/bootstrap-diagnostics/SKILL.md +++ b/.agents/skills/bootstrap-diagnostics/SKILL.md @@ -2,7 +2,7 @@ name: bootstrap-diagnostics description: >- Agent-only handling playbook for session-start bootstrap diagnostics. - Use whenever the session-start digest's bootstrap section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, FLEET_SYNC, PR_CHECK_MIGRATION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, NUDGE_SECONDMATES, or FMX - or when a standalone bin/fm-bootstrap.sh run prints one of those lines. + Use whenever the session-start digest's bootstrap section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, MODEL_REGISTRY, MODEL_PRICE, MODEL_VERIFY, ADMISSION_CONTROL, WAKE_LEDGER, FLEET_SYNC, PR_CHECK_MIGRATION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, or FMX - or when a standalone bin/fm-bootstrap.sh run prints one of those lines. A silent bootstrap section, or a BOOTSTRAP_INFO fact, means no skill load. user-invocable: false metadata: @@ -19,7 +19,8 @@ When any diagnostic needs captain attention, report the plain consequence and re - `MISSING: (install: )` - list the missing tools to the captain with a one-line purpose each plus the printed install commands, wait for consent (one approval may cover the list), then run `bin/fm-bootstrap.sh install `. For `treehouse`, this also covers an installed version whose `treehouse get` lacks `--lease`; treat it as an upgrade request. For `no-mistakes`, this also covers an installed version older than 1.31.2, because crewmate validation briefs delegate gate mechanics to no-mistakes' version-matched guidance. - For `tasks-axi`, this also covers an installed build that fails the compatibility probe (`docs/configuration.md` "Backlog backend" owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report. + For `gh-axi`, this also covers an installed version below the bootstrap-owned floor; treat it as an upgrade request so non-interactive PR merges keep a working bare `--squash` shorthand. + For `tasks-axi`, this also covers an installed build that fails the compatibility probe (`bin/fm-tasks-axi-lib.sh` owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report. For `quota-axi`, bootstrap requires it because firstmate reads its current output directly before resolving every crew-dispatch profile array; without it, report the missing requirement and do not choose around an unexamined candidate. - `MISSING_MANUAL: (instructions: )` - tell the captain why the tool is required and give them the printed instructions URL, but do not pass the tool to `bin/fm-bootstrap.sh install`; wait for the captain to complete the manual installation, then rerun session start to confirm the dependency is present. - `BACKEND_INVALID: (known: )` - the resolved runtime backend has no verified dependency or lifecycle contract, so do not dispatch work until the invalid `FM_BACKEND` or `config/backend` value is corrected to one of the listed backends. @@ -27,8 +28,28 @@ When any diagnostic needs captain attention, report the plain consequence and re - `TANGLE: ` - the primary checkout is stranded on a feature branch instead of its default branch; `AGENTS.md` section 8 explains why this guard exists and what it protects. The work is safe on that branch ref; restore the primary to its default branch with the printed `git -C checkout `, then re-validate that branch in a proper worktree. This is the only sanctioned firstmate-initiated git write to the primary, and it is a non-destructive branch switch that strands nothing. -- `STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - ` - the visible startup-memory budget is not a safe one-line positive decimal file; do not infer the default or propagate it. Correct the local primary file, then rerun session start so the normal convergence path can deliver the validated value to secondmate homes. +- `STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - ` - the visible startup-memory budget is not a safe one-line positive decimal file; do not infer the default or propagate it. + Correct the local primary file, then rerun session start so the normal convergence path can deliver the validated value to secondmate homes. - `CREW_DISPATCH: invalid config/crew-dispatch.json - ` - the optional dispatch profile file exists but failed low-cost bootstrap validation; stop profile-based dispatch, report the actionable error, and require correction of the malformed schema, unverified harness name, or invalid harness/effort pair rather than falling back around it or selecting a bad profile. +- `MODEL_REGISTRY: invalid config/models.json - ` - the model registry exists but failed schema validation, so every provider-prefixed model is now refused at spawn until it is corrected. + Fix the registry; never delete it to clear the error, because deleting it silently disables zero-budget enforcement rather than restoring it. +- `MODEL_REGISTRY: ` - the dispatch config and the registry disagree: the model is unregistered, carries a non-approved status, or has no current live-probe record. + This is the check that catches a bad model before any worker is launched against it, so correct the dispatch rule or complete the model's admission; do not weaken the check to make the line go away. +- `MODEL_REGISTRY: no config/models.json, ...` - routed provider models exist but nothing enforces the zero-budget rule for them. + This is a standing gap rather than a failure, and it stays inert by design; raise it with the captain rather than treating it as a startup blocker. +- `MODEL_PRICE: ... no longer zero` - an allowlisted model is no longer free at its provider, which is the exposure a name-only allowlist cannot see. + Suspend that route immediately, then re-verify its cost class before it is routed to again. +- `MODEL_PRICE: price drifted ...` or `... catalogue source is unreadable` - re-verify the cost class and update the recorded price, or repair the declared catalogue path so the check stops being blind. +- `MODEL_VERIFY: ...` - a live probe failed. Load `model-onboarding` and read the shape before reacting: a provider refusal means this account can never use that model and it must leave routing, while a local failure is a configuration error on this machine and not a provider fact. +- `ADMISSION_CONTROL: invalid config/crew-dispatch.json _scheduling.admission_control - ` - the optional fleet-admission policy exists but failed schema validation, so the fleet's admission layer cannot resolve a band. + Do not dispatch new work in this home until the named field is corrected; an unknown field is refused rather than ignored precisely so a typo cannot silently disable a safety condition. + `docs/configuration.md` "Fleet admission control" owns the schema, and `fleet-admission` owns what firstmate does with a resolved band. +- `WAKE_LEDGER: outcome record(s) join no wake record` - that many recorded supervision costs point at a wake this home never drained, so every figure drawn from the ledger overcounts by up to that number. + Treat it as a measurement defect, never as supervision work: report the count rather than any rate or total computed from the file, until the captain decides what to purge. + The records are append-only evidence, so never rewrite or migrate the file to clear the count; `bin/fm-wake-ledger.sh reconcile` restates it on demand and that script owns the join. + A count that grows during a session means something is still recording against unresolvable sequences, which is a bug to escalate rather than a backlog of old damage. +- `WAKE_LEDGER: the wake ledger could not be read ...` - the file exists but could not be opened, so the count above is unavailable rather than zero. + Repair its permissions or path before quoting any supervision-cost figure; an unreadable ledger is reported precisely so it cannot pass as a clean one. - `FLEET_SYNC: : skipped: ` - a benign one-off skip (offline, no origin, local-only); bootstrap continued, investigate only if it blocks work. A skip can also report the bounded fleet-refresh timeout (`FM_FLEET_SYNC_BOOTSTRAP_TIMEOUT`, or a fleet-size-aware default with a 20 second floor); a timeout never blocks startup. - `FLEET_SYNC: : recovered: ` - the clone had drifted onto a clean detached HEAD holding no unique commits and the sync self-healed it (re-attached the default branch and fast-forwarded); no action needed, it is reported only so the self-heal is visible. @@ -44,10 +65,13 @@ When any diagnostic needs captain attention, report the plain consequence and re Resume the emitted supervision protocol after finishing the session-start wake handling. - Any other `PR_CHECK_MIGRATION:` refusal means migration did not complete safely, whether because watcher exclusion, a private path, a diagnostic, quarantine validation, or marker publication could not be proved. Keep each affected poll unavailable, inspect the named private state path, and do not bypass the migration or execute a quarantined artifact; a completed safe-scan marker allows unrelated authenticated polls to continue while private repair remains pending. -- `SECONDMATE_SYNC: secondmate : skipped: ` - the local-HEAD secondmate sync left a live secondmate home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing the primary target commit, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. +- `SECONDMATE_SYNC: secondmate : skipped: ` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update. - `SECONDMATE_LIVENESS: secondmate : skipped: |respawn failed after : ` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process. Investigate the reason because that secondmate is not guaranteed live. -- `NUDGE_SECONDMATES: secondmate : send failed: ` - the secondmate sweep fast-forwarded a running secondmate home and its loaded instruction surface (`AGENTS.md`, `bin/`, or `.agents/skills/`) changed, but the deterministic `fm-send.sh fm-` re-read nudge failed. - Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send. +- `SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox. + Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after same-host connectivity returns; never re-add or dispatch the items from the main backlog. + An unsafe-outbox variant requires path and file-type inspection before any retry. +- `NUDGE_SECONDMATES: secondmate : send failed: ` - secondmate convergence changed a running home's loaded instructions or inherited config, but the deterministic `fm-send.sh fm-` re-read nudge failed. + Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send on the same local or remote route. - `FMX: X mode on ...` / `FMX: X mode off ...` - bootstrap confirmed or removed the local X-mode poll artifacts (`docs/configuration.md` "X mode (.env)"). Only when a running watcher needs the cadence transition applied immediately, restart the home-scoped watcher through the emitted harness supervision protocol; bootstrap deliberately never restarts the watcher itself. diff --git a/.agents/skills/firstmate-coding-guidelines/SKILL.md b/.agents/skills/firstmate-coding-guidelines/SKILL.md index ac1da34dd84..c6ab80a2c9b 100644 --- a/.agents/skills/firstmate-coding-guidelines/SKILL.md +++ b/.agents/skills/firstmate-coding-guidelines/SKILL.md @@ -73,6 +73,11 @@ Briefs for tasks that touch firstmate's own tracked material should tell the cre Firstmate adds this skill's load instruction to firstmate-repo briefs by hand instead. `CONTRIBUTING.md`'s "Development" section carries the same instruction as a durable reminder. +## Generated worker discipline + +For firstmate-repo work, the generated `# Branch conflict resolution` and `# Verification discipline` sections in [`bin/fm-brief.sh`](../../../bin/fm-brief.sh) own the full worker rules. +Follow both for branch shipping and verification even if the current task was scaffolded before those sections existed. + ## Compatibility and enforcement Before changing shared tracked behavior, review every affected supported primary harness and runtime backend rather than checking only the adapters active in the current fleet. @@ -81,6 +86,29 @@ Mark an axis not applicable only after inspecting its integration surface, and u For critical safety, routing, startup, and supervision infrastructure, prefer deterministic and idempotent enforcement over relying on agent memory alone. Keep instructions as the authority and discovery layer, but make repeated execution converge safely and make invalid or unsafe states fail closed wherever the runtime can enforce them. +### Harness-dependent checks + +This section is the single owner of the rule and of how to satisfy it. + +A check is harness-dependent when its verdict comes from something the vendor emits: a process name, rendered output, a spinner or keybind glyph, a banner, or a key the harness binds. +Anything in that class must be proven end to end against the real harness, because a stub or fake agent can only confirm the assumption already written into the stub. +That proof is authorized to spend tokens; the cost is small against a check that silently stops working. + +Build the check on the most structural signal that answers the question, and prefer a kernel or protocol fact over anything a release note could change. +When a rendered surface is genuinely the only source, read more than one independent signal and let any of them carry a positive verdict, so no single vendor string is load-bearing. +Where a surface signal is unavoidable, back it with a guard that fails loudly naming the harness and version rather than degrading quietly. + +Every such check needs two tests, because they fail for different reasons: + +- A portable regression in `tests/` that pins the logic with real processes and no harness, so CI enforces the classifier everywhere it runs tmux. + Drive the signals apart deliberately and assert the verdict survives losing one; assert the divergence itself so the case cannot go quietly vacuous. + Confirm which signal a given construction actually blinds on each supported platform rather than assuming, because the same trick can break different sources on macOS and Linux. +- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`), env-gated and self-skipping, that exercises every INSTALLED harness for real and fails naming the harness and version. + Report an absent harness explicitly rather than passing silently over it, and refuse a pass that checked nothing. + This guard is opt-in and on-demand because standard CI has neither harness binaries nor credentials; run it after every harness upgrade and before trusting refreshed per-harness evidence. + +Record the dated per-harness result in `docs/verification/runtime-backends.md`, and point at the live guard as the command that refreshes it, rather than leaving a version-scoped observation to rot into a false claim. + ## Documentation change review For every changed maintained prose surface, identify its inventory audience, authoritative owner, current-behavior relevance, destination for supporting evidence, and any unique safety fact that removal could lose. @@ -98,6 +126,7 @@ Run `bin/fm-doc-audience-check.sh`; it enforces classification, README setup rou - Run `bin/fm-lint.sh` before treating a script change as done; it is the single owner of the lint definition (file set, config, and pinned shellcheck version) that CI and the no-mistakes pre-push gate both invoke, and it refuses to run under any other shellcheck version. - Colocate tests with the existing pattern in `tests/`, name them `.test.sh`, and extend an existing script rather than inventing a new runner. - Tests must exercise behavior through an executable or public interface and must never assert implementation-source bytes, including through parsers, regexes, snapshots, or indirect wrappers. +- Hand every background process a test launches to `fm_test_reap` as soon as its pid is known, because a case that only kills on its happy path orphans that process on every failing or signalled path. - A maintainer-verification record under `docs/verification/` records active empirical facts, not assumptions or task chronology. - Include the date, version, exact commands run, and exact output needed to support the current guarantee. - Keep incident chronology and delivery evidence in private task reports or PR evidence unless a concise rationale is required to maintain a current safety boundary. diff --git a/.agents/skills/fleet-admission/SKILL.md b/.agents/skills/fleet-admission/SKILL.md new file mode 100644 index 00000000000..a165ffaba27 --- /dev/null +++ b/.agents/skills/fleet-admission/SKILL.md @@ -0,0 +1,111 @@ +--- +name: fleet-admission +description: >- + Agent-only procedure for fleet admission control, the third layer above routing and scheduling. + Use at intake before dispatching new work in a home with an active admission policy, and whenever an admission band other than preferred is returned, released, or overridden. +user-invocable: false +metadata: + internal: true +--- + +# fleet-admission + +Admission control answers one question: should the fleet accept another task right now? +It is the third layer above routing (who is capable) and scheduling (when accepted work runs). + +`docs/configuration.md` "Fleet admission control" owns the policy schema, `bin/fm-admission.sh`'s header owns the command contract and its decision record, and `bin/fm-admission-lib.sh` owns the executable schema check. +This skill owns what firstmate does with the result. + +## The boundary that makes this a real layer + +Admission inspects only properties of the fleet snapshot. +It never inspects the incoming task's tier, project, model, priority, urgency, estimated tokens, or file overlap, so the same snapshot returns the same band for every incoming task. +`bin/fm-admission.sh` enforces this structurally by refusing a task argument. + +A task's criticality is not an admission input. +An explicit captain exception is an override applied *after* a refusal, never a value that changes the fleet's band. +If you ever find yourself wanting a different band because of what the task is, you are doing scheduling, and it belongs in the scheduling layer. + +The other invariants that bind every decision: + +- Most restrictive wins; no combination rule may average a hard result away. +- Unknown is explicit: missing required evidence maps to the configured band, never to zero or "probably fine". +- Admission never stops running work and never mutates a task's tier or priority. +- Deferred and refused requests stay in the owning backlog under `hold_kind=load`; there is no second queue. +- Capacity is released by successful cleanup, not by a worker's `done` line. +- Every threshold, freshness limit, and band mapping is read from configuration; a number this layer cannot read from config does not exist. + +## At intake + +A home with no configured policy prints `admission: not configured` and dispatch is unchanged; do not run this procedure. + +When the policy is active, capture the request durably in its owning backlog first, so nothing is lost regardless of the band. +Then run `bin/fm-admission.sh` once, before routing and before scheduling. +Read the band, not the exit code alone, and act on it: + +| Band | Action | +|---|---| +| `preferred` | Admit this task, then continue to routing and scheduling as normal. Silent. | +| `soft` | Do not dispatch. Apply a `load` hold and tell the captain immediately. | +| `hard` | Do not route or spawn. Apply a `load` hold, tell the captain immediately, and name the condition that must clear. | + +An unknown required signal has already been mapped to its configured band by the evaluator; treat the result as that band and name the missing or contradictory evidence when you report it. + +Admitting several requests in one intake means evaluating them one at a time: each admission changes the snapshot before the next is evaluated. + +## Queueing a deferred or refused request + +`tasks-axi hold --kind load --reason ""` is the whole queue substrate. +The reason is one line and must not contain parentheses. +Keep it to the controlling observation, its configured limit, and the config digest, for example: + +```text +Fleet admission hard: census integrity unreadable against configured unknown_band hard; config sha256:... +``` + +Never put credentials, temporary paths, an opaque score, or an inferred cause in the reason; the full structured evidence belongs in the decision record. +Never use any hold kind other than `load` for an admission outcome - `captain` holds mean a captain decision is pending, and a load hold must never masquerade as one. + +## Releasing + +Re-examination happens at exactly two points, and both already exist. + +1. **Successful cleanup.** `bin/fm-teardown.sh` prints the admission release reminder after a successful teardown. Recompute the band with `bin/fm-admission.sh`, and only if it is `preferred` release load-held requests - one at a time, re-evaluating between each, because each admission changes the snapshot. Scheduling then chooses which released request goes first, using priority and urgency; admission only decides how many new obligations the fleet can take. +2. **Session start.** The startup digest prints the same band for an admission-active home. + +A failed cleanup releases nothing. +A worker's `done` status alone releases nothing. + +A fleet that is already empty while a load-held request waits will not get a teardown event and waits until the next session start. +That gap is deliberate and named; do not add a daemon, timer, or second supervision loop for it. +If you observe it happening - the fleet returned to `preferred` while a load-held request sat untouched until session start - that is one measured missed release event, and it belongs in the captain's next fleet review as evidence, not as a reason to build machinery on the spot. + +## Reporting to the captain + +Follow `AGENTS.md` section 9's translation contract: no band names, config paths, digests, hold kinds, or signal identifiers in captain-facing chat. + +- `preferred` is silent, and so is a repeated observation that changes nothing. +- A newly deferred or refused request is immediate, every time. Several requests deferred together in one intake may share one message that names each of them, but never delay one behind a digest. +- A recovered band that materially affects no waiting request is digest-level. + +The captain-facing shape is the plain consequence: what the fleet observed, what it means, how many requests are waiting, that no worker was started, that existing work continues, and when it will be reconsidered. + +## Overrides + +A captain may direct an exception after a refusal. +Record the original band and its explanation, the exact override authority, the affected task, the configuration digest, and the expected consequence. +No standing priority or urgency field silently overrides a hard fleet decision. + +## Telemetry seam + +The record printed by `bin/fm-admission.sh --json` is the unit of admission telemetry. +It carries the decision identity, the snapshot and configuration digests, every observation with its source and freshness, every rule with its exact JSON config path and configured value, and the controlling rules. +The wake-outcome ledger is being built in parallel and does not expose an extension seam yet, so nothing is persisted today. +When that seam lands, append this record in the ledger owner's format; admission control owns these semantic fields and never opens a competing evidence store. +If a telemetry write ever fails, surface the degraded evidence and apply the configured unknown policy - never a hidden retry loop that blocks intake. + +## What stays dormant + +Do not build a distributed registry, reservation store, lease daemon, RPC service, or remote agent while one local session is the only intake authority; the per-home session lock already serializes the real fleet. +Do not populate a numeric threshold from intuition: the configured `enforcement_mode` refuses it, and the report checkpoint for each dormant mechanism is the place to revisit it with evidence. +An admission daemon, a task-weighted admission score, and provider quota inside the admission band are permanently rejected designs, not deferred ones. diff --git a/.agents/skills/harness-adapters/SKILL.md b/.agents/skills/harness-adapters/SKILL.md index 03735fefb19..22b60e13d30 100644 --- a/.agents/skills/harness-adapters/SKILL.md +++ b/.agents/skills/harness-adapters/SKILL.md @@ -26,7 +26,8 @@ If `config/crew-harness` is unset or `default`, there is no concrete value to in Inheritance also copies the literal `config/crew-dispatch.json` file, so secondmates apply the same best-fit profile rules for their own crewmates. Each adapter splits into mechanics and knowledge. -The per-task mechanics, including launch command, autonomy flag, and any enabled crewmate turn-end hook, live in `bin/fm-spawn.sh`. +The per-task mechanics, including any enabled crewmate turn-end hook, live in `bin/fm-spawn.sh`. +`bin/fm-launch-lib.sh` is the single owner of every verified launch command and the autonomy, model, and effort flags it carries, for crewmate, scout, secondmate, and primary sessions alike; never hand-write one. The primary-session "no turn ends blind" guard contract and harness hook installation paths live in `docs/turnend-guard.md`. The primary-session watcher wake protocols are rendered from `docs/supervision-protocols/` by `bin/fm-supervision-instructions.sh`. The supervision knowledge lives here: busy state, exit command, interrupt, dialogs, resume behavior, skill invocation, and quirks. @@ -35,7 +36,7 @@ Each adapter's `Busy state` row names only which semantic source that harness us Never dispatch a crewmate or secondmate on an unverified adapter. If `config/crew-harness` or `config/secondmate-harness` names an unverified adapter, tell the captain under `AGENTS.md` section 9 that the requested worker runtime is not verified yet, use firstmate's own verified runtime for current work, and ask only whether to verify the requested runtime before future use. Do not pause current work for that future-verification choice, and never launch an unverified adapter. -If the captain asks for a new harness, propose verifying it first: spawn a trivial supervised task using `fm-spawn`'s raw-launch-command escape hatch, confirm every fact empirically, then record the mechanics in `fm-spawn`, its semantic busy source and trust gate in `bin/fm-busy-lib.sh`, any needed `FM_COMPOSER_IDLE_RE` empty-composer override plus any novel bare agent prompt glyph in `bin/fm-composer-lib.sh`'s shared composer classifier (the one fleet-wide owner of the empty/dead-shell/pending decision, so a new harness's own idle composer is not misread as a dead shell), the tmux agent-process liveness classification in `bin/backends/tmux.sh` when the harness can launch a secondmate, and the verified knowledge here. +If the captain asks for a new harness, propose verifying it first: spawn a trivial supervised task using `fm-spawn`'s raw-launch-command escape hatch, confirm every fact empirically, then record the verified launch command in `bin/fm-launch-lib.sh` and the remaining per-task mechanics in `fm-spawn`, its semantic busy source and trust gate in `bin/fm-busy-lib.sh`, any needed `FM_COMPOSER_IDLE_RE` empty-composer override plus any novel bare agent prompt glyph in `bin/fm-composer-lib.sh`'s shared composer classifier (the one fleet-wide owner of the empty/dead-shell/pending decision, so a new harness's own idle composer is not misread as a dead shell), the tmux agent-process liveness classification in `bin/backends/tmux.sh` when the harness can launch a secondmate, and the verified knowledge here. ## Detection @@ -125,7 +126,7 @@ The supported launch-profile flags below are verified locally; each row records | codex | `--model ` | `-c 'model_reasoning_effort=""'` | Verified on codex-cli 0.142.1. The installed binary schema contains `model_reasoning_effort`, the active config uses it, and the bundled model catalog advertises only low/medium/high/xhigh. `max` is omitted. | | grok | `--model ` | `--reasoning-effort ` | Verified on grok 0.2.99 (2026-07-13). `--effort` is an alias, but firstmate's profile axis is reasoning effort. As of 0.2.99 the ceiling is `high`; both `xhigh` and `max` are rejected with `use one of: high, medium, low`, so firstmate omits them. | | pi / pi-signed | `--model ` | `--thinking ` | Verified 2026-07-27 on Pi and pi-signed 0.82.0. Both expose the same accepted thinking levels and completed the same model-qualified max-thinking smoke. | -| opencode | `--model ` | none for firstmate's interactive launch | Verified on opencode 1.17.6. `opencode run` has `--variant`, but firstmate launches the interactive `opencode --prompt` path, which has no verified effort flag. | +| opencode | `--model ` | none for firstmate's interactive launch | Verified on opencode 1.17.6. `opencode run` has `--variant`, but firstmate launches opencode's interactive TUI, which has no verified effort flag. | | kimi | `--model ` | none | Verified 2026-07-25 on Kimi Code CLI 0.29.1. | The concrete `harness` field owns adapter identity independently of the model provider: `harness=pi` with `model=xai/grok-*` is Pi using xAI, not `harness=grok`, and does not require Grok CLI login; `harness=grok` remains the standalone Grok Build CLI adapter. @@ -156,6 +157,7 @@ This preserves launch success instead of passing a known-bad value. Send the validation skill using the target harness's skill invocation form. Natural language is acceptable if uncertain. +A harness recognizes these forms only at the very start of the composer line, which is why `fm-send` leaves a command-shaped crewmate steer unmarked; if a new adapter introduces another leading sigil, teach `bin/fm-send.sh`'s command-shaped classification about it or its steers will arrive as prose. - claude: `/`, for example `/no-mistakes`. - codex: `$`, for example `$no-mistakes`; `/` is claude-only and codex rejects it as "Unrecognized command". @@ -185,12 +187,14 @@ If such a dialog is showing, accept it from an active firstmate session using `F Claude renders a predicted-next-prompt suggestion as dim/faint text inside an otherwise-empty composer after a turn completes. A plain `tmux capture-pane` cannot tell that ghost text apart from typed text. -Firstmate launches every claude crewmate and secondmate with `CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false`, scoped to firstmate-launched agents through `bin/fm-spawn.sh`, so it never touches the captain's global config. +Firstmate launches every claude crewmate and secondmate with `CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false`, scoped to firstmate-launched agents through the launch templates in `bin/fm-launch-lib.sh`, so it never touches the captain's global config. The CLI's `--prompt-suggestions` flag is print/SDK-mode only and does not suppress the interactive composer ghost text, verified empirically on v2.1.186. As defense in depth for any pane that flag cannot reach, including the captain's own firstmate composer that away-mode reads, the shared `fm_composer_strip_ghost` extractor in `bin/fm-composer-lib.sh` removes dim/faint SGR 2 ghost runs before pending-input classification on both ANSI-capable readers (tmux and herdr). Its broader dark-TRUECOLOR placeholder handling and dark-theme tradeoff are documented in `docs/herdr-backend.md` "Composer and injection safety", with active captures in `docs/verification/runtime-backends.md`. That styled capture is internal to the boolean detector only. `fm-peek` and every other human or LLM-facing capture path stays plain `tmux capture-pane` with no escape codes. +Claude's `CTX ... | COMPACT NOW: /compact` row is host-computed context-pressure telemetry, not a busy signature or supervision event; [`docs/configuration.md`](../../../docs/configuration.md#claude-context-pressure-telemetry) owns its display, task snapshot, and 70-percent advisory behavior. +The underlying `statusLine` `context_window` payload contract was verified 2026-07-24 against the installed Claude Code 2.1.219 bundle (`/home/shane/.local/share/claude/versions/2.1.219`): a grep of that bundle confirmed `context_window` carries `remaining_percentage`, `used_percentage`, `total_tokens`, and `current_usage`. **Primary-session guard fact (verified 2026-07-04, Claude Code 2.1.201; preserved 2026-07-08, Claude Code 2.1.204; Stop-owned auto-arm revalidated 2026-07-24, Claude Code 2.1.219).** This is separate from the per-task crewmate turn-end hook above (that one just `touch`es a marker file in a task's own `.claude/settings.local.json`). @@ -283,7 +287,7 @@ The observed signed process tree is an exact `pi-signed` wrapper parent with the The installed plain `pi` command also execs that signed launcher, so `FM_PI_HARNESS=pi-signed` is the authoritative selection marker and shared unmarked ancestry remains `pi`. Firstmate sets `FM_PI_HARNESS` explicitly for both worker launch identities, and a signed primary uses the README launch command to establish the same boundary. Keep the brief as one positional argument. -Multiple positional args become separate queued messages; `fm-spawn`'s template already does this correctly. +Multiple positional args become separate queued messages; the verified pi template in `bin/fm-launch-lib.sh` already does this correctly. Project trust dialog can appear on the first pi run in any not-yet-trusted directory, observed even on clean worktrees. Accept with Enter. @@ -304,7 +308,7 @@ When a secondmate is launched on Pi or pi-signed, `fm-spawn.sh --secondmate` lau ## grok (VERIFIED 2026-06-29, grok 0.2.73; slash-submit re-verified 2026-07-03 on 0.2.82; reasoning-effort ceiling re-verified 2026-07-13 on 0.2.99; exit paths re-verified 2026-07-19 on grok 0.2.103) Grok Build TUI (`grok`), a Claude-Code-compatible CLI from xAI. -Launch with a positional prompt: `grok --always-approve "$(cat )"`. +A positional prompt starts the supervised interactive session; the verified command itself lives in `bin/fm-launch-lib.sh`, never here. For Grok's supported reasoning-effort values and omission behavior, see the [launch-profile-axes table](#launch-profile-axes). | Fact | Value | diff --git a/.agents/skills/loopspec/SKILL.md b/.agents/skills/loopspec/SKILL.md new file mode 100644 index 00000000000..d7116e32ba8 --- /dev/null +++ b/.agents/skills/loopspec/SKILL.md @@ -0,0 +1,56 @@ +--- +name: loopspec +description: >- + Agent-only procedure for firstmate's canonical LoopSpec primitive, the reusable answer to "under what conditions do we do this again, what may execute, how is progress verified, and when must it stop". + Use before authoring or changing a LoopSpec, before selecting one for an event, before driving an iteration to a terminal state, and before describing any loop's status to the captain. +user-invocable: false +metadata: + internal: true +--- + +# loopspec + +`bin/fm-loopspec.sh` is the single owner of validation, deterministic selection and persistent loop state, and its header plus `--help` own every command, flag and refusal token. +`loopspecs/schema.json` owns the field contract, `loopspecs/triggers.json` owns the trigger register, and each `loopspecs/.json` owns one loop. +This skill owns only what an agent must decide, and it never restates those files. + +## Keep the primitive separate + +A LoopSpec is temporal recurrence and nothing else. +It is not a ticket, a plan, a skill, an ExecutionUnit, a state machine, or a Fractal Node. +Decomposition is not looping: a node that produces three children has decomposed, and the schema recurring at another depth does not make that a loop. +Reach for a LoopSpec only when the real question is whether to do the same work again later. + +## Do not build around it + +The intended path reuses what already exists: an existing detector raises a durable wake, selection is deterministic, authorised skills do the work inside a model turn, a verifier declares progress, state persists, and the loop reaches a terminal state or a bounded next iteration. +Never add another watcher, scheduler, wake queue, supervision daemon, skill registry, evidence ledger or accounting source of truth to serve a loop. +Never write a private loop runner to bypass this substrate, and never schedule work from a detector: the detection plane may detect, classify, deduplicate, enqueue and write evidence, and everything beyond that belongs inside a model turn. + +## Operating sequence + +1. On a wake that might continue a loop, run `select` with the trigger the wake actually carries, never a trigger you infer. +2. Treat exactly one selection line as authority to proceed, and carry its reported version into `claim`. +3. Work only inside the selected spec's permitted skills and actions, and stop at its authority class. +4. Record the outcome with `finish`, naming the terminal the evidence supports. +5. Read the resulting terminal state, then follow the spec's escalation entry for it. + +## A refusal is a stop + +Every refusal token is a fail-closed result, never an obstacle to route around. +Do not retry a refusal with different arguments to obtain a different answer, do not edit persistent state by hand, and do not hand-write a spec version to make a claim succeed. +An unavailable verifier can never become a pass; the only route out of one is the spec's failure terminal. +When a refusal names a real gap, fix the gap through the ordinary delivery path or escalate it, and leave the loop stopped meanwhile. +Missing capacity evidence is unknown capacity, not spare capacity, so supply the measured headroom or accept the refusal. + +## Enabling a loop is a separate, evidenced decision + +A spec may be enabled only when its trigger is implemented, verified and enabled in the register, its permitted skills are installed, and its authority class permits it. +The validator enforces all of that, so a spec that will not enable is telling you which of those is absent. +Marking a trigger `verified` requires witnessed end-to-end execution; a complete description is not an implementation, and neither is a passing description of one. + +## Report loops honestly + +Say what the loop actually reached, in the captain's nouns, and never announce that a loop is live on the strength of a validated spec. +A spec that validates and cannot be selected is ready and inactive, which is a real and reportable outcome. +Reserve any claim that a loop is running for production evidence that it ran. diff --git a/.agents/skills/model-onboarding/SKILL.md b/.agents/skills/model-onboarding/SKILL.md new file mode 100644 index 00000000000..1ef1d080b60 --- /dev/null +++ b/.agents/skills/model-onboarding/SKILL.md @@ -0,0 +1,252 @@ +--- +name: model-onboarding +description: >- + Agent-only admission policy for routing a new model, and the promotion and demotion authority for models already routed. + Load before adding or changing a model in routing config, before probing a model, before acting on an entitlement or price-drift alarm, and before deciding a promotion or demotion. +user-invocable: false +metadata: + internal: true +--- + +# model-onboarding + +The admission gate for any model firstmate may route work to, and the authority model for moving one between reasoning tiers. +`docs/configuration.md` owns the `config/models.json` schema; `bin/fm-model-registry-lib.sh` owns the enforcement mechanics; this skill owns the decision procedure. + +## Why the ordering is what it is + +Three failures produced this policy, and every one of them was a question answered with an instrument that could not answer it. + +| The question actually being asked | Instrument used | Instrument that answers it | +|---|---|---| +| Entitlement - will *this account* get a response? | a catalogue listing | a live probe | +| Freshness - is this answer current? | a local cache read | nothing; a cache cannot answer it | +| Cost class - will this call cost money? | provider identity | per-model price plus explicit free-tier terms | + +Entitlement, cost class and capability are independent axes. +A model can be entitled and billable, unentitled and free-if-it-worked, or catalogued and neither. + +The single most important consequence: + +> **The instrument that catches an entitlement error is itself a billable act on a metered provider.** + +Probing a metered model "just to check it works" is a charge, and a *successful* probe then invites routing to it. +So cost class is established before entitlement, and a probe is authorized only after the model is shown to be subscription-flat or verified-free. +A framework ordered the other way would have prevented the entitlement incident and caused the billing one. + +## The freshness rule + +Every gate answer carries a source kind and a verified-at timestamp. +In descending authority: `probe`, `provider-entitlement`, `provider-doc`, `harness-static-catalogue`, `harness-fetched-cache`, `third-party`, `inference`. + +> **A `harness-fetched-cache` answer may never be the sole evidence for cost class or entitlement.** + +That source is refreshed by the provider underneath you. + +Applied to price, the rule is narrower than it first looks, and the narrowing matters. +A **probe does not establish a price.** +It proves the account gets an answer; it says nothing about what that answer costs, and a model can answer perfectly while being metered. +So an allowlist entry needs a genuinely price-bearing source - `provider-doc` or `harness-static-catalogue` - and the validator refuses one whose price rests on a cache plus a probe. +Treating "it responded" as "it is free" is the shape of the billing incident, not a defence against it. + +## The admission gate + +Eight ordered stages, fail-closed. Refusal at stage N forbids stage N+1 from running. +Most stages are a checklist, not machinery. + +**G0 Identity.** What exactly will the harness send? +Record the provider id and model id resolved through the harness's own resolution path, verbatim. +Refuse a name supplied by the captain, a vendor page, a benchmark table, or another agent that does not resolve. +Treat all model names and access claims from users or third parties as unverified until confirmed. + +**G1 Cost class.** Can this call cost money? +Output exactly one of `subscription-flat`, `verified-free`, `metered`, `unknown`, **plus the price recorded numerically**. +Evidence must include a provider doc *and* the harness's price metadata, and they must agree; a cache alone is insufficient. +`metered` is rejected by cost policy. `unknown` is **refused, not deferred** - an unresolved free-tier boundary is a refusal. +Also record: billing required at signup, overage possible, free tier promotional or withdrawable, card attached that could silently absorb overage. + +> Subscription is not automatically zero-marginal. +> A flat plan can carry a metered dimension on top, enabled per model by account policy. +> Classify `subscription-flat` only after confirming the *specific* usage sits inside the flat allowance. + +**G2 Probe authorization.** May I make a live request? +Authorized if and only if G1 returned `subscription-flat` (with that flat-allowance confirmation) or `verified-free`. +Otherwise the probe needs the captain's explicit word, and the candidate is recorded as *researched, not probed*. +This stage exists only because of the ordering rule above, and it is enforced in code rather than by this prose: `bin/fm-model-verify.sh` consults the zero-budget decision before issuing any live request, on the automatic sweep and on an explicit `--model` alike, and refuses to probe a model the decision refuses. +The captain's explicit word takes the concrete form of `--force-probe`, the only override, and a forced billable probe announces itself on stdout so it is never invisible. + +**G3 Entitlement and liveness.** Will this account actually get an answer? +Run `bin/fm-model-verify.sh --model /`. +The command itself re-checks G2 first: a model the zero-budget decision refuses is not probed, and only `--force-probe` overrides that refusal. +Where the provider offers real entitlement data, read it first and probe second - an empty entitlement set is a refusal that costs nothing to detect. + +Four distinguishable response shapes: + +| Shape | rc | Meaning | Handler | +|---|---|---|---| +| `ok` | 0 | entitled and live | admit to the next gate | +| `entitlement-refused` | 1 | server-side refusal naming the account type | reject; never route | +| `unknown-model` | 1 | unknown id, still sent upstream | reject; identity error at G0 | +| `client-error` | 1 | request never left the machine | configuration error, not a provider fact | + +A client-side failure returns in well under a second; a server-side refusal takes seconds. +That separation matters: a local typo must never be recorded as a provider outage. +Every probe closes stdin - `pi -p` can hang unbounded otherwise, and a wedged probe presents to supervision as a stale worker, making the monitor the fault. + +**G4 Harness expressibility.** Can the harness express what the route requires *on this model*? +Check the effort band, tool calling, structured output, streaming, whether the context ceiling can be pinned, and whether the credential can be referenced by environment-variable name. +Refuse when the route is defined by a control the model cannot accept. + +> Effort bands are **not portable**, and this is a gate rather than a footnote. +> Some models accept no reasoning-effort setting at all; others silently map low and medium onto the provider default. +> A substitution that looks capability-equivalent can therefore change reasoning depth silently - the exact failure the captain ruled against. + +**G5 Competence.** Does it clear the floor for the *lowest* route proposed? +A candidate is never admitted on one successful task. +The evaluation suite is deliberately dormant until a candidate exists; see "What stays dormant". + +**G6 Route assignment.** Which routes, and which are forbidden? +Record eligible routes, **explicitly prohibited routes**, the operational context ceiling (never the advertised maximum), rate and concurrency policy, failover position, and admission status. + +**G7 Reversibility and entry state.** Can this be undone in one step? +The change must be a single declarative edit with a stated rollback, the model enters at observation level **O1**, never straight into the general pool, and the routing rule that will use it is named in advance so its blast radius is known. + +## Cost policy + +> **The monthly paid-usage budget for API-key providers is $0.00.** +> This is a safety rule, not a preference. + +Four access classes: + +- **A Subscription-backed.** Confirm the subscription permits use through the intended harness. Exhaustion is an availability event, not a semantic routing change. +- **B API key with a genuine free tier.** Verify the tier exists now, whether a card is required, whether enabling the API auto-enables overage, the limits and reset behaviour, and that the provider can be configured without enabling paid billing. +- **C Metered with no qualifying free tier.** Rejected by default. Record as researched, not enabled. Do not configure credentials or billing preemptively. +- **D Self-hosted or open-weight.** **Never labelled "free."** Report infrastructure cost separately from the $0.00 API budget. + +> **The mixed-key rule.** Where one credential reaches both free and paid models, **the credential is not the unit of authorization - the model is.** +> Route only to an explicit allowlist; never to a provider generally; never to a provider default; **never substitute an unlisted sibling during failover.** + +Standing prohibitions: do not enable pay-as-you-go, attach a payment method, consume prepaid credits, allow automatic overage, fall through from free into paid, or retry in a way that could trigger billable usage. +Do not treat free credits as a permanent free tier. +When free quota is exhausted, stop rather than continue. + +Keys are referenced by environment-variable name only; no key is ever printed, logged, or committed. + +## Admission statuses + +| Status | Produced by | Routing | Constraints | +|---|---|---|---| +| **Rejected** | a G0/G1/G3/G4/G5 refusal | none, permanently | Record the refusal and its evidence so it is not rediscovered. | +| **Researched but blocked** | G1 `metered`, G2 unauthorized, or unresolved credential, region or terms | none | The dossier stands; eligible only on captain approval. | +| **Experimental** | a thin G5 pass, or uncertain limits | allowlisted tasks only | Limited concurrency, explicit task allowlist, no Tier 1 unless specifically approved, immediate fallback, one-step disablement. Enters at O1. | +| **Approved fallback** | meets a route's floor but is not preferred | pool member for that route | Must meet the **same** floor as the primary. | +| **Approved specialist** | strong on a narrow workload | that workload only | Prohibited routes stated explicitly. | +| **Approved primary** | strong evidence, clean evaluation, dependable access, understood quota, and it **improves on the current primary** | preferred for its routes | "As good as" is not promotion-worthy. | + +Status is orthogonal to availability: a rate-limited approved primary is still an approved primary. + +## Runtime failure policy + +Firstmate observes availability **unevenly**, and the design says so rather than assuming symmetry. +Authentication failure, model unavailability and provider unavailability are cheaply observable; rate limiting, daily and free-tier exhaustion, degradation and context incompatibility are visible only from a dispatch failure. +Build any quota-aware failover for that asymmetry, not around an assumption of symmetric telemetry. + +Discriminate a model outage from a provider outage with two probes: a sibling that succeeds means a model outage, so substitute within the tier; a sibling that also fails means a provider outage, so promote a tier. + +**On free-tier exhaustion:** mark unavailable, do not cross into paid usage, do not repeatedly retry, fall back only to a same-route candidate on the allowlist that meets the floor, otherwise stop and escalate. + +**The terminal state:** + +> **When no candidate in a route's pool meets that route's floor and is available, firstmate stops, queues the work, and reports to the captain immediately, naming the route, the floor, every candidate considered, and why each was unavailable.** +> **It does not degrade, does not substitute below the floor, and does not cross into paid usage to keep working.** + +Throughput reaching zero on a route is an acceptable outcome; silent degradation is not. + +## Promotion and demotion + +**Authority, ruled 2026-07-28.** The registry validator enforces this as a ceiling - a home may be more conservative, never more permissive. + +| Transition | Authority | +|---|---| +| Tier 4 -> Tier 3 | **Automatic** once configured thresholds are met, surfaced in the immediate notification band: the captain is told at once, not asked. | +| Tier 3 -> Tier 2 | **Captain confirmation.** Firstmate proposes with the accumulated evidence; the captain approves or declines. | +| Tier 2 -> Tier 1, Tier 1 -> Tier 0 | **Never entered by accumulated evidence.** | + +The hard ceiling is not a tuning choice. +Tier 1 is triggered by **risk**, not capability rank, and a model that has completed two hundred clean Tier 2 tasks has demonstrated nothing whatsoever about credential handling or destructive-operation judgment. +Entry to Tier 1 or Tier 0 is a captain decision informed by a domain-specific risk evaluation. + +**Evidence.** A bounded fixture suite systematically flatters the cheaper candidate, so it can *reject* but never *promote*. +Promotion therefore draws on real dispatch history. +A candidate test for tier N+1 is **a real task whose discriminator is tier N+1's**, not a harder task of tier N's shape. + +Three rules on any deliberate candidate test: + +1. **The candidate never judges itself.** Models are unreliable judges of the abstraction level of their own work; this mirrors the rule that an implementation worker never answers its own ask-user finding. +2. **It is real work, never a fixture.** +3. Every threshold is configuration, and firstmate must explain any promotion by naming the current value it acted on. + +**Demotion is deliberately faster than promotion**, because reasoning quality must never silently degrade. +A landed defect at the promoted tier that the tier's own verification path should have caught, or work performed across a tier boundary it should not have crossed, demotes on a single occurrence. +Lost entitlement, a changed cost class, or a harness that can no longer express the route's effort band demotes immediately **and suspends the route**. + +On demotion the model returns to its previous tier, promotion evidence **resets to zero rather than decaying** so a marginal model cannot oscillate across the threshold, a cooldown blocks re-promotion, and observation re-escalates. + +> **Demotion is a routing change, not an availability change.** +> A rate-limited, quota-exhausted or cooling-down model is **not demoted** - it is unavailable, which the failure policy handles on an independent axis. +> Conflating them would make every outage permanently degrade the routing table. +> The separation is structural: availability lives in `state/model-health.json`, routing status in `config/models.json`, written by different code. + +## Observation levels and the floor that never reaches zero + +Monitoring intensity is defined concretely as how often the gate is re-verified plus how deeply evidence is reviewed. + +| Level | Gate re-verification | Evidence review | Entry condition | +|---|---|---|---| +| **O1 Probation** | probe before first dispatch each day | every task reviewed | newly admitted, or newly promoted | +| **O2 Watch** | probe at session start | weekly | a run of clean tasks at O1 | +| **O3 Routine** | probe at session start | monthly | a longer clean run at O2 | +| **O4 Maintenance floor** | probe at session start **plus a price-drift check** | on trigger only | stable **and structurally promotion-ineligible** | + +Tapering stops at O4, entered only when the model is stable and **structurally** ineligible for further promotion. +That is checkable rather than a judgment call, and holds in exactly two ways: the ladder has no next rung (the model sits at the top of the automatic ladder), or a floor it cannot meet (G4 shows it cannot express a control the next tier requires). +"We decided not to promote it" is **not** structural ineligibility. + +**Observation never reaches zero.** At O4 two checks still run at every session start, forever, because **the two things that decay are not properties of the model** - they are properties of the account and of the provider's price list, and both change without warning. +A model can be perfectly stable while its entitlement is revoked and its price is raised. + +Re-escalate O4 -> O2 (never back to O1, which is for unproven models) on any demotion trigger, a failed entitlement probe, **any price drift**, a harness upgrade that changes the model's catalogue entry, a provider incident affecting that model, or a terminal task line recording a repair or escalation. + +## What stays dormant, and what activates it + +| Component | Status | Activated by | +|---|---|---| +| The gate, statuses, floors, failure policy | live now | - | +| Session-start probe and price-drift floor | live now | - | +| Registry integrity and spawn refusal | live now | - | +| Evaluation suite | **dormant** | a candidate actually reaching G5 | +| Automatic promotion thresholds | **dormant** | the evidence instrument producing terminal task lines, then `promotion.enabled` | +| Shadow dual-dispatch evaluation | **dormant** | the free natural experiment proving insufficient | + +Promotion activation is a **config and data condition, never a code change**: the named instrument must be producing terminal task lines and `promotion.enabled` must be true. +`fm_model_promotion_state` reports which of the two is unmet, because a trigger nobody can check is indistinguishable from a rejected one. + +Do not build per-model capability scores, a provider-health abstraction layer, or automatic quota-aware profile arrays without fresh evidence that the simpler mechanism failed. +Scoring a handful of models on several axes from a short suite manufactures precision the evidence cannot support; pass/fail against a floor is the honest granularity. + +## Operating checklist + +Adding a model: + +1. Load this skill and walk G0-G7 in order. Do not skip to the probe. +2. Record the dossier under `data/model-onboarding-policy/dossiers/`, including the exact rollback edit - admission time is the only time that is cheap to know. +3. Add the registry entry, and the `zero_budget.allowlist` entry when the provider is API-key backed. +4. Only then add the dispatch rule. Bootstrap refuses a rule naming an unregistered, non-approved, or unprobed model. + +Acting on an alarm: + +- `MODEL_VERIFY: ... REFUSED by the provider` - the account cannot use it. Remove it from routing; this is a rejection, not an outage. +- `MODEL_VERIFY: ... failed locally` - a configuration error on this machine, not a provider fact. +- `MODEL_PRICE: ... no longer zero` - suspend the route immediately, then re-verify the cost class at G1. +- `MODEL_PRICE: ... price drifted` - re-verify G1 and update `price_at_verification`. +- `MODEL_REGISTRY: ...` - the dispatch config and the registry disagree. Fix the config; do not weaken the check. diff --git a/.agents/skills/process-event-sources/SKILL.md b/.agents/skills/process-event-sources/SKILL.md index 8bff80d8a62..8b02dc4b504 100644 --- a/.agents/skills/process-event-sources/SKILL.md +++ b/.agents/skills/process-event-sources/SKILL.md @@ -28,7 +28,11 @@ For a Lavish review artifact: bin/fm-procevent-lavish.sh arm ``` -`bin/fm-procevent.sh --help` and `bin/fm-procevent-lavish.sh --help` own the exact commands and flags. +A configured remote secondmate reply source is armed and handled through `bin/fm-procevent-remote-reply.sh`. +Its header owns exact commands, while the adapter owns cursor continuity, validated deduplicated status ingest, path-confined document fetch, acknowledgement, and re-arming after a good delta. +A continuity break is escalated once and stays unarmed until an operator deliberately rebases it. + +`bin/fm-procevent.sh --help`, `bin/fm-procevent-lavish.sh --help`, and `bin/fm-procevent-remote-reply.sh --help` own the exact commands and flags. Two rules the commands cannot enforce for you: @@ -54,6 +58,7 @@ Two rules the commands cannot enforce for you: Supported by tests: - output that reached the runner is stored atomically at mode `0600` **before** any event referencing it is published; +- the remote-reply adapter reads its append-only source non-destructively from an offset plus prefix hash, so a pre-capture retry can derive the same bytes again, while source truncation or replacement is detected rather than silently rebased; - proactive delivery and adapter-owned terminal retirement follow the operating contract in [`docs/configuration.md`](../../../docs/configuration.md); - a durably captured result with no handled acknowledgement remains eligible for bounded re-announcement across any number of drains and restarts, and repeat wakes retain the same source and sequence for deduplication; - the handled acknowledgement is generation-keyed to the exact source and sequence, private, path-safe, durable, and idempotent, and is the only thing that stops re-announcement; @@ -67,7 +72,8 @@ Supported by tests: The currently published `lavish-axi poll` destructively clears feedback before returning it. A result lost after that clearing and before the runner reads the process output is unrecoverable, and no firstmate wrapper can close that source-side window. -Say this plainly wherever the behavior is described. +The remote-reply adapter removes that particular pre-capture window by never consuming its source, but it cannot recover bytes truly lost from the remote log itself. +Say these boundaries plainly wherever the behavior is described. ## Talking to the captain about it diff --git a/.agents/skills/project-management/SKILL.md b/.agents/skills/project-management/SKILL.md index f4c62a57901..8feb522bd0c 100644 --- a/.agents/skills/project-management/SKILL.md +++ b/.agents/skills/project-management/SKILL.md @@ -29,43 +29,50 @@ Apply `AGENTS.md` section 7's authoritative secondmate routing rules; if an exis Absence from the main `data/projects.md` registry is never evidence that no second mate owns the domain. If the owning second mate cannot accept the route, report that concrete blocker or obtain an explicit captain redirection rather than silently duplicating the project in the main home. -Resolve the project name, destination, delivery mode, and autonomy posture before changing local or remote state. +Resolve the project name, destination, delivery posture, and autonomy posture before changing local or remote state. Keep a newly added clone and its registry entry consistent, and roll back only artifacts created by the incomplete operation when a later initialization step fails and that rollback is safe. Do not overwrite or repurpose an existing path. ## Delivery posture -Choose the delivery mode when adding or creating the project: +The registry records the project's standing posture, which is the captain's default for the work rather than any task's answer; `AGENTS.md` section 7 owns how each task's concrete mode and yolo are resolved at intake and passed explicitly to the brief, the spawn, and any promotion. +Choose that posture when adding or creating the project: -- `no-mistakes` runs the full validation pipeline before a PR and is the default when the captain does not specify a mode. +- `no-mistakes` runs the full validation pipeline before a PR. - `direct-PR` pushes and opens a PR without the no-mistakes pipeline. - `local-only` has no required remote or PR and lands only through the approved local fast-forward path. +- `no-mistakes-prod-only` is a conditional policy rather than one flat mode: genuinely internal-only tooling, automation, contributor or operator process, and release or submission work ships `direct-PR`, while product-facing, mixed, and uncertain work ships `no-mistakes`. + +`no-mistakes-prod-only` is the default for a newly added or created remote-backed project when the captain specifies nothing, and a project with no remote defaults to `local-only`. +State that resolved default while confirming the source, local name, and posture instead of asking the captain to choose from scratch, and record a flat mode instead whenever they ask for one. +Existing registry entries keep the meaning they already have and are never migrated or reinterpreted, so a legacy entry with no bracket stays `no-mistakes`. +Registering a conditional policy is a one-time choice and never requires classifying any change; the per-task surface classification happens at each task's intake, and internal-only is never inferred from file location or project name. The optional `+yolo` posture changes routine approval authority but does not change the delivery mode. -Default it off, and enable it only on the captain's explicit instruction. +Default it off for every project and every posture, and enable it only on the captain's explicit instruction. `AGENTS.md` section 7 owns the complete authority boundary and exceptions when it is on. ## Add or clone an existing project -Confirm the source URL, local project name, delivery mode, and autonomy posture. +Confirm the source URL, local project name, delivery posture, and autonomy posture, stating the resolved default for each rather than asking the captain to invent one. Clone into `projects/` and add the registry entry only after the destination is known to be unused. -A `no-mistakes` project must have an `origin` remote and must complete the initialization procedure below. +A `no-mistakes` or `no-mistakes-prod-only` project must have an `origin` remote and must complete the initialization procedure below, because a conditional policy's product-facing work runs the pipeline while its internal-only work still takes the direct PR. A `direct-PR` project needs an `origin` remote but skips no-mistakes initialization. A `local-only` project may have no remote and skips no-mistakes initialization. ## Create a project Creating a GitHub repository is outward-facing. -Before making that remote change, propose the repository name, owner or organization, visibility, and delivery mode, defaulting visibility to private and delivery mode to `no-mistakes`, then obtain the captain's explicit consent for those values. +Before making that remote change, propose the repository name, owner or organization, visibility, and delivery posture, defaulting visibility to private and the posture to `no-mistakes-prod-only`, then obtain the captain's explicit consent for those exact values; a stated default never replaces that consent. Use `gh-axi` for the approved GitHub operation and consult its current help rather than relying on remembered flags. -After remote creation succeeds, clone it locally, add the registry entry, and initialize it according to its delivery mode. +After remote creation succeeds, clone it locally, add the registry entry, and initialize it according to its delivery posture. For a purely `local-only` project, create a local Git repository under its unused `projects/` path, add the registry entry, and make no GitHub call. The captain's request to create that local project authorizes this local initialization, but it does not authorize an unmentioned remote repository. ## Initialize -Run no-mistakes initialization only for `no-mistakes` projects: +Run no-mistakes initialization only for `no-mistakes` and `no-mistakes-prod-only` projects: ```sh cd projects/ && no-mistakes init && no-mistakes doctor diff --git a/.agents/skills/research-approved-work/SKILL.md b/.agents/skills/research-approved-work/SKILL.md new file mode 100644 index 00000000000..5acd50c4f6a --- /dev/null +++ b/.agents/skills/research-approved-work/SKILL.md @@ -0,0 +1,105 @@ +--- +name: research-approved-work +description: >- + Answer "which reported work was genuinely approved and is still unimplemented?" over this home's scout-report corpus without loading it. + Use when the captain asks what approved work is outstanding, still owed, or unfinished, when reconciling a recommendation register against reality, and before commissioning work that an earlier investigation may already have covered. + Read-only: it produces classified evidence, never a change to code, reports, or decision state. +user-invocable: true +metadata: + internal: true +--- + +# research-approved-work + +This skill answers one recurring question over a corpus too large to read: **which reported work was genuinely approved and remains unimplemented?** + +`bin/fm-research-scan.sh` owns every deterministic step and runs with no model involvement. +This file owns the judgement the scanner is not allowed to make. + +## Read-only boundary + +Producing this answer authorises nothing. +Do not implement anything you find, edit or annotate a report, close or reopen a backlog item, register or resolve a held decision, open or merge a pull request, or change any decision record. +An item classified `approved-unimplemented` is a finding to relay, and commissioning it is a separate captain decision under the ordinary task lifecycle. + +## Procedure + +**1. Scan first, always.** + +``` +bin/fm-research-scan.sh +``` + +If it prints `verdict=no_delta`, the corpus, the durable decision records, and every implementation HEAD are all unchanged since the last run. +**Stop reading here and answer from the previous run's findings.** +Do not open a report, do not re-derive a classification, do not "just check one thing". +That terminal is the entire point of the scanner: reaching it must cost no model turn. + +If it prints `verdict=delta`, only the reports on its `changed=` lines need fresh attention. +Every other report's evidence is already cached and unchanged; reopening one is wasted context. + +**2. Work from bounded extractions, not reports.** + +`bin/fm-research-scan.sh show ` prints a report's cached projection: headings, recommendation identifiers, and decision-language excerpts. +Open the underlying `report.md` only when a specific classification turns on wording the projection genuinely does not carry, and then read only the section you need. +Run `bin/fm-research-scan.sh schema` when, and only when, you need to parse the index yourself. + +**3. Prove approval and implementation separately.** + +``` +bin/fm-research-scan.sh evidence --token --token +``` + +Approval and implementation are different questions and neither prover may stand in for the other. +Never conclude "unimplemented" from one absent filename: pass at least two concrete artifacts the work would have created - a config key, a recorded field, a function name, a flag - and the prover refuses an absence verdict below that threshold. +Add `--landing` to check delivery, which is a separate question again. + +**The provers locate evidence; they do not grade it, and you must.** +`approval=mentions-found` means a durable record contains the identifier, nothing more - a commission asking an investigation to examine `LC-R4` mentions it exactly as a ruling approving it would. +Read every `approval_hit=` excerpt and decide whether it approves, commissions, cites, or declines. +`implementation=matches-at-head` means a token appeared in a tracked file - `route=` matches a local shell variable as readily as the recorded dispatch field a recommendation asked for. +Open the `impl_match=` paths and confirm the match is the artifact before calling anything implemented. +`landing=no-title-match` searched only a bounded window of pull request titles, so it is never proof that nothing was delivered, and `landing=unavailable-listing-failed` means the forge could not be read at all. + +Treating any of these three as a verdict reproduces exactly the false answers this skill exists to prevent. + +## Three facts that decide most classifications + +**Approval evidence is fragmented, and one source is not durably recorded.** +Approvals in this home live in ruling documents, in backlog task notes, and in direct captain instructions given in chat. +The scanner sweeps the first two. +The third leaves no durable trace at all, so `approval=no-mentions-in-durable-sources` means *the durable sources are silent*, never *this was never approved*. +Treat it as `insufficient-evidence` and ask the captain. + +**Delivered is not landed.** +Completed work can sit in an unmerged pull request and be absent from every HEAD. +A verdict built only on the working tree re-commissions finished work, so check delivery before calling anything unimplemented. + +**A recommendation is not an approval.** +A numbered register inside a report reads like a work list and authorises nothing. + +## Classes + +Assign the first class whose evidence is satisfied, in this order. +Every class needs a *graded* excerpt or path, never a bare prover verdict. + +| Class | Required evidence | +|---|---| +| `duplicate` | grouped with another item in `duplicates.tsv`; classify the group once | +| `contradicted-by-evidence` | the recommendation rests on a specific claim that current evidence refutes; cite both | +| `superseded` | a later durable record replaces it; cite the successor | +| `implemented-register-stale` | a confirmed artifact at HEAD while the source report still lists it outstanding | +| `approved-blocked` | a ruling or instruction that approves it, no confirmed artifact at HEAD, plus a named blocker: delivery awaiting merge authority, an unmet dependency, or a recorded external wait | +| `partially-implemented` | an approving record, some artifacts confirmed at HEAD and others absent; name which | +| `approved-unimplemented` | an approving record, two or more artifacts absent at HEAD, and no delivery found | +| `proposed-never-approved` | a durable record explicitly declines or rules against it - **never assignable from silence** | +| `insufficient-evidence` | anything else, including every case where only durable-source silence stands against approval, and every case where a match was found but not confirmed | + +## Evidence discipline + +State the evidence, not your confidence in it. +"Likely unimplemented" and "high confidence" are not findings; "no matches at HEAD across 3 signals in 2 repositories, and no delivery in the searched window" is. +Every classification must cite the approval evidence and the implementation evidence that produced it, separately, and name any source that was silent. +Where a captain instruction conflicts with a report's recommendation, the instruction governs and the conflict is reported, not quietly resolved. + +Report what you found and stop. diff --git a/.agents/skills/secondmate-provisioning/SKILL.md b/.agents/skills/secondmate-provisioning/SKILL.md index 598d3e1b975..577eb8d20f3 100644 --- a/.agents/skills/secondmate-provisioning/SKILL.md +++ b/.agents/skills/secondmate-provisioning/SKILL.md @@ -3,7 +3,7 @@ name: secondmate-provisioning description: >- Agent-only reference for persistent secondmate setup and retirement. Use when creating, seeding, validating, launching, recovering, handing backlog to, pushing inherited local material into, or retiring a secondmate home, or when editing data/secondmates.md. - Covers home leases, transactional seeding, project clone restrictions, secondmate harness pins, inherited local-material push, idle charter, handoff helper, and teardown safety. + Covers local leases, whole-home remote routes, transactional seeding, project clone restrictions, secondmate harness pins, inherited local-material push, idle charter, handoff helper, and teardown safety. user-invocable: false metadata: internal: true @@ -17,15 +17,26 @@ Keep the always-inline routing rules in `AGENTS.md` authoritative: route by natu ## Routing table -`data/secondmates.md` has one parser-compatible line per persistent second mate: +`data/secondmates.md` has one parser-compatible line per persistent second mate. +A local route uses: ```markdown - - (home: ; scope: ; projects: , ; added ) ``` +A whole-home remote route uses: + +```markdown +- - (host: ; root: ; home: ; scope: ; projects: , ; added ) +``` + Each registry entry stays concise and single-line: the summary is one sentence naming the durable charter, `scope:` is the natural-language intake responsibility, `projects:` is the non-exclusive clone list, and any extra prose is limited to genuinely domain-specific hard rules that change routing or safety for that secondmate. Natural-language summary and `scope:` text may contain parentheses and semicolons; keep the generated `(home: ...; scope: ...; projects: ...; added ...)` suffix intact so operational consumers resolve its explicit field markers. The `home:` path points to the seeded home containing `data/charter.md`; no extra registry pointer field is needed. +For a remote route, `host:` is an OpenSSH config alias and `root:` is that host's separate tracked Firstmate code root. +A remote second-mate agent always runs on the Herdr backend and every seed, launch, and liveness relaunch first gates its host on `bin/fm-remote-doctor.sh` readiness, so an unready host refuses with that doctor's own gap text rather than half-creating a route; the workers that second mate supervises keep the home's ordinary backend selection. +This release places whole secondmate homes remotely and never individual workers. +[`docs/remote-secondmates.md`](../../../docs/remote-secondmates.md) owns current operator setup and transport behavior. The home-seeded `data/charter.md` is the sole owner of boilerplate idle-by-default behavior, the normal delegation lifecycle, and standard escalation contracts, so point to that charter rather than restating those contracts in the registry entry. The `scope:` field is used during intake. The `projects:` field is a non-exclusive clone list, not ownership. @@ -49,12 +60,20 @@ Keep custom charter text focused on the persistent responsibility, available pro The scaffolded charter, later copied to `data/charter.md`, owns the standard lifecycle and escalation wording. Preserve the generated charter sections unless the domain genuinely needs a hard rule. -Provision the persistent home and registry entry after the charter is filled: +Provision a local persistent home and registry entry after the charter is filled: ```sh bin/fm-home-seed.sh {...|--no-projects} ``` +Provision a whole remote home through its configured SSH host with: + +```sh +bin/fm-remote-home-seed.sh {...|--no-projects} +``` + +The remote command transfers a bounded charter and project-origin manifest, then the remote host clones its own Firstmate home and project origins. +It never copies a project tree or the primary process environment. Pass `--no-projects` in the project position to seed the project-less home described above; the same mutual-exclusion and fail-loud-on-omission rules apply. It may only seed a home with no project clones or project-registry entries, and refuses conversion of populated homes without changing them. `-` durably leases a fresh firstmate worktree via `treehouse get --lease` under the secondmate id. @@ -70,16 +89,21 @@ It also writes the required `.fm-secondmate-home` identity marker, which is giti A bare `` (today's format, e.g. `claude`) behaves exactly as before - harness only, no model/effort flag - so this is fully backward-compatible. `bin/fm-harness.sh secondmate-model` and `bin/fm-harness.sh secondmate-effort` print the optional 2nd/3rd tokens (empty when absent, or when the file is absent/`default`/harness-only); they read only `config/secondmate-harness`, never `config/crew-harness`, which stays a bare adapter name. For a `--secondmate` spawn, `bin/fm-spawn.sh` populates `MODEL`/`EFFORT` from those tokens only when the harness itself came from the secondmate config path for that spawn. -An explicit per-spawn `--harness` flag, positional harness arg, or raw launch command starts clean on model and effort too, unless the caller also passes explicit `--model` or `--effort`. +For a local route, an explicit per-spawn `--harness` flag, positional harness arg, or raw launch command starts clean on model and effort too, unless the caller also passes explicit `--model` or `--effort`. +A remote route accepts only a verified harness adapter and refuses a raw launch command at the host boundary. When the file's tokens do apply, an explicit per-spawn `--model` or `--effort` flag always wins over the file's token for that axis. Because this resolves from the file on every spawn, the pin is durable across every respawn (recovery, `/updatefirstmate`, restart) exactly like the harness axis itself - e.g. `config/secondmate-harness` containing `claude opus` keeps a secondmate pinned to Opus even if the primary's own default model later changes. This is secondmate-only: crewmate/scout model resolution is untouched by this file. This section is the single owner of the secondmate sync and inherited-local-material propagation contract; `AGENTS.md` sections 3 and 4 point here. -Before launch, `fm-spawn.sh --secondmate` locally fast-forwards the home to the primary firstmate checkout's current default-branch commit when it is safe; dirty, diverged, or in-flight homes launch unchanged with a warning. -The locked session-start bootstrap sweep runs the same guarded fast-forward for every live secondmate home, discovered from `state/.meta` records with `kind=secondmate` (`data/secondmates.md` only backfills `home=` for older records). +Before a local launch, `fm-spawn.sh --secondmate` locally fast-forwards the home to the primary firstmate checkout's current default-branch commit when it is safe; dirty, diverged, or in-flight homes launch unchanged with a warning. +The locked session-start bootstrap sweep runs the same guarded fast-forward for every live local secondmate home, discovered from `state/.meta` records with `kind=secondmate` (`data/secondmates.md` only backfills `home=` for older records). That no-fetch path is a purely local fast-forward of tracked files, never an origin fetch, and it never touches the gitignored operational dirs, so a secondmate's backlog, projects, and in-flight work are never disturbed; a linked worktree advances immediately, while a standalone clone that lacks the target receives firstmate updates through `/updatefirstmate`'s origin refresh. -The same launch and the same locked bootstrap sweep also propagate the primary's declared inherited local material: `config/crew-dispatch.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`, and the one shared captain-preference file `data/captain-shared.md`. +A remote launch and locked bootstrap sweep ask the configured host to fast-forward its persistent home to that host's code-root commit under the same clean and ancestry guards. +`/updatefirstmate` first updates the remote code root from its own origin, then runs that guarded home sync. +SSH exit 255 preserves the route and reports unknown completion; it never triggers local respawn or failover. +The same placement-specific launch and locked bootstrap sweep also propagate the primary's declared inherited local material: `config/crew-dispatch.json`, `config/models.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`, and the one shared captain-preference file `data/captain-shared.md`. +`config/models.json` travels with `config/crew-dispatch.json` and is never separated from it: inheriting the dispatch rules without the registry would leave a secondmate's own crewmates outside zero-budget enforcement and make every inherited model read as unregistered in that home. Because these paths are gitignored, that propagation is a separate, primary-authoritative copy independent of the tracked-files fast-forward: it re-converges every live home whether or not its tracked files advanced, and it touches only the declared items. Propagation failures warn without blocking secondmate launch or session-start continuation, and the destination keeps whatever safely validated state the helper left behind. Inheritance copies the literal `config/crew-harness` file, so a secondmate's own crewmates use the primary's crewmate harness only when it names a concrete adapter such as `codex`; an unset or `default` value has nothing concrete to inherit, and the secondmate's own crewmates fall back to the secondmate's own or detected harness instead. @@ -100,7 +124,7 @@ Keep every `data/learnings.md` fully local by captain decision; route fleet-gene No AGENTS.md reread nudge is needed at spawn or respawn because the agent reads instructions fresh on launch; only the bootstrap sweep's running-home instruction-surface advance needs that AGENTS.md re-read. Bootstrap reports successful AGENTS.md re-read sends as `BOOTSTRAP_INFO:` and only emits `NUDGE_SECONDMATES:` when that send fails and needs retry. A separate, literal-content config reread is required whenever inherited `config/*` material changes under an already-running secondmate. -After each successful allowlisted config write, both the locked bootstrap convergence path and mid-session `bin/fm-config-push.sh` use the shared propagation report to build one per-home generation-specific private instruction file from the validated destination post-write bytes for only the allowlisted config items that actually changed for that home (`config/crew-dispatch.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`), in deterministic allowlist order. +For a local home, after each successful allowlisted config write, both the locked bootstrap convergence path and mid-session `bin/fm-config-push.sh` use the shared propagation report to build one per-home generation-specific private instruction file from the validated destination post-write bytes for only the allowlisted config items that actually changed for that home (`config/crew-dispatch.json`, `config/models.json`, `config/crew-harness`, `config/backlog-backend`, `config/backend`, `config/herdr-presentation-spaces`, `config/startup-memory-budget`), in deterministic allowlist order. Each changed path is printed with clear begin/end delimiters and the destination file's full exact new bytes unparsed, or the explicit token `ABSENT` when propagation removed the destination copy. The instruction uses only minimal framing that these are defaults/rules and do not remove judgment; it never includes SHA values, selected profiles, parsed summaries, or any other generated interpretation. `data/captain-shared.md` is not a config file and is never inlined into this instruction file or message. @@ -111,6 +135,9 @@ The propagation, generation publication, and pointer-delivery sequence holds one A newly launched or relaunched secondmate already reads its files at launch, so its pending config-reread generations are discarded or quarantined after cleanup failure and it needs no redundant live-agent config nudge unless propagation changes files after launch. Quarantined pre-relaunch generations are retained in bounded private history, and cleanup skips creating an empty quarantine generation. Successfully delivered generations are retained only within a bounded per-home state history, while pending generations remain until delivery succeeds or a launch supersedes them. +A remote home receives the same allowlisted bytes through `fm-remote-inherit.sh` and gets one marked re-read instruction after a changed transfer. +The parent records that nudge before delivery, retains it after a failed send, and retries the exact same route during locked bootstrap convergence. +It does not receive a pointer to a primary-local generation path that cannot exist on that host. These config values remain defaults and rules only; they must not harden `fm-spawn` to reject a deliberate runtime choice that differs from the configured defaults. For already-live secondmates, use `bin/fm-config-push.sh` to push a mid-session inherited local-material change without running the tracked-file fast-forward. It uses the same live-home discovery and propagation helper as bootstrap, reports each item as `pushed`, `unchanged`, `skipped`, or `error`, and follows the config-reread contract above for changed or pending generations. @@ -138,7 +165,11 @@ bin/fm-backlog-handoff.sh ... ``` After seeding, run this handoff for the new secondmate's in-scope queued items. -The helper resolves and validates the secondmate home from `data/secondmates.md`, then delegates the item move to `tasks-axi mv` (the single owner of the backlog format), which moves each named item - and a whole connected set, blocker plus dependents, atomically - from the main `data/backlog.md` into the secondmate home's `data/backlog.md`. +For a local route, the helper resolves and validates the secondmate home from `data/secondmates.md`, then delegates the item move to `tasks-axi mv` (the single owner of the backlog format), which moves each named item - and a whole connected set, blocker plus dependents, atomically - from the main `data/backlog.md` into the secondmate home's `data/backlog.md`. +For a remote route, the same helper first moves the dependency-closed set atomically from the main backlog into `data/handoff/.outbox.md`, then transfers that backlog-format outbox through `fm-on.sh` and lets the remote home's `fm-backlog-receive.sh` move every not-already-present key under the destination lock. +The outbox is the whole recovery record: its presence means delivery is unfinished, `--resume-pending` safely re-delivers it, and confirmed receipt removes it. +There is no two-phase handoff journal and no tasks-axi release beyond the already-required atomic `mv` capability. +Bootstrap retries pending outboxes when mutation is authorized and emits `SECONDMATE_HANDOFF:` for any that remain. This delegated route remains required when `config/backlog-backend=manual`, which controls only routine firstmate backlog edits. It moves each queued item's whole block - the `- [ ] ...` header plus every following two-or-more-space-indented body line and blank separator, up to the next item or column-0 section heading - byte-exact under the same section, treating an indented `## ...` line as body rather than a section boundary, so neither the header nor its body is duplicated or orphaned. It refuses a selected item with a single-space or tab-indented continuation rather than risk leaving content orphaned in the main backlog. @@ -150,14 +181,16 @@ Do not hand off `local-only` items. ## Recovery -For `kind=secondmate` meta with no window, treat the secondmate as a dead persistent direct report and respawn it with: +For local `kind=secondmate` meta with no window, treat the secondmate as a dead persistent direct report and respawn it with: ```sh bin/fm-spawn.sh --secondmate ``` Use the recorded `home=` in meta. -If meta is missing but `data/secondmates.md` still registers the secondmate, respawn from the registry entry and its persistent on-disk home. +If meta is missing but `data/secondmates.md` still registers the secondmate, respawn from the registry entry and its persistent home. +For a remote route, the same command probes and relaunches only on the configured host. +An SSH transport failure or unreadable remote endpoint remains unknown and must be reconciled on that host; never launch a local replacement. Respawn re-resolves the secondmate harness from current config, uses the same guarded pre-launch sync, and re-propagates inherited local material, so recovered secondmates converge inherited config items and shared captain preferences whenever their home validates; tracked-file sync remains guarded separately. If the secondmate is already running and only inherited local material changed, prefer `bin/fm-config-push.sh` over respawning. @@ -175,7 +208,9 @@ Run `bin/fm-teardown.sh ` for `kind=secondmate` only when the captain or mai The safety check is the secondmate's own home. Teardown refuses while its `state/*.meta` contains in-flight work. -When safe, teardown kills the direct tmux window, removes the `data/secondmates.md` route, clears the main home metadata, and removes the retired secondmate home. +A remote route delegates the same guard to its configured host and additionally refuses while the primary has a pending handoff outbox or unresolved routed reply. +SSH exit 255 preserves the route and local records because remote completion is unknown. +When safe, teardown kills the direct endpoint, removes the `data/secondmates.md` route, clears the main home metadata, and removes the retired secondmate home. Removing a leased home releases its durable treehouse lease via `treehouse return`, so the pool slot is freed for reuse rather than left leased forever. A plain-clone home with no pool slot is simply removed. If `treehouse return` fails for a leased home, teardown stops with state intact rather than raw-removing the directory and hiding a held lease. diff --git a/.agents/skills/updatefirstmate/SKILL.md b/.agents/skills/updatefirstmate/SKILL.md index de95ed35302..0230b31f073 100644 --- a/.agents/skills/updatefirstmate/SKILL.md +++ b/.agents/skills/updatefirstmate/SKILL.md @@ -1,6 +1,9 @@ --- name: updatefirstmate -description: Self-update a running firstmate and its secondmates to the latest from origin. Use when the captain invokes /updatefirstmate (e.g. "/updatefirstmate", "update firstmate", "pull the latest firstmate"). Fast-forwards this firstmate repo's default branch and every secondmate home from origin (fast-forward only, never forced, never disruptive), then re-reads AGENTS.md and nudges each updated secondmate to do the same, so the whole tree runs the latest bin/ and instructions. +description: >- + Self-update a running firstmate and its secondmates to the latest from origin. + Use when the captain invokes /updatefirstmate (e.g. "/updatefirstmate", "update firstmate", "pull the latest firstmate"). + Fast-forwards this firstmate repo's default branch and every local or remote secondmate through its guarded update path (never forced, never disruptive), then re-reads AGENTS.md and nudges each updated secondmate to do the same, so the whole tree runs the latest bin/ and instructions. user-invocable: true metadata: internal: true @@ -14,6 +17,7 @@ Only `AGENTS.md`, `bin/`, and `.agents/skills/` are a running firstmate instruct This skill performs that pull for the running main firstmate and every secondmate, without disturbing any in-flight work. The update is **fast-forward only** - the same sanctioned self-write as the fleet sync firstmate already runs. +For a remote route, it updates the configured Firstmate code root on that host from its own origin, then guardedly fast-forwards the persistent home to that code-root commit. It never forces, never creates a merge commit, never stashes, and advances a target only on a clean fast-forward; anything dirty, diverged, offline, or on the wrong branch is skipped and reported. A tracked-files fast-forward leaves the gitignored operational dirs (data/, state/, config/, projects/, .no-mistakes/) untouched, so a secondmate's in-flight work is never disrupted. This touches only the firstmate repo and its own worktrees, never anything under `projects/`. @@ -24,7 +28,7 @@ This touches only the firstmate repo and its own worktrees, never anything under ```sh bin/fm-update.sh ``` - It fast-forwards this firstmate repo's default branch from origin, then fast-forwards every registered secondmate home (each a treehouse worktree of this same repo, leased at a detached HEAD on the default branch) the same way. + It fast-forwards this firstmate repo's default branch from origin, then updates every registered local or remote secondmate home through its placement-specific guarded path. It prints one status line per target (`updated ..` / `already current` / `skipped: `), followed by two action lines that tell you exactly what to do next: - `reread-firstmate: yes|no` - `nudge-secondmates: fm-...|none` @@ -56,5 +60,5 @@ This touches only the firstmate repo and its own worktrees, never anything under - **Only the firstmate repo and its worktrees** are touched, never `projects/`. It is the same sanctioned self-write as the fleet sync. - **Secondmates are never disrupted.** - A secondmate gets a tracked-files fast-forward (safe while it is mid-task, since its work lives in gitignored operational dirs and separate project worktrees) plus a gentle re-read nudge. + A local or remote secondmate gets a tracked-files fast-forward only when its own checkout is safe to advance, plus a gentle re-read nudge when it changed. It is never torn down, interrupted, or forced. diff --git a/.claude/settings.json b/.claude/settings.json index 0be379c46b7..cd724bf5919 100644 --- a/.claude/settings.json +++ b/.claude/settings.json @@ -1,4 +1,8 @@ { + "statusLine": { + "type": "command", + "command": "\"$CLAUDE_PROJECT_DIR\"/bin/fm-context-statusline.sh" + }, "hooks": { "SessionStart": [ { diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 00000000000..2451ceb0603 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,3 @@ +*.sh text eol=lf +*.bat text eol=crlf +*.cmd text eol=crlf diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 064f1c16131..82061b8d5ff 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -24,6 +24,16 @@ jobs: # re-spell the shellcheck command here; keep CI and the pre-push gate on it. - run: bin/fm-lint.sh + windows-launcher: + name: Windows launcher bridge + runs-on: windows-latest + timeout-minutes: 5 + steps: + - uses: actions/checkout@v6 + - name: Exercise the batch command and WSL entry contracts + shell: bash + run: bash tests/fm-wsl-entry.test.sh + # Deterministic proof that portable parallel shards + portable serial + Herdr # equal the complete tests/*.test.sh inventory with no missing or duplicates, # and that the portable serial CI shards partition that serial lane exactly. @@ -350,8 +360,8 @@ jobs: snapshot_output=$(/bin/bash tests/fm-fleet-snapshot-view.test.sh) printf '%s\n' "$snapshot_output" snapshot_count=$(printf '%s\n' "$snapshot_output" | grep -c '^ok - ') - [ "$snapshot_count" -eq 15 ] || { - echo "::error::expected 15 snapshot/fleet-view tests, got $snapshot_count" + [ "$snapshot_count" -eq 20 ] || { + echo "::error::expected 20 snapshot/fleet-view tests, got $snapshot_count" exit 1 } diff --git a/AGENTS.md b/AGENTS.md index 6e8eea580ad..15de980aa85 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -33,12 +33,13 @@ Hard rules, in priority order: A scout worktree is declared scratch and may be discarded only after its report exists and the shared unresolved-decision completion gate passes. 4. **Crewmates never address the captain.** All crewmate communication flows through firstmate. + Firstmate's own messages to another agent carry the from-firstmate marker owned by `bin/fm-operational-input.sh`, so an unmarked message in an agent's pane is a human at the keyboard; each surface's own contract names the consequence. Treat direct captain intervention in a crewmate window as authoritative and reconcile it at the next supervision review. 5. **Report outcomes faithfully.** If work failed, say so plainly with the evidence. You may maintain this repo's private operational state directly. -Shared tracked material is `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `bin/`, `.agents/skills/`, and public `skills/`. +Shared tracked material is `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `firstmate.bat`, `bin/`, `.agents/skills/`, `loopspecs/`, and public `skills/`. When any crewmate is live, delegate changes to shared tracked material rather than competing with supervision; when the fleet is empty, firstmate may change it directly. This repo is a shared template, while `.env`, `data/`, `state/`, `config/`, `projects/`, and `.no-mistakes/` are captain-private and gitignored. Ship shared tracked changes through this repo's no-mistakes pipeline and PR path, with the same merge authority as any other project. @@ -62,16 +63,20 @@ README.md public overview and development notes .agents/skills/ firstmate-loaded internal skills, committed; each carries metadata.internal=true for installers .claude/skills symlink to .agents/skills for claude compatibility skills/ standalone public installer-facing skills, committed; not loaded by firstmate +loopspecs/ canonical LoopSpec registry, committed: schema.json (field contract), triggers.json (the sixteen-trigger register), and one .json per loop; bin/fm-loopspec.sh is their only interpreter (section 13) +firstmate.bat Windows-to-WSL launcher bridge, committed; docs/windows-launcher.md owns setup bin/ helper scripts, committed; read each script's header before first use .env optional X-mode pairing token; LOCAL, gitignored; presence-gates section 14 config/crew-harness crewmate harness override; LOCAL, gitignored; absent or "default" = same as firstmate. Inherited as the literal file: a concrete primary adapter value also controls a secondmate home's own crewmates (section 4) -config/crew-dispatch.json optional crewmate dispatch profiles; LOCAL, gitignored; firstmate-maintained but human-editable natural-language rules that choose a per-task harness/model/effort profile (section 4). Inherited by secondmate homes +config/crew-dispatch.json optional crewmate dispatch profiles; LOCAL, gitignored; firstmate-maintained but human-editable natural-language rules that choose a per-task harness/model/effort profile (section 4). Also carries the optional inert-by-default `_scheduling.admission_control` fleet-admission policy (section 7). Inherited by secondmate homes +config/models.json optional model registry and enforced zero-budget allowlist; LOCAL, gitignored; binds every routed model to verified cost-class and entitlement evidence (section 4; docs/configuration.md). Inherited by secondmate homes config/secondmate-harness harness the PRIMARY uses to launch SECONDMATE agents, optionally followed by a model and effort token on the same line (" [] []"; section 4); LOCAL, gitignored; absent or "default" harness falls back to config/crew-harness then firstmate's own. The primary's own setting; NOT inherited into secondmate homes (secondmates do not spawn secondmates) config/backlog-backend backlog backend override; LOCAL, gitignored; absent or "tasks-axi" = default tasks-axi backend, "manual" = force routine backlog updates to hand-editing; inherited by secondmate homes (section 10) config/backend runtime session-provider backend override for new tasks; LOCAL, gitignored; absent = falls through to runtime auto-detection (the runtime firstmate itself is executing inside), then tmux; tmux is the verified reference backend (docs/tmux-backend.md), while herdr, zellij, orca, and cmux are experimental spawn backends (docs/herdr-backend.md, docs/zellij-backend.md, docs/orca-backend.md, docs/cmux-backend.md) - herdr and cmux can also be selected by runtime auto-detection, zellij and orca never are (always explicit), and codex-app is not accepted; see docs/codex-app-backend.md; inherited by secondmate homes under the primary-authoritative contract in secondmate-provisioning config/calm Pi Calm presentation preference; LOCAL, gitignored, and not inherited; see docs/configuration.md "Pi Calm preference" config/startup-memory-budget primary-authoritative per-home startup-memory budget; LOCAL, gitignored, materialized as 7,500 estimated tokens by locked primary bootstrap and inherited into secondmate homes; see docs/configuration.md "Startup memory budget" -config/herdr-presentation-spaces optional presence flag for Herdr's default-off disposable single-task visual projection; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Optional presentation spaces" +config/herdr-presentation-spaces optional "off" opt-out from Herdr's default-on disposable single-task visual projection; LOCAL, gitignored; inherited by secondmate homes; see docs/herdr-backend.md "Presentation spaces" +config/trace-context optional presence flag enabling default-off native W3C trace-context propagation to spawned agents; LOCAL, gitignored; inherited by secondmate homes; see docs/configuration.md "Trace context propagation" and docs/trace-context.md config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup") config/wedge-alarm optional away-mode wedge-alarm active-alert directives; LOCAL, gitignored; absent means auto (macOS Notification Center when available); see docs/wedge-alarm.md config/x-mode.env generated X-mode watcher cadence; LOCAL, gitignored; source before arming watcher when present @@ -80,18 +85,21 @@ data/ personal fleet records; LOCAL, gitignored as a whole captain.md this home's domain-local captain preferences and working style; LOCAL, gitignored, canonical even if harness memory mirrors it, and updated with inspect-then-update captain-shared.md main-authoritative shared captain preferences propagated read-only to secondmate homes; LOCAL, gitignored, owned by secondmate-provisioning learnings.md fleet-local operational facts and gotchas; LOCAL, gitignored; dated, evidence-backed, curated, and updated with inspect-then-update - rewrite and prune rather than append forever, the same contract as captain.md; created lazily, absent until this home has a learning to store - projects.md thin fleet navigation registry; firstmate-private, parsed by fm-project-mode.sh (section 6) - secondmates.md secondmate routing table; firstmate-private, maintained by fm-home-seed.sh (section 6) + projects.md thin fleet navigation registry recording each project's standing delivery posture; firstmate-private, parsed for mechanical sync and seeding by fm-project-mode.sh (section 6) + secondmates.md local and remote secondmate routing table; firstmate-private, maintained by the secondmate seed helpers (section 6) + wake-ledger.tsv append-only wake-outcome and terminal-task evidence; bin/fm-wake-ledger.sh owns its format, vocabulary, and append semantics /brief.md per-task crewmate brief, or per-secondmate charter brief when kind=secondmate /report.md scout task deliverable, written by the crewmate; survives teardown projects/ cloned repos; gitignored; read-only except under hard rule 1's concrete captain-approved project operation exception state/ volatile runtime signals; gitignored .status appended by crewmates: ": " wake-event lines, not current-state truth .turn-ended touched by turn-end hooks + .childcpu identity-bound sample of the CPU consumed by the task agent's descendants, so supervision can see work happening in a child process; written by the watcher, removed by teardown .grok-turnend-token firstmate-owned grok hook registry token for the task; removed by teardown .kimi-turnend-token firstmate-owned Kimi hook registry token for the task; removed by teardown - .meta written by fm-spawn: window=, endpoint_task_id=, worktree=, project=, harness=, model=, effort=, kind=, mode=, yolo=, tasktmp=; kind=secondmate also records home= and projects=; a non-default runtime backend records further backend-specific fields (docs/configuration.md "Runtime backend"; bin/fm-backend.sh, section 8); fm-pr-check, including through fm-pr-merge, records one canonical pr= and the forge's pr_head= when available (GitHub pull requests and GitLab merge requests; docs/gitlab-merge-watch.md); fm-x-link appends x_request=, x_request_ts=, x_followups=, and optional x_platform=/x_reply_max_chars= for an X-mode-originated task (section 14) - .herdr-presentation quarantinable attempt and restart-binding journal for Herdr's optional visual projection; never task or endpoint authority; see docs/herdr-backend.md "Optional presentation spaces" + .meta written by fm-spawn: window=, endpoint_task_id=, worktree=, project=, harness=, model=, effort=, kind=, mode=, yolo=, tasktmp=; a ship or scout also records the task's two base references as slot_base=, contribution_target=, and base_state= (bin/fm-task-base-lib.sh); an optional traceparent= only when trace context is enabled (docs/configuration.md "Trace context propagation"); kind=secondmate also records home= and projects=, plus remote_host=/remote_root=/remote_backend=/remote_herdr_session=/remote_target= for a remote route; a non-default runtime backend records further backend-specific fields (docs/configuration.md "Runtime backend"; bin/fm-backend.sh, section 8); fm-pr-check, including through fm-pr-merge, records one canonical pr= and the forge's pr_head= when available (GitHub pull requests and GitLab merge requests; docs/gitlab-merge-watch.md); fm-pr-merge records merge_verification= plus merge_verified_head= for the head it re-verified, or merge_verification=override for an explicitly unverified merge; fm-x-link appends x_request=, x_request_ts=, x_followups=, and optional x_platform=/x_reply_max_chars= for an X-mode-originated task (section 14) + .herdr-presentation quarantinable attempt and restart-binding journal for Herdr's optional visual projection; never task or endpoint authority; see docs/herdr-backend.md "Presentation spaces" + .landing private minimal landing record (pr=, forge pr_head=, project=) written by fm-teardown when a ship task is released before its PR lands; stands in for the removed meta so fm-pr-merge can still land that PR and fm-pr-check can rearm its merge watch .check.sh authenticated slow poll; the watcher dispatches validated PR data and the byte-identified X shim through trusted repository scripts, runs registered custom checks from hash-validated private snapshots, and rejects every other state check without execution .check-trust private content binding created by fm-check-register.sh for an intentional custom check .pr-poll private validated data sidecar for the byte-static PR merge poll @@ -101,9 +109,11 @@ state/ volatile runtime signals; gitignored .pr-check-migration.log private per-task outcomes distinguishing rebuilt or canonically registered replacement polls, quarantined unarmed polls, and incomplete migrations .pr-check-migration-scan-v1 private marker proving the non-executing scan disabled every unsafe legacy check; .pr-check-migration-v1 separately records completed private repairs x-watch.check.sh generated X-mode relay poll shim; present only when opted in (section 14) + loopspec/ persistent per-loop iteration state written only by bin/fm-loopspec.sh; never hand-edited (section 13) pending-replies/ parent-owned secondmate pending-reply records (correlation id, delivery vs reply, recovery, escalation); fm-pending-reply-lib.sh procevent/ registered process-to-event sources, one private record per canonical source id; written only by bin/fm-procevent.sh, and their presence alone keeps supervision required (section 13) procevent-inbox/ private captured results and their durable handled-acknowledgement markers; source output lives here and never in an event line + research-index/ derived, content-addressed prefilter over data/**/report.md; never approval or implementation authority, always safe to delete, rebuilt by bin/fm-research-scan.sh x-inbox/ generated X-mode pending mention payloads; fmx-respond drains it (section 14) x-context/ generated X-mode durable per-request reply context and one-wake offer markers, keyed by request_id; survives inbox cleanup and expires within seven days (section 14; bin/fm-x-lib.sh) x-outbox/ generated X-mode dry-run reply and dismiss previews; inspect it when FMX_DRY_RUN is set (section 14) @@ -113,7 +123,7 @@ state/ volatile runtime signals; gitignored .afk durable away-mode flag; present = sub-supervisor may inject escalations (set by /afk, cleared on user return) .watch.lock .wake-queue.lock watcher singleton and queue serialization locks .claude-autoarm.lock .claude-autoarm-epoch .claude-autoarm-failure-notified .claude-autoarm-failure-alarmed .turnend-claude-blocks .turnend-claude-blocks.lock Claude Stop auto-arm single-flight, epoch, failure-episode, attended-alarm, guard-budget, and budget-lock records; never touch - .hash-* .count-* .stale-* .stale-since-* .paused-* .wedge-escalations-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak watcher internals; never touch + .hash-* .count-* .stale-* .stale-since-* .paused-* .pr-dirty-* .settled-* .wedge-escalations-* .seen-* .hb-surfaced-* .last-* .heartbeat-streak watcher internals; never touch .watch-triage.log watcher's absorbed-wake debug log (size-capped); never relied on, safe to delete .last-watcher-beat watcher liveness beacon, touched every poll (including while absorbing benign wakes); guard scripts read it .subsuper-* .supervise-daemon.* sub-supervisor internals; never touch @@ -141,9 +151,10 @@ A lock-refused session must not spawn, steer, merge, drain the wake queue, repai 1. **Lock** - acquires the per-home session lock first, before anything mutates shared state. 2. **Bootstrap** - detect-only checks (tool/version problems, GitHub auth, the worktree-tangle check, harness override, dispatch-profile validation, backlog-backend status) always run, but routine confirmations stay silent by default. When the lock could not be acquired, the worktree-tangle check uses read-only advisory wording without a checkout repair command. - Home-local stale Herdr projection cleanup and the five bootstrap MUTATING sweeps - non-executing legacy PR-check migration, fleet sync, the local secondmate fast-forward sweep, the secondmate liveness sweep, and X-mode artifact writes - run only when this session actually holds the lock from step 1. - The secondmate liveness sweep deterministically accounts for every registered secondmate: it relaunches only from the recovery-grade `dead` or `missing` states, preserves ambiguous or unreadable targets, and reports skipped or failed guarantees as `SECONDMATE_LIVENESS:` lines (`bin/fm-bootstrap.sh`; `bin/fm-backend.sh`'s `fm_backend_agent_state`). + Home-local stale Herdr projection cleanup and the six bootstrap MUTATING sweeps - non-executing legacy PR-check migration, fleet sync, secondmate convergence, secondmate liveness, pending remote handoff retry, and X-mode artifact writes - run only when this session actually holds the lock from step 1. + The secondmate liveness sweep deterministically accounts for every registered secondmate: it relaunches only from the recovery-grade `dead` or `missing` states, preserves ambiguous, unreadable, or unreachable remote targets, and reports skipped or failed guarantees as `SECONDMATE_LIVENESS:` lines (`bin/fm-bootstrap.sh`; `bin/fm-backend.sh`'s `fm_backend_agent_state`; `docs/remote-secondmates.md`). 3. **Wake queue** - when locked, drains the durable wake queue and prints the raw records prominently as this turn's first work queue; a bounded, clearly labeled historical status-event annotation may follow a valid `signal` record but never replaces it or current-state reconciliation, and a lapsed watcher chain still surfaces here via the same guard alarm. + Every locked drain also prints a bounded fleet-wide `OPEN DECISIONS` section when durable decision records remain open, including when the queue itself is empty; reconcile those entries before continuing. When the lock could not be acquired and verified, the queue is left untouched because no session mutation is authorized, and the guard's tangle/watcher-liveness alarms still print in read-only advisory mode without drain, supervision repair, or checkout repair commands. 4. **Context digest** - the full contents of `data/projects.md`, `data/secondmates.md`, `data/captain.md`, `data/captain-shared.md`, and `data/learnings.md`, each clearly delimited. A file that does not exist prints an explicit `ABSENT` marker, never confused with an empty-but-present file: absence is meaningful (`captain.md` absent means use the firstmate repo's built-in defaults, `projects.md` absent means rebuild it from the clones under `projects/`, etc.). @@ -166,7 +177,7 @@ Load `harness-adapters` before every spawn or recovery and before trust handling The verified harnesses are `claude`, `codex`, `opencode`, `pi`, `pi-signed`, `grok`, and `kimi`; never dispatch on an unverified adapter. If static `config/crew-harness` or `config/secondmate-harness` names an unverified adapter, report it and fall back only to a verified adapter rather than launching it. -`docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation. +`docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-launch-lib.sh` owns the verified launch commands, launch flags, and the fail-closed unverified-adapter guard that `bin/fm-spawn.sh` sources and enforces at spawn. When dispatch profiles exist, consult them at every crewmate or scout intake and pass the resolved concrete profile required by `fm-spawn`. Routing precedence is an explicit per-task captain override, then the best-fit configured rule, then the configured default, then the static crewmate harness. Firstmate alone resolves a matched profile array: run `quota-axi --json` at that intake, evaluate every configured candidate against that current output, and choose with inspectable effective headroom and usable runway, using pace and reserve only later when needed. @@ -184,6 +195,7 @@ Do not add model-specific versions of that policy. `secondmate-provisioning` owns secondmate harness pins and inherited local material, while `harness-adapters` owns the harness consequences. Dispatch only on a backend that `fm-spawn` validates as spawn-capable; pass an explicit per-spawn `--backend` only under that exact task's own authority, never as later-task precedent (selection contract: [`docs/configuration.md`](docs/configuration.md) "Runtime backend"). +When `config/models.json` exists, `fm-spawn` refuses a model whose API-key provider is not on its verified-free allowlist, whose provider cost posture is unclassified, whose registry status is rejected or blocked, or whose recorded concurrency cap is already met; that file is the enforced copy of the zero-budget rule and `model-onboarding` owns the admission policy. A missing dependency, authentication failure, unsupported backend, or version refusal is a blocker; never silently retry on another backend. ## 5. Recovery @@ -248,6 +260,7 @@ For one-off or infrequent operational work, start with the simplest direct end-t Do not build wrappers, control planes, policy layers, custom verifiers, or automation unless the direct path exposes a concrete blocker or repeated need that justifies the added machinery. Before commissioning an investigation, consult existing reports and established evidence. +When that consultation is the question "what approved work is still unimplemented?", or the captain asks what work is outstanding or still owed, load the `research-approved-work` skill rather than reading the report corpus. Classify the deliverable: - **Ship** is the default and produces a project change through the selected delivery mode; once implementation is authorized, dispatch a ship and keep any remaining bounded research inside it unless unresolved uncertainty could materially change whether or what to build. @@ -258,6 +271,12 @@ Never both present a likely-enough solution and launch a parallel design exercis A diagnostic request, report, recommendation, or implementation-ready finding is evidence, not authorization to change code. Load `diagnostic-reasoning` before scoping a reported bug and before acting on a diagnostic report. +Resolve every ship task's concrete delivery mode and yolo posture at intake, and pass both explicitly to the brief, the spawn, and any scout promotion, which all refuse to guess. +A current explicit captain instruction wins; otherwise the project's registry entry is the captain's standing posture, and dropping below its rigor needs a reason you can state. +On a `no-mistakes-prod-only` project, classify the task's surface: internal-only tooling, automation, contributor or operator process, and release or submission work ships `direct-PR`, while product-facing, mixed, and uncertain work ships `no-mistakes`; never infer internal-only from file location or project name. +An unregistered project or absent registry resolves to `no-mistakes` with yolo off, and the registration gap goes to the captain. +Record the resulting mode, yolo, and the one-line reason for any deviation in the backlog item note. + Treat file or subsystem overlap as a risk signal rather than an automatic reason to wait, and dispatch isolated work immediately with no concurrency cap when each change can be independently implemented and validated and the selected delivery path can reconcile ordinary rebases or conflicts. Serialize only for a true semantic dependency, shared mutable external state, incompatible concurrent migration, or another concrete condition that makes independent progress or reconciliation unsafe; same-file editing alone is insufficient, and genuine blockers remain durable. Write the task-specific brief under section 11 before spawning. @@ -266,6 +285,7 @@ Write the task-specific brief under section 11 before spawning. Spawn only through `bin/fm-spawn.sh` after the profile and backend checks in section 4. The spawn must resolve a genuine isolated task worktree distinct from the primary checkout; a failed isolation assertion stops the task. +A spawn also refuses before allocating when a pool slot that the worktree pool would hand out still holds live work; the refusal names the slot, the evidence, and the apparent owner, and is a stop-and-investigate result rather than an obstacle to bypass. After spawning, confirm the worker is processing the brief, handle any trust dialog through `harness-adapters`, and record ship or scout work as under way. A persistent secondmate is recorded in the secondmate registry and runtime state, never as a backlog work item. @@ -296,6 +316,7 @@ Before deciding any ask-user finding, load `ask-user-authority`; the implementat Never merge a red PR. Without a current explicit captain instruction that states the concrete merge, that default stands, and standing `yolo` cannot authorize a red merge; section 1 owns when such an instruction overrides a Firstmate-written standing rule within its exact scope. Use `bin/fm-pr-merge.sh` for every task PR merge so merge metadata is recorded, and use `bin/fm-merge-local.sh` for approved local-only landing; never call a lower-level merge command around their guards. +`bin/fm-pr-merge.sh` re-checks the pull request's current head and refuses a merge it cannot confirm is green, mergeable, and unblocked by review, naming the concrete failing condition and the head whenever GitHub supplies one; treat that refusal as the state to fix, and use its recorded `--allow-unverified` override only on a current explicit captain instruction for that concrete merge. After an autonomous merge, give the captain a one-line full-URL or local-main outcome. ### Validate @@ -318,19 +339,21 @@ Require the matching `resolved` event, forbid `--yes`, and require the worker to Resume fleet supervision immediately after the decision lands. Judge validation by the current-code-matched run step through `bin/fm-crew-state.sh`, not by shell liveness or the last status event. -Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed or checks-passed is done; failed or cancelled is failed. +Running, fixing, or CI states remain working; parked approval or fix-review states require the worker to follow the active gate help; passed is done, and checks-passed is done only when the run's own evidence records it and blocked otherwise, because a head no check examined is unverified rather than green; failed or cancelled is failed. A worker hand-editing, committing, aborting, or restarting during an active validation run duplicates pipeline ownership outside the supersession sequence above; steer it back to the gate response flow. +The other exception is a rebase the pipeline hands back, which the ship brief requires the worker to resolve and commit itself before returning to the gates. The worker reports the PR when CI first becomes green rather than waiting for merge monitoring to finish. ### PR ready, landing, and teardown For PR-based ship tasks, the ready signal depends on mode: `no-mistakes` reports `done: PR checks green` after CI is green, while `direct-PR` reports `done: PR ` after opening the PR. -Run `bin/fm-pr-check.sh ` - it records `pr=` and the forge's `pr_head=` when available in the task's meta and arms the watcher's merge poll. +Run `bin/fm-pr-check.sh ` - it records `pr=` and the forge's `pr_head=` when available in the task's live meta or released-task landing record and arms the watcher's PR poll. Tell the captain the PR's full URL, always the complete `https://...` link rather than a bare `#number`, a concise outcome summary, and the no-mistakes risk level when applicable. A captain instruction to merge is explicit authority; `yolo` is the only standing routine authority. For any custom `state/.check.sh` you write yourself, keep it an ordinary single-link mode-`0700` file, print one line only when firstmate should wake, print nothing otherwise, finish before `FM_CHECK_TIMEOUT`, then bind its current bytes with `bin/fm-check-register.sh ` before the watcher may execute it. Tear down a ship task only after landing is confirmed. +When the captain releases one before its PR lands, that cleanup leaves a durable landing record so `bin/fm-pr-merge.sh` can still land the PR and `bin/fm-pr-check.sh` can rearm its merge watch afterwards. A teardown refusal for uncommitted or unlanded work is a stop-and-investigate result, never an obstacle to bypass. Never force teardown without explicit discard authority. After successful teardown, record completion, retain only the configured recent Done history, and re-evaluate queued work whose blockers and time gates have cleared. @@ -358,6 +381,7 @@ No turn ends blind while work is under way, including turns described as holding At the start of every wake-handling turn, drain the durable wake queue before peeking, reading beyond the reason line, steering, or starting work. Session start is the only exception because its one-shot digest already drained while locked or deliberately left the queue untouched in lock-refused read-only mode. +Treat any `OPEN DECISIONS` section from the drain as actionable reconciliation input even when no wake record was queued. A status line is a wake event, not current state; use `bin/fm-crew-state.sh` when current state matters, especially before re-escalating an old decision, blocker, or pause. A declared `paused:` event means a bounded external wait expected to clear on its own, while `blocked:` means firstmate action is needed. @@ -365,9 +389,11 @@ Handle actionable wakes as follows: 1. For `signal:`, read the listed event lines first, then reconcile current state only where action depends on it. 2. For `stale:`, inspect the recorded endpoint and load `stuck-crewmate-recovery` for a stopped, looping, confused, or unresponsive worker; a deep-inspection reason also requires current-state and validation-log inspection. -3. For `check:`, act on the named poll result, including merges, X-mode events, and process-to-event source results. +3. For `check:`, act on the named poll result, including merges, X-mode events, process-to-event source results, and a conflicted PR, whose worker is steered to rebase immediately rather than left waiting. 4. For `heartbeat:`, review the whole fleet from the structured fleet view, reconcile suspicious tasks and PR state, update the backlog, and never report an unchanged fleet as progress. +After handling each drained wake, record what it cost with `bin/fm-wake-ledger.sh outcome `, naming only the cost and never a wake sequence, which that script resolves; it owns the closed outcome vocabulary and its mechanics. + When any wake reports a merged PR for a project cloned in this home, refresh that clone through the guarded fleet-sync path. When X-linked work reaches a milestone or terminal state, load `fmx-respond`; before terminal teardown, use its promised-final reconciliation when a typed public commitment exists, otherwise post the final completion follow-up so the link clears even if earlier follow-ups were spent. @@ -411,7 +437,7 @@ When evidence uses an internal label, rewrite it before sending: - teardown -> cleanup. - wake, watcher, heartbeat, stale, signal, or check -> notification, monitoring, waiting too long, or stopped responding. - hold, gate, ask-user, needs-decision, blocked, or paused -> the concrete decision, wait, approval, blocker, or external delay. -- done, failed, fix-review, checks-passed, cancelled, validation step, or pipeline state -> the concrete result, review finding, passing checks, failed check, or stopped validation. +- done, failed, fix-review, checks-passed, cancelled, validation step, or pipeline state -> the corroborated concrete result, review finding, verified checks, unverified claim, failed check, or stopped validation. - brief -> instructions. - crewmate -> worker, only when naming the helper matters. - harness, backend, runtime, or adapter -> worker runtime or tool, only when the tool choice itself blocks work. @@ -488,10 +514,12 @@ It performs guarded fast-forward updates of firstmate and registered secondmate These skills are not captain-invocable; load them only at their precise triggers. -- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `PR_CHECK_MIGRATION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `NUDGE_SECONDMATES:`, or `FMX:`); silence and `BOOTSTRAP_INFO:` need no load. +- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `MODEL_REGISTRY:`, `MODEL_PRICE:`, `MODEL_VERIFY:`, `ADMISSION_CONTROL:`, `WAKE_LEDGER:`, `FLEET_SYNC:`, `PR_CHECK_MIGRATION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`); silence and `BOOTSTRAP_INFO:` need no load. - `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report. - `ask-user-authority` - load before deciding any ask-user finding, regardless of the project's `yolo` posture. - `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi output. +- `model-onboarding` - load before adding or changing a model in routing config, before probing a model, before acting on an entitlement or price-drift alarm, and before deciding a model promotion or demotion. +- `fleet-admission` - load at intake before dispatching new work in a home with an active admission policy, and whenever an admission band other than preferred is returned, released, or overridden. - `harness-adapters` - load before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. - `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata. - `project-management` - load before adding, creating, removing, or initializing a project. @@ -501,6 +529,8 @@ These skills are not captain-invocable; load them only at their precise triggers - `decision-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a decision, and when recording or routing the captain's answer. - `process-event-sources` - load before arming a long-polling source, and on any `procevent ` check wake. Never run a registered source's blocking command yourself in a conversational turn. +- `loopspec` - load before authoring or changing a LoopSpec, before selecting one for an event, before driving an iteration to a terminal state, and before describing any loop's status to the captain. + A LoopSpec is temporal recurrence only; never build a private loop runner, and never report a loop as live without production evidence that it ran. - `fmx-respond` - load on an `x-mention ` `check:` wake to handle the mention, on an `x-mode-error ...` `check:` wake to report the X-mode configuration blocker, on a `public-followup ...` `check:` wake or a startup-surfaced public commitment, and on any milestone or terminal wake for an X-mode-linked task before posting its completion follow-up; relevant only when X mode is on. - `firstmate-codexapp` - load before coordinating a visible Codex Desktop thread, evaluating a Codex App backend request, or reconciling Codex Desktop host-tool smoke evidence for Firstmate work. - `firstmate-coding-guidelines` - load before changing firstmate's shared, tracked material, as defined by section 1's list, whether editing directly or briefing a crewmate for a firstmate-repo task. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index e6633ffc920..19b21f2a002 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -35,7 +35,7 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star - This repo is a template for running a firstmate orchestrator agent. `AGENTS.md` is the agent's main job description and names when to load bundled firstmate skills; `CLAUDE.md` is a symlink to it, and `.claude/skills` is a symlink to `.agents/skills`. -- Only shared material is tracked: `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `bin/`, `.agents/skills/`, and `skills/`. +- Only shared material is tracked: `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `firstmate.bat`, `bin/`, `.agents/skills/`, and `skills/`. `.agents/skills/` holds agent-loaded skills that assume a live firstmate home and carry `metadata.internal: true` so installers such as [skills.sh](https://skills.sh) hide them from discovery; `skills/` holds standalone, installer-facing public skills with no firstmate dependency (see the README's "Two-tier skill layout"). Everything personal to one captain's fleet (`.env`, `data/`, `state/`, `config/`, `projects/`, `.no-mistakes/`) is gitignored; never commit it. The root `.tasks.toml` is tracked `tasks-axi` config for `data/backlog.md`; compatible `tasks-axi` is the default backend for routine backlog mutations, with the compatibility definition owned by [`docs/configuration.md`](docs/configuration.md) ("Backlog backend"). @@ -47,7 +47,11 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star Test scripts and helpers in `tests/` are plain bash too. `bin/fm-lint.sh` must pass: it is the single owner of the lint definition (the shellcheck file set, config, and pinned shellcheck version), and both CI and the no-mistakes pre-push gate run it, so local and CI can never diverge. It pins one exact shellcheck version and refuses to run under any other; print it with `bin/fm-lint.sh --required-version` and install that build locally. -- Changes to harness adapters (detection in `bin/fm-harness.sh`, launch and hook mechanics in `bin/fm-spawn.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-tmux-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in `.agents/skills/harness-adapters/SKILL.md`) must be verified empirically against the real harness, never written from documentation alone. +- Treehouse has no such local/CI equivalence, so a locally green suite does not prove CI green. + CI installs the version pinned in `bin/fm-install-treehouse.sh`, while a developer machine runs whatever `treehouse` is on its `PATH`, and nothing reconciles the two. + That gap is not hypothetical: a change verified locally against a newer Treehouse failed CI because the pinned build lacked the flag it used. + When a change touches spawn worktree acquisition, run the affected suites against the pinned build too - `bin/fm-install-treehouse.sh ` installs it beside your own without replacing it, and prefixing `PATH` with that directory reruns any suite under it. +- Harness-adapter ownership spans detection in `bin/fm-harness.sh`, verified launch commands in `bin/fm-launch-lib.sh`, remaining launch and hook mechanics in `bin/fm-spawn.sh`, semantic busy sources and trust gates in `bin/fm-busy-lib.sh`, delivery-only rendered guards in `bin/fm-tmux-lib.sh`, cleanup in `bin/fm-teardown.sh`, and facts in `.agents/skills/harness-adapters/SKILL.md`; the `firstmate-coding-guidelines` skill owns the validation policy for checks that depend on those harnesses. - Changes to runtime session backends (`bin/fm-backend.sh`, `bin/backends/`, and the scripts that dispatch through them) keep current setup and limits in the relevant backend guide and active empirical evidence in [`docs/verification/runtime-backends.md`](docs/verification/runtime-backends.md). - [`docs/documentation-audiences.md`](docs/documentation-audiences.md) and its machine-consumed inventory own prose classification; run `bin/fm-doc-audience-check.sh` after documentation changes. - In Markdown, put each full sentence on its own line. @@ -56,7 +60,7 @@ See the [no-mistakes quick start](https://kunchenguid.github.io/no-mistakes/star ## Development -Tracked changes to firstmate itself - `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `bin/`, `.agents/skills/`, and `skills/` - ship through the `no-mistakes` pipeline on a feature branch and require an explicit merge approval. +Tracked changes to firstmate itself - `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `.github/workflows/`, `firstmate.bat`, `bin/`, `.agents/skills/`, and `skills/` - ship through the `no-mistakes` pipeline on a feature branch and require an explicit merge approval. Before making any such change, load the agent-only `firstmate-coding-guidelines` skill (`.agents/skills/firstmate-coding-guidelines/SKILL.md`). It has the knowledge-placement rules that keep `AGENTS.md` from regrowing after each diet pass. There is no reliable way for `bin/fm-brief.sh`'s scaffold to detect that a task's repo is firstmate itself, so firstmate adds this skill's load line to firstmate-repo briefs by hand. diff --git a/README.md b/README.md index ac54cf70259..09909d41a47 100644 --- a/README.md +++ b/README.md @@ -31,7 +31,7 @@ But the moment you want three project tasks done in parallel - fixes, investigat firstmate flips the model. You talk to a single agent - the first mate - and it runs the crew for you: spawning autonomous agents in a visible session backend, giving each a clean git worktree, supervising them to completion, and handing you finished PRs, approved local merges, or standalone investigation reports. -For larger fleets, you can opt in to persistent secondmates: second mates that are still ordinary direct reports, but run from their own isolated firstmate homes. +For larger fleets, you can opt in to persistent secondmates: second mates that are still ordinary direct reports, but run from their own isolated firstmate homes on this machine or another SSH-reachable host. firstmate is not a model, not a harness, not a skill, not an MCP server, and not a CLI. firstmate is an agent distro for running a crew of agents. @@ -46,7 +46,7 @@ Launching a supported harness inside it instantiates your first mate - and makes - **Disposable worktrees** - each task runs in a clean [treehouse](https://github.com/kunchenguid/treehouse) git worktree, or an Orca-managed worktree when `backend=orca`, so parallel work on one repo never collides. - **Two task shapes** - ship tasks deliver authorized changes; scout tasks leave standalone investigation reports when the intake contract warrants separate research. - **Explicit project modes** - each project ships via `no-mistakes`, `direct-PR`, or `local-only`, with an optional `+yolo` autonomy flag. -- **Optional secondmates** - opt in to persistent second mates that run from isolated firstmate homes with their own `FM_HOME`, state, projects, and session lock, supervising project clones or a project-less firstmate-repo domain, kept on the primary firstmate version by guarded local fast-forwards and checked for live agent processes at session start. +- **Optional secondmates** - opt in to persistent second mates that run from isolated firstmate homes with their own `FM_HOME`, state, projects, and session lock, either locally or as a whole home on an SSH-reachable host, with guarded updates and recovery that never turns an unavailable remote route into a local replacement. - **Event-driven, zero-token supervision** - a bash watcher sleeps on the fleet and wakes the first mate only when something needs you; verified primary harnesses also get a turn-end backstop that blocks or follows up on a blind stop when work is under way and supervision is not live. - **Optional X mode** - opt in with one local `.env` token so firstmate can answer your public `@myfirstmate` mentions, act on normal reversible mention requests through the same lifecycle as chat requests, acknowledge spawned work, and post up to three public-safe completion follow-ups within seven days for genuine milestones and the final outcome without changing non-X behavior; a final reply promised in a thread becomes durable state that is reconciled from disk, so a restart or a compacted conversation cannot lose it; dry-run preview records would-be replies and dismissals locally before go-live. - **Strict project boundary** - the first mate is read-only over your projects except for the narrow guarded and captain-approved operations authorized by [hard rule 1](AGENTS.md#1-identity-and-prime-directives), including fleet sync's guarded safe branch pruning; crewmates make every other project change behind the configured merge authority. @@ -106,11 +106,19 @@ FM_PI_HARNESS=pi-signed pi-signed For Grok, `--trust` is needed once per clone so project hooks and the turn-end guard load; `/hooks-trust` inside Grok works too. For Pi, approve the project trust prompt once per clone on first launch so the tracked `.pi/extensions/*.ts` files auto-load. +On Windows with WSL and Herdr, the repository-root `firstmate.bat` opens the optional harness menu without requiring a remembered terminal command; see [the Windows launcher guide](docs/windows-launcher.md). Pi's `/calm` toggle hides supported transcript chrome, including canonically classified Firstmate operational user rows, and uses a Calm-only animated working boat during active runs while preserving all model context and session data. The hidden operational inputs remain ordinary user-role messages with unchanged delivery, ordering, authority, persistence, and exports. The preference persists for the effective Firstmate home, and toggling it off restores ordinary rendering. [Calm's current behavior and supported limits](docs/calm.md) are separate from its [version-scoped maintainer evidence](docs/calm-mode-feasibility.md). +### Optional: launch from a menu + +If you run the Herdr backend, `bin/fm-launch.sh` offers the same launch as a short menu instead of a remembered command, starts the session in this home's Herdr workspace, and attaches to it. +It shows every configured harness with honest availability, remembers your last choice, and refuses rather than starting a second session in a home that already has one. +Herdr is required for this path; launching a harness directly, as above, stays fully supported on every backend. +See [docs/launcher.md](docs/launcher.md). + ### Talk to it ```sh @@ -157,7 +165,7 @@ Setup guides for tmux (the default) and every other supported backend (herdr, ze You chat with the first mate. It routes each request to a crewmate in its own session endpoint and git worktree, supervises the fleet with a zero-token event-driven watcher, and brings you finished PRs, approved local merges, or investigation reports. -Optional secondmates extend this to persistent second mates, dispatch profiles let you steer which harness handles which task, and an opt-in X mode lets the same fleet answer public mentions. +Optional secondmates extend this to persistent local or whole-home remote second mates, dispatch profiles let you steer which harness handles which task, and an opt-in X mode lets the same fleet answer public mentions. `codex-app` is not a runtime backend yet; [docs/codex-app-backend.md](docs/codex-app-backend.md) owns the Codex App boundary. Full architecture - the supervision engine, worktree isolation, secondmates, dispatch profiles, project modes, optional X mode, fleet sync, and self-update - is in [docs/architecture.md](docs/architecture.md). @@ -198,6 +206,9 @@ Firstmate's skills live in two separate places with different audiences: - [docs/architecture.md](docs/architecture.md) - maintainer architecture for the crew, supervision, worktrees, secondmates, and project modes. - [docs/configuration.md](docs/configuration.md) - environment variables, `FM_HOME`, runtime backend selection, optional X mode, the files you set, and harness support. +- [docs/launcher.md](docs/launcher.md) - the optional harness menu that starts and attaches one primary session in a home. +- [docs/windows-launcher.md](docs/windows-launcher.md) - launch the Herdr-backed harness menu from Windows through WSL. +- [docs/remote-secondmates.md](docs/remote-secondmates.md) - current setup, routing, transfer, recovery, and safety behavior for whole-home remote second mates. - [docs/calm.md](docs/calm.md) - current Pi `/calm` behavior and supported presentation limits. - [docs/wedge-alarm.md](docs/wedge-alarm.md) - configure the active alert for an away-mode escalation delivery that gets stuck. - [docs/tmux-backend.md](docs/tmux-backend.md) - current setup and limits for the tmux reference backend. diff --git a/bin/backends/cmux.sh b/bin/backends/cmux.sh index 12dc7629eb6..745450b11c1 100644 --- a/bin/backends/cmux.sh +++ b/bin/backends/cmux.sh @@ -501,14 +501,12 @@ fm_backend_cmux_send_key() { # [expected-label] fm_backend_cmux_cli send-key --workspace "$FM_BACKEND_CMUX_WORKSPACE" --surface "$FM_BACKEND_CMUX_SURFACE" "$key" >/dev/null 2>&1 } -# fm_backend_cmux_send_text_line: send one line of TEXT then submit. cmux has -# no single-call atomic "run and submit" primitive (like herdr's `pane run`), -# so this composes send (literal) + send-key enter, exactly like zellij's -# equivalent - used for the fixed spawn-time commands (treehouse get, the -# GOTMPDIR export). +# fm_backend_cmux_send_text_line: send one line of TEXT then submit. fm_backend_cmux_send_text_line() { # [expected-label] fm_backend_cmux_send_literal "$1" "$2" "${3:-}" || return 1 - fm_backend_cmux_send_key "$1" Enter "${3:-}" + fm_backend_cmux_send_key "$1" Enter "${3:-}" && return 0 + fm_backend_cmux_send_key "$1" C-c "${3:-}" >/dev/null 2>&1 && return 1 + return 2 } # fm_backend_cmux_capture: bounded plain-text surface capture. No herdr-style diff --git a/bin/backends/herdr.sh b/bin/backends/herdr.sh index cb677be0cfd..a84c71a3bd0 100644 --- a/bin/backends/herdr.sh +++ b/bin/backends/herdr.sh @@ -15,8 +15,9 @@ # herdr-verification-p2.md "Task container shape", refined by # docs/herdr-backend.md "Default task container shape"): ONE herdr workspace PER # FIRSTMATE HOME (the primary, and each secondmate, gets its own), ONE herdr TAB -# per task inside its home's workspace. An optional, default-off presentation -# flag creates a disposable workspace for a clean fresh task instead. That +# per task inside its home's workspace. The default-on presentation projection +# creates a disposable workspace for a clean fresh task instead unless the home +# opts out. That # workspace is a non-authoritative visual projection containing only the normal # task pane. Its random token and mutable label never authorize lookup, # adoption, reuse, closure, deletion, task ownership, or endpoint selection. @@ -108,8 +109,8 @@ FM_BACKEND_HERDR_ESCALATED_PREFIX=".herdr-escalated-" # at a seeded secondmate home's root, containing exactly that secondmate's id. # The primary firstmate home never carries this marker. FM_BACKEND_HERDR_SECONDMATE_MARKER=".fm-secondmate-home" -# The default-off presentation projection is intentionally separate from the -# authoritative task endpoint record. +# The presentation projection is intentionally separate from the authoritative +# task endpoint record. # A per-task journal lives under state/ as .herdr-presentation. # Version 1 records only the attempted projection's random correlator. # Version 2 additionally binds the successful projection's exact home, @@ -118,6 +119,36 @@ FM_BACKEND_HERDR_SECONDMATE_MARKER=".fm-secondmate-home" # No send, capture, Treehouse, or general task-ownership path reads it. FM_BACKEND_HERDR_PRESENTATION_JOURNAL_SUFFIX=".herdr-presentation" +# The config item a home writes to opt OUT of the projection. +FM_BACKEND_HERDR_PRESENTATION_CONFIG="herdr-presentation-spaces" + +# fm_backend_herdr_presentation_enabled : true when this home's +# children should be projected into disposable one-task workspaces +# (docs/herdr-backend.md "Presentation spaces" owns the full contract). +# Projection is ON by default, so an absent config file enables it; a home opts +# out by writing "off". Values are read with the whole-file whitespace-stripped +# convention the other scalar config items already use (config/backlog-backend, +# config/crew-harness), plus case folding. An empty file is the historical +# presence-based opt-in form and still means on, so no home that had the +# projection enabled can be turned off by the default flip. An unrecognized +# value warns and keeps the default rather than failing a spawn over a purely +# visual setting, so a typo is visible instead of silently disabling anything. +fm_backend_herdr_presentation_enabled() { # + local config_dir=${1:-} file value + [ -n "$config_dir" ] || return 0 + file="$config_dir/$FM_BACKEND_HERDR_PRESENTATION_CONFIG" + [ -f "$file" ] || return 0 + value=$(tr -d '[:space:]' < "$file" 2>/dev/null | tr '[:upper:]' '[:lower:]') || value="" + case "$value" in + off) return 1 ;; + ''|on) return 0 ;; + *) + echo "warning: $file: unrecognized value \"$value\"; herdr presentation spaces stay on (write \"off\" to opt out)" >&2 + return 0 + ;; + esac +} + # fm_backend_herdr_workspace_label: the per-firstmate-HOME herdr workspace # label (docs/herdr-backend.md "Default task container shape"). The PRIMARY home (no # secondmate marker) resolves to the constant "firstmate", byte-identical to diff --git a/bin/backends/tmux.sh b/bin/backends/tmux.sh index f8da21bf0de..454f8405942 100644 --- a/bin/backends/tmux.sh +++ b/bin/backends/tmux.sh @@ -20,6 +20,8 @@ # duplicating it, so the two consumers cannot drift apart. # shellcheck source=bin/fm-tmux-lib.sh . "$FM_BACKEND_LIB_DIR/fm-tmux-lib.sh" +# shellcheck source=bin/fm-session-lock-lib.sh +. "$FM_BACKEND_LIB_DIR/fm-session-lock-lib.sh" # fm_backend_tmux_resolve_bare_selector: the live-window-listing fallback for a # selector that is neither an explicit target nor a task selector routed @@ -148,6 +150,80 @@ fm_backend_tmux_current_command() { # tmux display-message -p -t "$1" '#{pane_current_command}' 2>/dev/null } +# fm_backend_tmux_classify_process_name: the single owner of the process-name +# vocabulary shared by every liveness signal below - `agent` for a verified +# harness, `shell` for an idle login/interactive shell, `other` for anything +# else. Keeping one classifier means the two independent name sources can never +# drift into disagreeing about what a given name means. +fm_backend_tmux_classify_process_name() { # [argv0] -> agent|shell|other + local path=$1 argv0=${2:-} base + base=${path##*/} + base=${base#-} + case "$base" in + *claude*|*codex*|*opencode*|*grok*|*kimi*|pi|pi-signed|pi-launcher|Pi) printf 'agent' ;; + zsh|bash|sh|dash|ash|ksh|mksh|tcsh|csh|fish) printf 'shell' ;; + *) + if fm_harness_path_name "$path" >/dev/null || fm_harness_path_name "$argv0" >/dev/null; then + printf 'agent' + else + printf 'other' + fi + ;; + esac +} + +# fm_backend_tmux_foreground_comms: the kernel-side names of every process in +# 's pane tty foreground process group, one full value per line. +# Empty on any failure. +# +# This is the foreground-process-group half of the liveness probe, and it exists +# because `#{pane_current_command}` and `ps -o comm=` expose different name +# fields whose roles vary by platform. On macOS the tmux field can carry a +# harness-rewritten title (Claude Code 2.1.220 reports `2.1.220`) while `comm` +# retains executable identity; the portable Linux regression observes the +# reverse for its version-named executable. Reading both `comm` and argv[0] +# preserves an identifying install path without making either platform's field +# assignment load-bearing. +# +# Scoping to the foreground process group rather than to the pane's descendants +# is what keeps the probe honest in the other direction: a harness-named process +# left running in the background of an otherwise idle pane is deliberately NOT +# reported, so a genuinely agent-free pane still classifies `dead`. It also +# reports every member of a multi-process launcher (the Pi Launcher path runs a +# `pi-signed` wrapper and a `pi` engine in one group), so no launcher needs its +# own special case here. +# +# Like fm_backend_tmux_current_command this is a RAW pane read: tmux answers an +# absent target from the client's active window rather than failing, so callers +# must confirm exact window membership first, exactly as the classifier below +# does, or they will describe some other pane entirely. +fm_backend_tmux_foreground_comms() { # + local target=$1 tty pid pgid tpgid comm + tty=$(tmux display-message -p -t "$target" '#{pane_tty}' 2>/dev/null) || return 0 + [ -n "$tty" ] || return 0 + LC_ALL=C ps -t "${tty#/dev/}" -o pid=,pgid=,tpgid=,comm= 2>/dev/null \ + | while read -r pid pgid tpgid comm; do + [ -n "$comm" ] || continue + [ "$pgid" = "$tpgid" ] || continue + printf '%s\n' "$comm" + done +} + +fm_backend_tmux_foreground_argv0s() { # + local target=$1 tty pid pgid tpgid comm args argv0 + tty=$(tmux display-message -p -t "$target" '#{pane_tty}' 2>/dev/null) || return 0 + [ -n "$tty" ] || return 0 + LC_ALL=C ps -t "${tty#/dev/}" -o pid=,pgid=,tpgid=,comm= 2>/dev/null \ + | while read -r pid pgid tpgid comm; do + [ -n "$comm" ] || continue + [ "$pgid" = "$tpgid" ] || continue + args=$(LC_ALL=C ps -p "$pid" -o args= 2>/dev/null) || continue + args=${args#"${args%%[![:space:]]*}"} + argv0=${args%%[[:space:]]*} + [ -n "$argv0" ] && printf '%s\n' "$argv0" + done +} + # fm_backend_tmux_agent_state: recovery-grade harness-agent state for one # recorded target. See bin/fm-backend.sh's fm_backend_agent_state for the # shared state vocabulary and docs/tmux-backend.md "Agent liveness probe" for @@ -157,8 +233,16 @@ fm_backend_tmux_current_command() { # # An omitted window or a definitive missing-session/server response is # `missing`; any other inventory or pane read failure is `unreadable`, so a # transient tmux problem never licenses a duplicate. +# +# The verdict combines two independent name sources rather than trusting either +# alone. Either source naming a verified harness is enough for `alive`, because +# a false `dead` is the one outcome that can launch a duplicate agent onto a +# live worktree, while the foreground process group - when it is readable - is +# authoritative for the negative verdicts, since it is the only source that can +# distinguish a truly idle pane from a rewritten process title. fm_backend_tmux_agent_state() { # local target=$1 comm session window windows inventory_status + local foreground argv0s name fg_seen=0 fg_shell=0 fg_other=0 case "$target" in *:*:*|'':*|*:'') printf 'unreadable'; return 0 ;; *:*) ;; @@ -187,15 +271,55 @@ fm_backend_tmux_agent_state() { # return 0 fi + foreground=$(fm_backend_tmux_foreground_comms "$target") + while IFS= read -r name; do + [ -n "$name" ] || continue + fg_seen=1 + case "$(fm_backend_tmux_classify_process_name "$name")" in + agent) printf 'alive'; return 0 ;; + shell) fg_shell=1 ;; + *) fg_other=1 ;; + esac + done < [expected-label] fm_backend_zellij_cli "$FM_BACKEND_ZELLIJ_SESSION" action send-keys --pane-id "$FM_BACKEND_ZELLIJ_PANE" "$key" >/dev/null 2>&1 } -# fm_backend_zellij_send_text_line: send one line of TEXT then submit, -# ATOMICALLY - mirrors tmux's `send-keys -t T text Enter` / herdr's `pane -# run`. Used for the fixed spawn-time commands (treehouse get, the GOTMPDIR -# export). Zellij has no single-call atomic "run and submit" action, so this -# composes paste (literal) + send-keys Enter, exactly like send_literal + -# send_key are composed elsewhere - the two-step form is the ONLY form for -# this adapter, unlike tmux/herdr which have a genuinely atomic primitive. +# fm_backend_zellij_send_text_line: send one line of TEXT then submit. fm_backend_zellij_send_text_line() { # [expected-label] fm_backend_zellij_send_literal "$1" "$2" "${3:-}" || return 1 - fm_backend_zellij_send_key "$1" Enter "${3:-}" + fm_backend_zellij_send_key "$1" Enter "${3:-}" && return 0 + fm_backend_zellij_send_key "$1" C-c "${3:-}" >/dev/null 2>&1 && return 1 + return 2 } # fm_backend_zellij_capture: bounded plain-text pane capture. Mirrors diff --git a/bin/fm-admission-lib.sh b/bin/fm-admission-lib.sh new file mode 100644 index 00000000000..9399f7b9a3c --- /dev/null +++ b/bin/fm-admission-lib.sh @@ -0,0 +1,356 @@ +#!/usr/bin/env bash +# fm-admission-lib.sh - shared fleet-admission policy reader and schema validator. +# +# ONE owner of "is this home's `_scheduling.admission_control` policy present, +# inert, or valid?" so bin/fm-bootstrap.sh's startup diagnostic and +# bin/fm-admission.sh's evaluator cannot drift apart on the same config bytes. +# The per-field schema and its operator-facing semantics are owned by +# docs/configuration.md "Fleet admission control"; this file owns the executable +# check of that schema. +# Sourced by scripts and has no side effects on source. +# +# Policy lives in the optional local `config/crew-dispatch.json` under +# `_scheduling.admission_control`. Keys beginning with `_` are operator notes and +# are ignored, matching the surrounding scheduling config's existing convention. +# +# Configured states (fm_admission_state): +# absent - no config file, no `_scheduling`, or no `admission_control` object. +# inert - the object carries only notes, or `enabled` is false. +# active - a valid, enabled policy. +# invalid - the object failed schema validation; the reason is on stdout of +# fm_admission_validate_reason. An invalid policy never resolves to a +# band: callers refuse rather than admitting on unreadable policy. + +# Print the resolved crew-dispatch config path for a config dir. +fm_admission_config_file() { # + printf '%s\n' "$1/crew-dispatch.json" +} + +# The single jq schema program. Prints one reason line when the policy is +# malformed and nothing when it is valid or absent. +# shellcheck disable=SC2016 # jq program, not shell expansion. +FM_ADMISSION_VALIDATE_JQ=' +def real_keys($o): ($o | keys_unsorted | map(select(startswith("_") | not))); +def unknown($o; $allowed): (real_keys($o) - $allowed); +def isbool($v): ($v | type) == "boolean"; +def isstr($v): (($v | type) == "string") and (($v | length) > 0); +def isobj($v): ($v | type) == "object"; +def nullable_num($v): ($v == null) or ((($v | type) == "number") and ($v >= 0)); +def band_value($v): ["soft","hard"] | index($v); + +def signal_source: { + census_integrity: "fresh-authority-census", + backlog_consistency: "main-inventory", + admission_queue_pressure: "tasks-axi-load-holds-plus-ledger", + coordination_debt: "wake-outcome-ledger", + active_workers: "fresh-authority-census", + host_resources: "node-summaries", + reservation_pressure: "admission-registry" +}; +def signal_keys: { + census_integrity: ["enabled","required","source","unknown_band","max_snapshot_age_seconds"], + backlog_consistency: ["enabled","enforce","source","unknown_band"], + admission_queue_pressure: ["enabled","enforce","source","queued_soft_count","queued_hard_count","oldest_wait_soft_seconds","oldest_wait_hard_seconds"], + coordination_debt: ["enabled","enforce","source","pending_wakes_soft_count","pending_wakes_hard_count","oldest_unhandled_wake_soft_seconds","oldest_unhandled_wake_hard_seconds","handled_wake_latency_window_seconds","handled_wake_latency_soft_seconds","handled_wake_latency_hard_seconds"], + active_workers: ["enabled","enforce","source","soft_count","hard_count"], + host_resources: ["enabled","enforce","source","metrics"], + reservation_pressure: ["enabled","enforce","source","soft_count","hard_count"] +}; +# Sources with no collector on this machine today. Enabling one would record an +# invented value instead of an observation, so the schema refuses it outright. +def uncollectable_signals: ["coordination_debt","host_resources","reservation_pressure"]; +def threshold_pairs: { + admission_queue_pressure: [["queued_soft_count","queued_hard_count"],["oldest_wait_soft_seconds","oldest_wait_hard_seconds"]], + coordination_debt: [["pending_wakes_soft_count","pending_wakes_hard_count"],["oldest_unhandled_wake_soft_seconds","oldest_unhandled_wake_hard_seconds"],["handled_wake_latency_soft_seconds","handled_wake_latency_hard_seconds"]], + active_workers: [["soft_count","hard_count"]], + reservation_pressure: [["soft_count","hard_count"]] +}; + +def top_keys: ["schema_version","enabled","enforcement_mode","fleet_id","combine","severity_order","unknown_band","bands","signals","authority","reservations","queue","notifications","telemetry","dormant_triggers"]; + +def signal_error($name; $sig): + if (isobj($sig) | not) then "signals.\($name) must be an object" + elif ((unknown($sig; signal_keys[$name])) | length) > 0 then + "unknown signals.\($name) field: " + ((unknown($sig; signal_keys[$name])) | join(", ")) + elif (isbool($sig.enabled) | not) then "signals.\($name).enabled must be a boolean" + elif ($sig.source != signal_source[$name]) then + "signals.\($name).source must be \"\(signal_source[$name])\"" + elif ($sig | has("enforce")) and (isbool($sig.enforce) | not) then + "signals.\($name).enforce must be a boolean" + elif ($sig | has("required")) and (isbool($sig.required) | not) then + "signals.\($name).required must be a boolean" + elif ($sig | has("unknown_band")) and (band_value($sig.unknown_band) | not) then + "signals.\($name).unknown_band must be soft or hard" + elif ($sig | has("metrics")) and (isobj($sig.metrics) | not) then + "signals.\($name).metrics must be an object" + elif ([signal_keys[$name][] | select(endswith("_count") or endswith("_seconds")) + | select(nullable_num($sig[.]) | not)] | length) > 0 then + "signals.\($name) thresholds must be null or a non-negative number: " + + ([signal_keys[$name][] | select(endswith("_count") or endswith("_seconds")) + | select(nullable_num($sig[.]) | not)] | join(", ")) + elif ([threshold_pairs[$name] // [] | .[] + | select(($sig[.[0]] != null) and ($sig[.[1]] != null) and ($sig[.[0]] > $sig[.[1]])) + | .[0]] | length) > 0 then + "signals.\($name) soft threshold must not be more restrictive than hard: " + + ([threshold_pairs[$name] // [] | .[] + | select(($sig[.[0]] != null) and ($sig[.[1]] != null) and ($sig[.[0]] > $sig[.[1]])) + | .[0]] | join(", ")) + elif ($sig.enabled == true) and (uncollectable_signals | index($name)) then + "signals.\($name) cannot be enabled: its \(signal_source[$name]) source is not collectable yet" + elif ($sig.enforce == true) and ($sig.enabled != true) then + "signals.\($name).enforce requires enabled" + else empty + end; + +def band_error($name; $band; $action): + if (isobj($band) | not) then "bands.\($name) must be an object" + elif ($band.action != $action) then "bands.\($name).action must be \"\($action)\"" + elif ($name == "preferred") and ((unknown($band; ["action"])) | length) > 0 then + "bands.preferred admits with no hold: only action is allowed, not " + + ((unknown($band; ["action"])) | join(", ")) + elif ($name != "preferred") and ($band.hold_kind != "load") then + "bands.\($name).hold_kind must be \"load\"" + elif ($band | has("auto_reconsider")) and (isbool($band.auto_reconsider) | not) then + "bands.\($name).auto_reconsider must be a boolean" + elif ((unknown($band; ["action","hold_kind","auto_reconsider"])) | length) > 0 then + "unknown bands.\($name) field: " + ((unknown($band; ["action","hold_kind","auto_reconsider"])) | join(", ")) + else empty + end; + +def structure_error: + . as $a + | if ((unknown($a; top_keys)) | length) > 0 then + "unknown field: " + ((unknown($a; top_keys)) | join(", ")) + elif ($a.schema_version != 1) then "schema_version must be 1" + elif (isbool($a.enabled) | not) then "enabled must be a boolean" + elif ($a.enforcement_mode != "safety-only") then + "enforcement_mode must be \"safety-only\" until an evidence-gated mode is added" + elif (isstr($a.fleet_id) | not) then "fleet_id must be a non-empty string" + elif ($a.combine != "most_restrictive") then + "combine must be \"most_restrictive\" so no averaging can hide a hard result" + elif ($a.severity_order != ["preferred","soft","hard"]) then + "severity_order must be [\"preferred\",\"soft\",\"hard\"]" + elif (band_value($a.unknown_band) | not) then "unknown_band must be soft or hard" + elif ($a.enabled == true) and ([("bands","signals","authority","queue","telemetry") + | . as $k | select(($a | has($k)) | not)] | length) > 0 then + "an enabled policy needs " + ([("bands","signals","authority","queue","telemetry") + | . as $k | select(($a | has($k)) | not)] | join(", ")) + else empty + end; + +def bands_error: + . as $a + | if ($a | has("bands") | not) then empty + elif (isobj($a.bands) | not) then "bands must be an object" + elif ((unknown($a.bands; ["preferred","soft","hard"])) | length) > 0 then + "unknown bands field: " + ((unknown($a.bands; ["preferred","soft","hard"])) | join(", ")) + elif ([("preferred","soft","hard") | . as $k | select(($a.bands | has($k)) | not)] | length) > 0 then + "bands needs " + ([("preferred","soft","hard") | . as $k | select(($a.bands | has($k)) | not)] | join(", ")) + else + first(band_error("preferred"; $a.bands.preferred; "admit"), + band_error("soft"; $a.bands.soft; "queue"), + band_error("hard"; $a.bands.hard; "refuse")) + // empty + end; + +def signals_error: + . as $a + | if ($a | has("signals") | not) then empty + elif (isobj($a.signals) | not) then "signals must be an object" + elif ((unknown($a.signals; (signal_keys | keys))) | length) > 0 then + "unknown signal: " + ((unknown($a.signals; (signal_keys | keys))) | join(", ")) + elif ($a.enabled == true) and (($a.signals | has("census_integrity")) | not) then + "an enabled policy needs signals.census_integrity" + else + (first(real_keys($a.signals)[] as $n | signal_error($n; $a.signals[$n])) // empty) + end; + +def enforcement_error: + . as $a + | [real_keys($a.signals // {})[] | select(($a.signals[.].enforce // false) == true)] as $enforcing + | if ($a.enforcement_mode == "safety-only") and (($enforcing | length) > 0) then + "enforcement_mode safety-only forbids enforce on " + ($enforcing | join(", ")) + + "; numeric enforcement is evidence-gated and not implemented yet" + else empty + end; + +def authority_error: + . as $a + | if ($a | has("authority") | not) then empty + elif (isobj($a.authority) | not) then "authority must be an object" + elif ((unknown($a.authority; ["mode","authority_id","config_mismatch_band","unreachable_band"])) | length) > 0 then + "unknown authority field: " + ((unknown($a.authority; ["mode","authority_id","config_mismatch_band","unreachable_band"])) | join(", ")) + elif ($a.authority.mode != "single-primary") then + "authority.mode must be \"single-primary\" until a second intake authority is registered" + elif ($a.authority | has("authority_id")) and (isstr($a.authority.authority_id) | not) then + "authority.authority_id must be a non-empty string" + elif (band_value($a.authority.config_mismatch_band) | not) then + "authority.config_mismatch_band must be soft or hard" + elif (band_value($a.authority.unreachable_band) | not) then + "authority.unreachable_band must be soft or hard" + else empty + end; + +def reservations_error: + . as $a + | if ($a | has("reservations") | not) then empty + elif (isobj($a.reservations) | not) then "reservations must be an object" + elif ((unknown($a.reservations; ["enabled","ttl_seconds","heartbeat_seconds","clock_skew_tolerance_seconds","release_on","reconcile_on"])) | length) > 0 then + "unknown reservations field: " + + ((unknown($a.reservations; ["enabled","ttl_seconds","heartbeat_seconds","clock_skew_tolerance_seconds","release_on","reconcile_on"])) | join(", ")) + elif (isbool($a.reservations.enabled) | not) then "reservations.enabled must be a boolean" + elif ([("ttl_seconds","heartbeat_seconds","clock_skew_tolerance_seconds") + | select(nullable_num($a.reservations[.]) | not)] | length) > 0 then + "reservations durations must be null or a non-negative number" + elif ($a.reservations.enabled == true) then + "reservations are dormant until a second intake authority or remote node is registered" + else empty + end; + +def queue_error: + . as $a + | if ($a | has("queue") | not) then empty + elif (isobj($a.queue) | not) then "queue must be an object" + elif ((unknown($a.queue; ["substrate","release_triggers","already_empty_fleet_recheck"])) | length) > 0 then + "unknown queue field: " + ((unknown($a.queue; ["substrate","release_triggers","already_empty_fleet_recheck"])) | join(", ")) + elif ($a.queue.substrate != "tasks-axi hold --kind load") then + "queue.substrate must be \"tasks-axi hold --kind load\" - admission adds no second queue" + elif ($a.queue.release_triggers != ["teardown","session-start"]) then + "queue.release_triggers must be [\"teardown\",\"session-start\"] - admission release triggers are fixed seams, not operator-selectable" + elif ($a.queue.already_empty_fleet_recheck != "session-start-only") then + "queue.already_empty_fleet_recheck must be \"session-start-only\"" + else empty + end; + +def notifications_error: + . as $a + | if ($a | has("notifications") | not) then empty + elif (isobj($a.notifications) | not) then "notifications must be an object" + elif ((unknown($a.notifications; ["policy_ref","episode_dedupe_seconds"])) | length) > 0 then + "unknown notifications field: " + ((unknown($a.notifications; ["policy_ref","episode_dedupe_seconds"])) | join(", ")) + elif (isstr($a.notifications.policy_ref) | not) then + "notifications.policy_ref must be a non-empty string" + elif (nullable_num($a.notifications.episode_dedupe_seconds) | not) then + "notifications.episode_dedupe_seconds must be null or a non-negative number" + else empty + end; + +def telemetry_error: + . as $a + | ["sink","record_every_decision","record_signal_values","record_config_paths","record_config_digest","credentials_forbidden"] as $keys + | if ($a | has("telemetry") | not) then empty + elif (isobj($a.telemetry) | not) then "telemetry must be an object" + elif ((unknown($a.telemetry; $keys)) | length) > 0 then + "unknown telemetry field: " + ((unknown($a.telemetry; $keys)) | join(", ")) + elif (isstr($a.telemetry.sink) | not) then "telemetry.sink must be a non-empty string" + elif ([$keys[] | select(. != "sink") | select(isbool($a.telemetry[.]) | not)] | length) > 0 then + "telemetry flags must be booleans" + elif ($a.enabled == true) and ($a.telemetry.record_every_decision != true) then + "telemetry.record_every_decision must be true while admission is enabled" + elif ($a.enabled == true) and ($a.telemetry.credentials_forbidden != true) then + "telemetry.credentials_forbidden must be true while admission is enabled" + else empty + end; + +def dormant_error: + . as $a + | if ($a | has("dormant_triggers") | not) then empty + elif (isobj($a.dormant_triggers) | not) then "dormant_triggers must be an object" + elif ([real_keys($a.dormant_triggers)[] + | select((isobj($a.dormant_triggers[.]) | not) + or (isstr($a.dormant_triggers[.].checkpoint) | not))] | length) > 0 then + "each dormant trigger needs a named checkpoint: " + + ([real_keys($a.dormant_triggers)[] + | select((isobj($a.dormant_triggers[.]) | not) + or (isstr($a.dormant_triggers[.].checkpoint) | not))] | join(", ")) + else empty + end; + +(._scheduling // {}) as $s +| if (($s | type) != "object") then "_scheduling must be an object" + elif (($s | has("admission_control")) | not) then empty + elif (($s.admission_control | type) != "object") then "admission_control must be an object" + else + $s.admission_control as $a + | if ((real_keys($a)) | length) == 0 then empty + else + first($a | structure_error, bands_error, signals_error, enforcement_error, + authority_error, reservations_error, queue_error, + notifications_error, telemetry_error, dormant_error) + // empty + end + end +' + +# Print the schema failure reason for , or nothing when the policy +# is valid, absent, or note-only. Returns 1 when the file cannot be inspected at +# all (missing jq or malformed JSON) so callers can distinguish "unreadable" from +# "well-formed but wrong"; the reason is still printed. +fm_admission_validate_reason() { # + local file=$1 reason + [ -f "$file" ] || return 0 + if ! command -v jq >/dev/null 2>&1; then + printf 'jq is required to validate admission policy\n' + return 1 + fi + if ! jq -e . "$file" >/dev/null 2>&1; then + printf 'malformed JSON\n' + return 1 + fi + reason=$(jq -r "$FM_ADMISSION_VALIDATE_JQ" "$file" 2>/dev/null) || { + printf 'admission_control could not be validated\n' + return 1 + } + [ -n "$reason" ] && printf '%s\n' "$reason" + return 0 +} + +# Print absent | inert | active | invalid for . +fm_admission_state() { # + local file=$1 reason enabled real + if [ ! -f "$file" ] || ! command -v jq >/dev/null 2>&1; then + printf 'absent\n' + return 0 + fi + reason=$(fm_admission_validate_reason "$file") || { printf 'invalid\n'; return 0; } + if [ -n "$reason" ]; then + printf 'invalid\n' + return 0 + fi + real=$(jq -r ' + (._scheduling.admission_control // null) + | if . == null then "absent" + else ([keys_unsorted[] | select(startswith("_") | not)] | length | tostring) + end' "$file" 2>/dev/null) || { printf 'absent\n'; return 0; } + case "$real" in + absent) printf 'absent\n'; return 0 ;; + 0) printf 'inert\n'; return 0 ;; + esac + enabled=$(jq -r '._scheduling.admission_control.enabled // false' "$file" 2>/dev/null) + if [ "$enabled" = true ]; then printf 'active\n'; else printf 'inert\n'; fi +} + +# Print the canonical admission policy object for , or "null". +fm_admission_object() { # + local file=$1 + if [ ! -f "$file" ] || ! command -v jq >/dev/null 2>&1; then + printf 'null\n' + return 0 + fi + jq -S -c '._scheduling.admission_control // null' "$file" 2>/dev/null || printf 'null\n' +} + +# Print a stable short digest of an exact byte string, so every decision can name +# the configuration and fleet-view identities it acted on. +fm_admission_digest() { # + local json=$1 sum + if command -v sha256sum >/dev/null 2>&1; then + sum=$(printf '%s' "$json" | sha256sum | awk '{print $1}') + elif command -v shasum >/dev/null 2>&1; then + sum=$(printf '%s' "$json" | shasum -a 256 | awk '{print $1}') + else + printf 'unavailable\n' + return 0 + fi + printf 'sha256:%s\n' "${sum:0:16}" +} diff --git a/bin/fm-admission.sh b/bin/fm-admission.sh new file mode 100755 index 00000000000..0fa91e82bc2 --- /dev/null +++ b/bin/fm-admission.sh @@ -0,0 +1,422 @@ +#!/usr/bin/env bash +# fm-admission.sh - read-only fleet admission decision from a fresh census. +# +# Admission control is the third layer above routing and scheduling: it answers +# "should the fleet accept another task right now?" from properties of the FLEET +# ALONE. It is deliberately task-independent - this command takes no task +# argument and refuses one - so the same snapshot returns the same band for every +# incoming task. Anything that varies per task (tier, project, model, priority, +# urgency, token estimate, file overlap) belongs to routing or scheduling. +# +# The command is read-only: it does not acquire the session lock, hold or unhold +# backlog items, spawn, mutate task state, or write any record. It observes and +# explains; the caller applies the outcome. The decision procedure for each band +# is owned by .agents/skills/fleet-admission/SKILL.md, and the policy schema by +# docs/configuration.md "Fleet admission control". +# +# Usage: +# fm-admission.sh human-readable band and per-rule explanation +# fm-admission.sh --brief one summary line (session-start/teardown use) +# fm-admission.sh --json the decision record (schema fm-admission.v1) +# fm-admission.sh --snapshot evaluate an already-taken fleet snapshot +# instead of running fm-fleet-snapshot.sh; +# its age is then checked against the +# configured freshness limit +# fm-admission.sh validate validate policy only; print the reason on +# stdout and exit 2 when malformed +# fm-admission.sh --help print this header +# +# Exit status is the band, so a caller that ignores the output still stops safely: +# 0 admit - preferred, or admission is not configured for this home +# 2 refuse - policy is malformed or the census could not be evaluated +# 3 defer - soft +# 4 refuse - hard +# +# Every non-preferred result names, for each rule: the observed value, its source +# and freshness, the exact JSON configuration path, the configured value, and the +# resulting band. Numbers are never embedded here; a threshold this file cannot +# read from configuration does not exist. +# +# Stage 1 scope: only deterministic safety conditions enforce (authority, +# census integrity, snapshot freshness). Every other signal is recorded as an +# observation and cannot change the band - see the policy schema's +# enforcement_mode. Backlog consistency is its own signal precisely so an +# unrelated backlog contradiction is reported and repaired without being +# misread as physical fleet saturation. +# +# Telemetry seam: the decision record printed by --json is the unit of admission +# telemetry. When the wake-outcome ledger exposes its extension seam, append this +# record in the ledger owner's format; admission never opens a competing store. +# +# Environment: +# FM_ADMISSION_NOW_EPOCH override the decision clock (tests) +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" + +# shellcheck source=bin/fm-admission-lib.sh +. "$SCRIPT_DIR/fm-admission-lib.sh" +# shellcheck source=bin/fm-session-lock-lib.sh +. "$SCRIPT_DIR/fm-session-lock-lib.sh" + +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$SCRIPT_DIR/fm-admission.sh" +} + +die() { printf 'error: %s\n' "$1" >&2; exit 2; } + +MODE=human +SNAPSHOT_FILE= +VALIDATE_ONLY=0 +while [ $# -gt 0 ]; do + case "$1" in + --json) MODE=json ;; + --brief) MODE=brief ;; + --snapshot) shift; [ $# -gt 0 ] || die "--snapshot needs a path"; SNAPSHOT_FILE=$1 ;; + --snapshot=*) SNAPSHOT_FILE=${1#--snapshot=} ;; + validate) VALIDATE_ONLY=1 ;; + -h|--help) usage; exit 0 ;; + -*) die "unknown option $1" ;; + # A task id here would make admission task-dependent, which is the one thing + # this layer must never be. Refuse loudly rather than silently ignoring it. + *) die "fm-admission.sh takes no task argument: admission inspects fleet state only, never the incoming task" ;; + esac + shift +done + +CONFIG_FILE=$(fm_admission_config_file "$CONFIG") +REASON=$(fm_admission_validate_reason "$CONFIG_FILE") || true +if [ -n "$REASON" ]; then + printf 'admission policy invalid: config/crew-dispatch.json - %s\n' "$REASON" + exit 2 +fi +[ "$VALIDATE_ONLY" -eq 1 ] && exit 0 + +STATE_KIND=$(fm_admission_state "$CONFIG_FILE") +if [ "$STATE_KIND" != active ]; then + case "$MODE" in + json) + printf '{"schema":"fm-admission.v1","record_kind":"admission","active":false,' + printf '"configured":"%s","decision_band":"preferred","action":"admit",' "$STATE_KIND" + printf '"config_path":"/_scheduling/admission_control",' + printf '"reason":"admission control is not configured for this home",' + printf '"notification_band":"silent","task_id":null}\n' + ;; + *) + printf 'admission: not configured (%s) - dispatch is unchanged\n' "$STATE_KIND" + ;; + esac + exit 0 +fi + +command -v jq >/dev/null 2>&1 || die "jq is required to evaluate admission policy" + +NOW_EPOCH=${FM_ADMISSION_NOW_EPOCH:-$(date -u +%s)} +TIMESTAMP=$(date -u -r "$NOW_EPOCH" +%Y-%m-%dT%H:%M:%SZ 2>/dev/null \ + || date -u -d "@$NOW_EPOCH" +%Y-%m-%dT%H:%M:%SZ 2>/dev/null \ + || date -u +%Y-%m-%dT%H:%M:%SZ) + +# --- fresh census ---------------------------------------------------------- +# bin/fm-fleet-snapshot.sh is the single owner of reading task metadata, current +# state, and backlog structure; admission composes its output rather than +# re-parsing state files. +if [ -n "$SNAPSHOT_FILE" ]; then + [ -f "$SNAPSHOT_FILE" ] || die "snapshot file not found: $SNAPSHOT_FILE" + SNAPSHOT=$(cat "$SNAPSHOT_FILE") || die "cannot read snapshot: $SNAPSHOT_FILE" +else + SNAPSHOT=$("$FM_ROOT/bin/fm-fleet-snapshot.sh" --json 2>/dev/null) || SNAPSHOT= +fi + +CENSUS_READABLE=true +if [ -z "$SNAPSHOT" ] || ! printf '%s' "$SNAPSHOT" | jq -e '.schema == "fm-fleet-snapshot.v1"' >/dev/null 2>&1; then + CENSUS_READABLE=false + SNAPSHOT='{"schema":"fm-fleet-snapshot.v1","generated":null,"tasks":[],"backlog":{"records":[]},"main_inventory":{"valid":null,"reason":"census unreadable"}}' +fi + +SNAPSHOT_GENERATED=$(printf '%s' "$SNAPSHOT" | jq -r '.generated // empty') +if [ -n "$SNAPSHOT_GENERATED" ]; then + SNAPSHOT_EPOCH=$(date -u -j -f '%Y-%m-%dT%H:%M:%SZ' "$SNAPSHOT_GENERATED" +%s 2>/dev/null \ + || date -u -d "$SNAPSHOT_GENERATED" +%s 2>/dev/null || printf '') +else + SNAPSHOT_EPOCH= +fi +case "$SNAPSHOT_EPOCH" in + ''|*[!0-9]*) SNAPSHOT_AGE=null ;; + *) SNAPSHOT_AGE=$(( NOW_EPOCH - SNAPSHOT_EPOCH )); [ "$SNAPSHOT_AGE" -lt 0 ] && SNAPSHOT_AGE=0 ;; +esac + +POLICY=$(fm_admission_object "$CONFIG_FILE") +CONFIG_DIGEST=$(fm_admission_digest "$POLICY") +SNAPSHOT_ID=$(fm_admission_digest "$(printf '%s' "$SNAPSHOT" | jq -S -c .)") + +# --- single-primary authority ---------------------------------------------- +# The per-home session lock already serializes intake, so it IS the admission +# authority. No new process, daemon, or reservation store is introduced. +if fm_session_lock_owned_by_self "$STATE"; then + AUTHORITY_HELD=true + AUTHORITY_DETAIL="session lock held by this session" +elif [ -f "$STATE/.lock" ]; then + AUTHORITY_HELD=false + AUTHORITY_DETAIL="session lock is not held by this session" +else + AUTHORITY_HELD=false + AUTHORITY_DETAIL="no session lock is held for this home" +fi + +DECISION_ID=$(fm_admission_digest "$SNAPSHOT_ID|$CONFIG_DIGEST|$TIMESTAMP") + +RECORD=$(printf '%s' "$SNAPSHOT" | jq \ + --argjson policy "$POLICY" \ + --arg timestamp "$TIMESTAMP" \ + --argjson snapshot_age "$SNAPSHOT_AGE" \ + --argjson census_readable "$CENSUS_READABLE" \ + --argjson authority_held "$AUTHORITY_HELD" \ + --arg authority_detail "$AUTHORITY_DETAIL" \ + --arg config_digest "$CONFIG_DIGEST" \ + --arg snapshot_id "$SNAPSHOT_ID" \ + --arg decision_id "$DECISION_ID" \ + --arg home "$FM_HOME" ' + def rank: {"preferred":0,"soft":1,"hard":2}; + def worse($a; $b): if rank[$a] >= rank[$b] then $a else $b end; + def cfg($path): "/_scheduling/admission_control" + $path; + + . as $snap + | $policy as $p + | ($p.signals // {}) as $sig + | ($snap.tasks // []) as $tasks + | ([$tasks[].id] | group_by(.) | map(select(length > 1) | .[0])) as $dupes + | ([$snap.backlog.records[]? | select(.hold_kind == "load")] | length) as $load_holds + | ([$tasks[] | select((.current_state.state // "") == "")] | length) as $unknown_state + | ($snapshot_age != null and ($sig.census_integrity.max_snapshot_age_seconds != null) + and ($snapshot_age > $sig.census_integrity.max_snapshot_age_seconds)) as $stale + | ($snapshot_age == null and ($sig.census_integrity.max_snapshot_age_seconds != null)) as $age_unmeasurable + + # Every rule carries the same five parts: observed value, source and + # freshness, config path, configured value, resulting band. + | ([ + { + rule_id: "authority.single_primary", + signal: "authority.session_lock", + observed: (if $authority_held then "held" else "not-held" end), + detail: $authority_detail, + unit: "state", + source: "per-home-session-lock", + observed_at: $timestamp, + freshness_seconds: 0, + valid: true, + config_path: cfg("/authority/unreachable_band"), + operator: "authority-held", + configured_value: ($p.authority.unreachable_band // "hard"), + result: (if $authority_held then "preferred" else ($p.authority.unreachable_band // "hard") end) + }, + { + rule_id: "census_integrity.inventory", + signal: "census_integrity.inventory_readable", + observed: $census_readable, + unit: "boolean", + source: ($sig.census_integrity.source // "fresh-authority-census"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: $census_readable, + config_path: cfg("/signals/census_integrity/unknown_band"), + operator: "readable", + configured_value: ($sig.census_integrity.unknown_band // $p.unknown_band), + result: (if $census_readable then "preferred" + else ($sig.census_integrity.unknown_band // $p.unknown_band) end) + }, + { + rule_id: "census_integrity.duplicate_identities", + signal: "census_integrity.duplicate_task_ids", + observed: ($dupes | length), + detail: ($dupes | join(", ")), + unit: "count", + source: ($sig.census_integrity.source // "fresh-authority-census"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: $census_readable, + config_path: cfg("/signals/census_integrity/unknown_band"), + operator: "==0", + configured_value: ($sig.census_integrity.unknown_band // $p.unknown_band), + result: (if ($dupes | length) == 0 then "preferred" + else ($sig.census_integrity.unknown_band // $p.unknown_band) end) + }, + { + rule_id: "census_integrity.snapshot_age", + signal: "census_integrity.snapshot_age_seconds", + observed: $snapshot_age, + unit: "seconds", + source: ($sig.census_integrity.source // "fresh-authority-census"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: ($snapshot_age != null), + config_path: cfg("/signals/census_integrity/max_snapshot_age_seconds"), + operator: ">", + configured_value: $sig.census_integrity.max_snapshot_age_seconds, + result: (if $stale or $age_unmeasurable then ($sig.census_integrity.unknown_band // $p.unknown_band) + else "preferred" end) + }, + { + rule_id: "backlog_consistency.main_inventory", + signal: "backlog_consistency.main_inventory_valid", + observed: ($snap.main_inventory.valid), + detail: ($snap.main_inventory.reason // ""), + unit: "boolean", + source: ($sig.backlog_consistency.source // "main-inventory"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: true, + config_path: cfg("/signals/backlog_consistency/enforce"), + operator: "observe", + configured_value: ($sig.backlog_consistency.enforce // false), + result: "preferred", + note: "backlog consistency is a separate health signal; a contradiction here must be repaired but is not physical fleet saturation" + }, + { + rule_id: "admission_queue_pressure.load_hold_depth", + signal: "admission_queue_pressure.queued_count", + observed: $load_holds, + unit: "count", + source: ($sig.admission_queue_pressure.source // "tasks-axi-load-holds-plus-ledger"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: true, + config_path: cfg("/signals/admission_queue_pressure/enforce"), + operator: "observe", + configured_value: ($sig.admission_queue_pressure.enforce // false), + result: "preferred" + }, + { + rule_id: "admission_queue_pressure.oldest_wait", + signal: "admission_queue_pressure.oldest_wait_seconds", + observed: null, + unit: "seconds", + source: ($sig.admission_queue_pressure.source // "tasks-axi-load-holds-plus-ledger"), + observed_at: $timestamp, + freshness_seconds: null, + valid: false, + unmeasured_reason: "backlog age is task age, not admission wait age; wait age needs the wake-outcome ledger", + config_path: cfg("/signals/admission_queue_pressure/oldest_wait_soft_seconds"), + operator: "observe", + configured_value: $sig.admission_queue_pressure.oldest_wait_soft_seconds, + result: "preferred" + }, + { + rule_id: "active_workers.count", + signal: "active_workers.count", + observed: ($tasks | length), + detail: ([$tasks[] | .kind // "crewmate"] | group_by(.) | map("\(.[0])=\(length)") | join(" ")), + unit: "count", + source: ($sig.active_workers.source // "fresh-authority-census"), + observed_at: ($snap.generated // $timestamp), + freshness_seconds: $snapshot_age, + valid: $census_readable, + unknown_state_count: $unknown_state, + config_path: cfg("/signals/active_workers/enforce"), + operator: "observe", + configured_value: ($sig.active_workers.enforce // false), + result: "preferred", + note: "observation only; an ambiguous worker is counted as present, never dropped" + } + ] + + [ ("coordination_debt", "host_resources", "reservation_pressure") + | . as $name + | { + rule_id: "\($name).unavailable", + signal: $name, + observed: null, + unit: "none", + source: ($sig[$name].source // "unmeasured"), + observed_at: $timestamp, + freshness_seconds: null, + valid: false, + unmeasured_reason: "signal source is not collectable in this home yet", + config_path: cfg("/signals/\($name)/enabled"), + operator: "observe", + configured_value: ($sig[$name].enabled // false), + result: "preferred" + } + ]) as $rules + | ($rules | map(.result) | reduce .[] as $r ("preferred"; worse(.; $r))) as $band + | ($rules | map(select(.result == $band and $band != "preferred") | .rule_id)) as $controlling + | { + schema: "fm-admission.v1", + record_kind: "admission", + active: true, + decision_id: $decision_id, + timestamp: $timestamp, + task_id: null, + task_independent: true, + fleet_id: $p.fleet_id, + authority_id: ($p.authority.authority_id // $home), + node_id: $home, + home: $home, + snapshot_id: $snapshot_id, + snapshot_generated: $snap.generated, + snapshot_freshness_seconds: $snapshot_age, + config_digest: $config_digest, + authority_held: $authority_held, + enforcement_mode: $p.enforcement_mode, + combine: $p.combine, + decision_band: $band, + action: ($p.bands[$band].action // "admit"), + hold_kind: ($p.bands[$band].hold_kind // null), + auto_reconsider: ($p.bands[$band].auto_reconsider // null), + notification_band: (if $band == "preferred" then "silent" else "immediate" end), + active_worker_count: ($tasks | length), + unknown_state_count: $unknown_state, + load_queue_depth: $load_holds, + oldest_load_wait_seconds: null, + override_authority: null, + release_triggers: ($p.queue.release_triggers // []), + rules: ($rules | map(. + {signal_band: .result, fleet_band: $band})), + controlling_rules: $controlling, + telemetry: { + sink: ($p.telemetry.sink // null), + integrated: false, + reason: "wake-outcome ledger extension seam is not available yet; this record is the unit to append when it lands" + } + } +') || die "admission evaluation failed" + +BAND=$(printf '%s' "$RECORD" | jq -r '.decision_band') + +case "$MODE" in + json) + printf '%s\n' "$RECORD" + ;; + brief) + printf '%s' "$RECORD" | jq -r ' + "admission: \(.decision_band) (\(.action); \(.active_worker_count) worker(s), " + + "\(.load_queue_depth) load-held; authority \(if .authority_held then "held" else "not held" end); config \(.config_digest))" + + (if .decision_band == "preferred" then "" else "\n controlling: \(.controlling_rules | join(", "))" end)' + ;; + *) + printf '%s' "$RECORD" | jq -r ' + "admission: \(.decision_band) -> \(.action) fleet=\(.fleet_id) authority=\(.authority_id) config=\(.config_digest)", + "snapshot: \(.snapshot_id) generated=\(.snapshot_generated) age=\(.snapshot_freshness_seconds)s mode=\(.enforcement_mode) combine=\(.combine)", + "", + (.rules[] + | " \(.rule_id): observed=\(.observed | tostring)\(if (.unit // "none") == "none" then "" else " \(.unit)" end)" + + " source=\(.source)/age=\(if .freshness_seconds == null then "unknown" else "\(.freshness_seconds)s" end) valid=\(.valid)" + + " \(.config_path) \(.operator) \(.configured_value | tostring)" + + " -> \(.signal_band)"), + "", + "notification: \(.notification_band)" + + (if .decision_band == "preferred" then "" else " controlling: \(.controlling_rules | join(", "))" end)' + ;; +esac + +case "$BAND" in + preferred) exit 0 ;; + soft) exit 3 ;; + hard) exit 4 ;; + *) exit 2 ;; +esac diff --git a/bin/fm-backend.sh b/bin/fm-backend.sh index e505b99f757..f404fbe384d 100644 --- a/bin/fm-backend.sh +++ b/bin/fm-backend.sh @@ -240,8 +240,14 @@ fm_backend_detect_cmux_app_is_ancestor() { # notice names the winning signal, so a fallback-detected cmux (bundle id or # ancestry, after the claude wrapper stripped CMUX_WORKSPACE_ID) is visibly # distinct from the primary-marker case. -fm_backend_name() { - local line v detected marker +# fm_backend_configured_name: the backend EXPLICITLY selected for this home - +# FM_BACKEND env, then config/backend's first non-empty line - or empty when +# nothing is configured and fm_backend_name would fall through to runtime +# auto-detection. Local reads only: it never touches a backend server, so a +# caller that must stay off every socket (bin/fm-launch.sh's front door) can +# consult it before painting anything. +fm_backend_configured_name() { + local line v if [ -n "${FM_BACKEND:-}" ]; then printf '%s' "$FM_BACKEND" return 0 @@ -255,6 +261,16 @@ fm_backend_name() { fi done < "$FM_BACKEND_CONFIG_DIR/backend" fi + return 0 +} + +fm_backend_name() { + local configured detected marker + configured=$(fm_backend_configured_name) + if [ -n "$configured" ]; then + printf '%s' "$configured" + return 0 + fi # Called directly (not in a command substitution) so the detect signal # globals survive into the notice below. if fm_backend_detect >/dev/null; then diff --git a/bin/fm-backlog-handoff.sh b/bin/fm-backlog-handoff.sh index 29ab9cd2d66..e83a857e8a3 100755 --- a/bin/fm-backlog-handoff.sh +++ b/bin/fm-backlog-handoff.sh @@ -42,7 +42,12 @@ # own hand-editing of its own backlog, not this validated helper. Idempotent: # re-running converges. Atomic: on any move failure nothing moves. # See AGENTS.md project management and task lifecycle. +# Remote routes use an outbox handoff: one atomic local tasks-axi mv removes the +# selected set from the dispatchable backlog into data/handoff/.outbox.md, +# then an idempotent confined transfer and fm-backlog-receive.sh deliver it. +# A present outbox is the whole recovery record. No two-phase journal exists. # Usage: fm-backlog-handoff.sh ... +# fm-backlog-handoff.sh --resume-pending set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -55,10 +60,39 @@ MAIN_BACKLOG="$DATA/backlog.md" . "$SCRIPT_DIR/fm-tasks-axi-lib.sh" # shellcheck source=bin/fm-secondmate-registry-lib.sh . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" +# shellcheck source=bin/fm-wake-lib.sh +. "$SCRIPT_DIR/fm-wake-lib.sh" -[ $# -ge 2 ] || { echo "usage: fm-backlog-handoff.sh ..." >&2; exit 1; } -ID=$1 -shift +ACTIVE_HANDOFF_LOCK= +ACTIVE_REGISTRY_LOCK= +release_remote_locks() { + if [ -n "$ACTIVE_HANDOFF_LOCK" ]; then + fm_lock_release "$ACTIVE_HANDOFF_LOCK" + ACTIVE_HANDOFF_LOCK= + fi + if [ -n "$ACTIVE_REGISTRY_LOCK" ]; then + fm_lock_release "$ACTIVE_REGISTRY_LOCK" + ACTIVE_REGISTRY_LOCK= + fi +} +trap release_remote_locks EXIT +trap 'exit 1' HUP INT TERM + +sha256_file() { + if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi +} + +RESUME_PENDING=0 +if [ "${1:-}" = --resume-pending ]; then + [ "$#" -eq 1 ] || { echo "usage: fm-backlog-handoff.sh --resume-pending" >&2; exit 1; } + RESUME_PENDING=1 + ID= + shift +else + [ "$#" -ge 2 ] || { echo "usage: fm-backlog-handoff.sh ..." >&2; exit 1; } + ID=$1 + shift +fi secondmate_home() { local id=$1 home @@ -227,6 +261,214 @@ backlog_key_noncanonical_body_lines() { ' "$file" } +seed_backlog_scaffold() { # + mkdir -p "$(dirname "$1")" + [ -f "$1" ] || printf '## In flight\n\n## Queued\n\n## Done\n' > "$1" +} + +outbox_item_count() { # + awk '/^- \[[ x]\] / { count++ } END { print count + 0 }' "$1" +} + +remote_deliver_outbox() { # + local id=$1 outbox=$2 remote_rel receive_out snapshot bytes hash generation counter counter_tmp current + [ -f "$outbox" ] && [ ! -L "$outbox" ] || { + echo "error: pending outbox is unavailable or unsafe: $outbox" >&2 + return 1 + } + snapshot=$(umask 077; mktemp "${TMPDIR:-/tmp}/fm-handoff-payload.XXXXXX") || return 1 + if ! cp -p -- "$outbox" "$snapshot"; then + rm -f -- "$snapshot" + return 1 + fi + bytes=$(LC_ALL=C wc -c < "$snapshot" | tr -d ' ') + hash=$(sha256_file "$snapshot") || { rm -f -- "$snapshot"; return 1; } + counter="$STATE/.remote-handoff-$id.generation" + current=0 + if [ -e "$counter" ] || [ -L "$counter" ]; then + [ -f "$counter" ] && [ ! -L "$counter" ] || { rm -f -- "$snapshot"; return 1; } + IFS= read -r current < "$counter" || { rm -f -- "$snapshot"; return 1; } + case "$current" in ''|*[!0-9]*) rm -f -- "$snapshot"; return 1 ;; esac + [ "${#current}" -le 17 ] || { rm -f -- "$snapshot"; return 1; } + fi + generation=$((current + 1)) + counter_tmp=$(umask 077; mktemp "$STATE/.remote-handoff-generation.XXXXXX") \ + || { rm -f -- "$snapshot"; return 1; } + printf '%s\n' "$generation" > "$counter_tmp" \ + || { rm -f -- "$snapshot" "$counter_tmp"; return 1; } + chmod 600 "$counter_tmp" \ + || { rm -f -- "$snapshot" "$counter_tmp"; return 1; } + mv -f -- "$counter_tmp" "$counter" \ + || { rm -f -- "$snapshot" "$counter_tmp"; return 1; } + remote_rel="state/handoff/$id.outbox.md" + if ! "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-file.sh put "$remote_rel" 1048576 \ + "$bytes" "$hash" "$generation" < "$snapshot"; then + rm -f -- "$snapshot" + echo "error: handoff transfer to $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 + return 1 + fi + rm -f -- "$snapshot" + if ! receive_out=$("$SCRIPT_DIR/fm-on.sh" "$id" fm-backlog-receive.sh \ + "$remote_rel" "$bytes" "$hash" "$generation" < /dev/null 2>&1); then + [ -z "$receive_out" ] || printf '%s\n' "$receive_out" >&2 + echo "error: handoff receipt by $id was unavailable or completion is unknown; outbox preserved at $outbox" >&2 + return 1 + fi + rm -f -- "$outbox" || { + echo "error: remote receipt was confirmed but local outbox cleanup failed: $outbox" >&2 + return 1 + } + printf '%s\n' "$receive_out" +} + +remove_interrupted_source_duplicates() { # + local outbox=$1 key progress remaining pass=0 + shift + while :; do + remaining=0 + progress=0 + for key in "$@"; do + backlog_key_section "$outbox" "$key" >/dev/null 2>&1 || continue + if backlog_key_section "$MAIN_BACKLOG" "$key" >/dev/null 2>&1; then + remaining=$((remaining + 1)) + if tasks-axi rm "$key" --file "$MAIN_BACKLOG" >/dev/null 2>&1; then + progress=$((progress + 1)) + fi + fi + done + [ "$remaining" -gt 0 ] || return 0 + [ "$progress" -gt 0 ] || { + echo "error: could not complete interrupted source removal; outbox remains authoritative at $outbox" >&2 + return 1 + } + pass=$((pass + 1)) + [ "$pass" -le "$#" ] || return 1 + done +} + +remote_handoff() { # + local id=$1 outbox section main_section out_section key mv_out + local -a requested to_move already missing in_flight done_items not_queued + shift + requested=("$@") + outbox="$DATA/handoff/$id.outbox.md" + validate_backlog_file "main backlog" "$MAIN_BACKLOG" || return 1 + validate_backlog_file "remote handoff outbox" "$outbox" || return 1 + fm_tasks_axi_compatible || { + echo "error: tasks-axi with atomic multi-ID mv support (0.2.2+) is required to stage remote handoffs" >&2 + return 1 + } + to_move=() + already=() + missing=() + in_flight=() + done_items=() + not_queued=() + for key in "${requested[@]}"; do + out_section=$(backlog_key_section "$outbox" "$key" 2>/dev/null || true) + main_section=$(backlog_key_section "$MAIN_BACKLOG" "$key" 2>/dev/null || true) + if [ -n "$out_section" ]; then + [ "$out_section" = '## Queued' ] || not_queued+=("$key") + already+=("$key") + continue + fi + case "$main_section" in + '## Queued') to_move+=("$key") ;; + '## In flight') in_flight+=("$key") ;; + '## Done') done_items+=("$key") ;; + '') missing+=("$key") ;; + *) not_queued+=("$key") ;; + esac + done + if [ "${#in_flight[@]}" -gt 0 ] || [ "${#done_items[@]}" -gt 0 ] \ + || [ "${#not_queued[@]}" -gt 0 ] || [ "${#missing[@]}" -gt 0 ]; then + [ "${#in_flight[@]}" -eq 0 ] || echo "error: refusing to hand off in-flight backlog items: ${in_flight[*]}" >&2 + [ "${#done_items[@]}" -eq 0 ] || echo "error: refusing to hand off Done backlog items: ${done_items[*]}" >&2 + [ "${#not_queued[@]}" -eq 0 ] || echo "error: refusing to hand off non-Queued outbox or backlog items: ${not_queued[*]}" >&2 + [ "${#missing[@]}" -eq 0 ] || echo "error: no backlog or pending outbox item matched: ${missing[*]}" >&2 + echo " nothing new was staged." >&2 + return 1 + fi + for key in "${to_move[@]}"; do + while IFS= read -r line; do + printf 'error: refusing to hand off %s: non-2-space continuation line: %s\n' "$key" "$line" >&2 + return 1 + done < <(backlog_key_noncanonical_body_lines "$MAIN_BACKLOG" "$key") + done + seed_backlog_scaffold "$outbox" + if [ "${#to_move[@]}" -gt 0 ]; then + if ! mv_out=$(tasks-axi mv "${to_move[@]}" --file "$MAIN_BACKLOG" --to "$outbox" 2>&1); then + [ -z "$mv_out" ] || printf '%s\n' "$mv_out" >&2 + echo "error: atomic outbox staging failed; nothing new was handed off" >&2 + return 1 + fi + fi + # A hard local kill can land tasks-axi's target persist before its source + # persist. The outbox is already authoritative in that state, so converge by + # deleting only duplicates that tasks-axi itself confirms are dependency-safe. + remove_interrupted_source_duplicates "$outbox" "${requested[@]}" || return 1 + remote_deliver_outbox "$id" "$outbox" || return 1 + echo "handed off ${#requested[@]} item(s) to remote secondmate $id: ${requested[*]}" + [ "${#already[@]}" -eq 0 ] || echo " already staged (recovered): ${already[*]}" +} + +with_remote_route_locks() { # + local id=$1 operation=$2 rc + shift 2 + case "$id" in ''|*[!A-Za-z0-9._-]*) echo "error: unsafe remote handoff id: $id" >&2; return 1 ;; esac + ACTIVE_REGISTRY_LOCK=$(secondmate_registry_lock_path "$STATE") + fm_lock_acquire_wait "$ACTIVE_REGISTRY_LOCK" + if [ "$(secondmate_registry_field "$REG" "$id" remote 2>/dev/null || true)" != 1 ]; then + echo "error: pending outbox has no matching remote secondmate route: $id" >&2 + release_remote_locks + return 1 + fi + ACTIVE_HANDOFF_LOCK="$STATE/.backlog-handoff-$id.lock" + fm_lock_acquire_wait "$ACTIVE_HANDOFF_LOCK" + if "$operation" "$@"; then rc=0; else rc=$?; fi + release_remote_locks + return "$rc" +} + +resume_remote_outbox() { # + local id=$1 outbox=$2 + [ -e "$outbox" ] || [ -L "$outbox" ] || return 0 + if [ ! -f "$outbox" ] || [ -L "$outbox" ]; then + echo "error: unsafe pending handoff outbox: $outbox" >&2 + return 1 + fi + remote_deliver_outbox "$id" "$outbox" +} + +resume_pending_outboxes() { + local outbox id failed=0 + [ -d "$DATA/handoff" ] || return 0 + for outbox in "$DATA/handoff"/*.outbox.md; do + [ -e "$outbox" ] || [ -L "$outbox" ] || continue + id=$(basename "$outbox" .outbox.md) + case "$id" in ''|*[!A-Za-z0-9._-]*) echo "error: unsafe pending handoff id: $id" >&2; failed=1; continue ;; esac + with_remote_route_locks "$id" resume_remote_outbox "$id" "$outbox" || failed=1 + done + return "$failed" +} + +if [ "$RESUME_PENDING" -eq 1 ]; then + resume_pending_outboxes + exit $? +fi + +ACTIVE_REGISTRY_LOCK=$(secondmate_registry_lock_path "$STATE") +fm_lock_acquire_wait "$ACTIVE_REGISTRY_LOCK" +REMOTE=$(secondmate_registry_field "$REG" "$ID" remote 2>/dev/null || true) +if [ "$REMOTE" = 1 ]; then + ACTIVE_HANDOFF_LOCK="$STATE/.backlog-handoff-$ID.lock" + fm_lock_acquire_wait "$ACTIVE_HANDOFF_LOCK" + if remote_handoff "$ID" "$@"; then rc=0; else rc=$?; fi + release_remote_locks + exit "$rc" +fi +release_remote_locks + RAW_HOME=$(secondmate_home "$ID") || exit 1 [ -n "$RAW_HOME" ] || { echo "error: secondmate $ID has no home in $REG" >&2; exit 1; } SUB_HOME=$(validate_secondmate_home "$ID" "$RAW_HOME") || exit 1 diff --git a/bin/fm-backlog-receive.sh b/bin/fm-backlog-receive.sh new file mode 100755 index 00000000000..b2aec10e104 --- /dev/null +++ b/bin/fm-backlog-receive.sh @@ -0,0 +1,187 @@ +#!/usr/bin/env bash +# Receive one delivered remote-secondmate outbox into this home's backlog. +# +# Usage: +# fm-backlog-receive.sh state/handoff/.outbox.md +# +# The delivered file must be a non-symlink backlog-format scratch file confined +# to FM_HOME/state/handoff. Every item must be Queued. Keys already present in +# data/backlog.md are skipped; every remaining key moves in one dependency-closed +# `tasks-axi mv` transaction under tasks-axi's own locks. On an ambiguous caller +# retry, destination-present classification makes this operation idempotent. +# +# If tasks-axi reports a lock failure, this host may remove and retry once only +# for its own backlog or delivered lock whose pid is dead and whose mtime is at +# least 30 seconds old. No live or uncertain lock is touched. On confirmed +# receipt the delivered scratch file is removed; no other path is deletable. +set -eu + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +DEST="$FM_HOME/data/backlog.md" +LOCK_STALE_SECS=30 + +# shellcheck source=bin/fm-tasks-axi-lib.sh +. "$SCRIPT_DIR/fm-tasks-axi-lib.sh" +# shellcheck source=bin/fm-wake-lib.sh +. "$SCRIPT_DIR/fm-wake-lib.sh" + +die() { printf 'error: %s\n' "$1" >&2; exit 1; } +usage() { sed -n '2,16p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } +sha256_file() { + if command -v shasum >/dev/null 2>&1; then shasum -a 256 "$1" | awk '{print $1}'; else sha256sum "$1" | awk '{print $1}'; fi +} + +backlog_key_section() { # + awk -v key="$2" ' + BEGIN { section = "## Queued" } + /^##[[:space:]]+/ { section=$0; sub(/^##[[:space:]]+/, "## ", section); sub(/[[:space:]]+$/, "", section); next } + /^- \[[ x]\] / { + rest=$0; sub(/^- \[[ x]\] +/, "", rest); id=rest; sub(/[ \t].*/, "", id) + if (id == key) { print section; found=1; exit } + } + END { exit found ? 0 : 1 } + ' "$1" +} + +list_keys() { # + awk ' + /^- \[[ x]\] / { + rest=$0; sub(/^- \[[ x]\] +/, "", rest); id=rest; sub(/[ \t].*/, "", id) + if (id != "" && !seen[id]++) print id + } + ' "$1" +} + +lock_age() { + local modified now + if [ "$(uname 2>/dev/null)" = Darwin ]; then + modified=$(stat -f '%m' "$1" 2>/dev/null) || return 1 + else + modified=$(stat -c '%Y' "$1" 2>/dev/null) || return 1 + fi + now=$(date +%s) || return 1 + case "$modified$now" in *[!0-9]*) return 1 ;; esac + printf '%s\n' "$((now - modified))" +} + +remove_dead_stale_lock() { # + local lock=$1 token pid age + [ -f "$lock" ] && [ ! -L "$lock" ] || return 1 + IFS= read -r token < "$lock" || return 1 + pid=${token%%:*} + case "$pid" in ''|*[!0-9]*) return 1 ;; esac + kill -0 "$pid" 2>/dev/null && return 1 + age=$(lock_age "$lock") || return 1 + [ "$age" -ge "$LOCK_STALE_SECS" ] || return 1 + rm -f -- "$lock" +} + +run_move() { # + tasks-axi mv "$@" --file "$DELIVERED" --to "$DEST" +} + +[ "$#" -eq 4 ] || usage +REL=$1 +EXPECTED_BYTES=$2 +EXPECTED_HASH=$3 +GENERATION=$4 +case "$EXPECTED_BYTES" in ''|*[!0-9]*) die "expected bytes must be a nonnegative integer" ;; esac +[ "${#EXPECTED_BYTES}" -le 10 ] || die "expected bytes are outside the supported range" +[ "$EXPECTED_BYTES" -le 1048576 ] || die "expected bytes are outside the supported range" +case "$EXPECTED_HASH" in ''|*[!A-Fa-f0-9]*) die "expected SHA-256 is invalid" ;; esac +[ "${#EXPECTED_HASH}" -eq 64 ] || die "expected SHA-256 has the wrong length" +EXPECTED_HASH=$(printf '%s' "$EXPECTED_HASH" | tr 'A-F' 'a-f') +case "$GENERATION" in ''|*[!0-9]*) die "generation must be a positive integer" ;; esac +[ "${#GENERATION}" -le 18 ] && [ "$GENERATION" -ge 1 ] || die "generation is outside the supported range" +case "$REL" in state/handoff/*.outbox.md) ;; *) die "delivered outbox path is outside state/handoff: $REL" ;; esac +case "/$REL/" in */../*|*/./*) die "delivered outbox path contains traversal" ;; esac +case "$REL" in *'//'*) die "delivered outbox path is malformed" ;; esac +[ -f "$FM_HOME/.fm-secondmate-home" ] && [ ! -L "$FM_HOME/.fm-secondmate-home" ] \ + || die "FM_HOME is not a seeded secondmate home" +[ -f "$FM_HOME/AGENTS.md" ] && [ -d "$FM_HOME/bin" ] || die "FM_HOME is not a Firstmate home" +HOME_REAL=$(CDPATH='' cd -- "$FM_HOME" 2>/dev/null && pwd -P) || die "FM_HOME cannot be resolved" +PARENT=$(dirname "$FM_HOME/$REL") +PARENT_REAL=$(CDPATH='' cd -- "$PARENT" 2>/dev/null && pwd -P) || die "delivered outbox parent is unavailable" +case "$PARENT_REAL" in "$HOME_REAL/state/handoff") ;; *) die "delivered outbox escapes the remote scratch directory" ;; esac +DELIVERED="$PARENT_REAL/$(basename "$REL")" +NAME=$(basename "$REL") +ID=${NAME%.outbox.md} +case "$ID" in ''|*[!A-Za-z0-9._-]*) die "delivered outbox id is unsafe" ;; esac +TRANSFER_LOCK="$PARENT_REAL/.$ID.upload.lock" +fm_lock_acquire_wait "$TRANSFER_LOCK" || die "cannot lock delivered outbox" +trap 'fm_lock_release "$TRANSFER_LOCK" || true' EXIT +[ -f "$DELIVERED" ] && [ ! -L "$DELIVERED" ] || die "delivered outbox is not a non-symlink regular file" +GENERATION_FILE="$PARENT_REAL/.$ID.upload-generation" +[ -f "$GENERATION_FILE" ] && [ ! -L "$GENERATION_FILE" ] || die "delivered outbox generation is unavailable or unsafe" +{ + IFS= read -r STORED_GENERATION \ + && IFS= read -r STORED_BYTES \ + && IFS= read -r STORED_HASH \ + && ! IFS= read -r +} < "$GENERATION_FILE" || die "delivered outbox generation is malformed" +case "$STORED_GENERATION" in ''|*[!0-9]*) die "delivered outbox generation is malformed" ;; esac +[ "${#STORED_GENERATION}" -le 18 ] || die "delivered outbox generation is malformed" +case "$STORED_BYTES" in ''|*[!0-9]*) die "delivered outbox generation is malformed" ;; esac +case "$STORED_HASH" in ''|*[!A-Fa-f0-9]*) die "delivered outbox generation is malformed" ;; esac +[ "${#STORED_HASH}" -eq 64 ] || die "delivered outbox generation is malformed" +[ "$STORED_GENERATION" = "$GENERATION" ] \ + && [ "$STORED_BYTES" = "$EXPECTED_BYTES" ] \ + && [ "$STORED_HASH" = "$EXPECTED_HASH" ] \ + || die "delivered outbox generation is superseded or conflicting" +ACTUAL_BYTES=$(LC_ALL=C wc -c < "$DELIVERED" | tr -d ' ') +[ "$ACTUAL_BYTES" -eq "$EXPECTED_BYTES" ] || die "delivered outbox length does not match its commitment" +ACTUAL_HASH=$(sha256_file "$DELIVERED") || die "cannot hash delivered outbox" +[ "$ACTUAL_HASH" = "$EXPECTED_HASH" ] || die "delivered outbox digest does not match its commitment" +[ ! -L "$DEST" ] || die "destination backlog must not be a symlink" +if [ -e "$DEST" ] && [ ! -f "$DEST" ]; then die "destination backlog is not a regular file"; fi + +KEYS=() +while IFS= read -r key; do + [ -n "$key" ] && KEYS+=("$key") +done < <(list_keys "$DELIVERED") +for key in "${KEYS[@]}"; do + section=$(backlog_key_section "$DELIVERED" "$key") || die "delivered key disappeared during classification: $key" + [ "$section" = '## Queued' ] || die "delivered outbox contains non-Queued item $key under $section" +done + +mkdir -p "$FM_HOME/data" +DEST_CREATED=0 +if [ ! -f "$DEST" ]; then + printf '## In flight\n\n## Queued\n\n## Done\n' > "$DEST" + DEST_CREATED=1 +fi +TO_MOVE=() +ALREADY=() +for key in "${KEYS[@]}"; do + if backlog_key_section "$DEST" "$key" >/dev/null 2>&1; then + ALREADY+=("$key") + else + TO_MOVE+=("$key") + fi +done + +if [ "${#TO_MOVE[@]}" -gt 0 ]; then + fm_tasks_axi_compatible || die "tasks-axi 0.2.2+ is required for atomic backlog receipt" + if ! MOVE_OUT=$(run_move "${TO_MOVE[@]}" 2>&1); then + recovered=0 + for lock in "$DELIVERED.lock" "$DEST.lock"; do + if remove_dead_stale_lock "$lock"; then recovered=1; fi + done + if [ "$recovered" -ne 1 ] || ! MOVE_OUT=$(run_move "${TO_MOVE[@]}" 2>&1); then + [ "$DEST_CREATED" -eq 0 ] || rm -f -- "$DEST" + [ -z "$MOVE_OUT" ] || printf '%s\n' "$MOVE_OUT" >&2 + die "atomic backlog receipt failed; delivered outbox is preserved for retry" + fi + fi +fi + +for key in "${KEYS[@]}"; do + backlog_key_section "$DEST" "$key" >/dev/null 2>&1 \ + || die "receipt verification failed for $key; delivered outbox is preserved" +done +rm -f -- "$DELIVERED" || die "receipt succeeded but delivered scratch cleanup failed" +fm_lock_release "$TRANSFER_LOCK" || die "receipt succeeded but transfer lock cleanup failed" +trap - EXIT +printf 'received: %s moved=%s already=%s\n' "$(basename "$REL" .outbox.md)" "${#TO_MOVE[@]}" "${#ALREADY[@]}" diff --git a/bin/fm-bootstrap.sh b/bin/fm-bootstrap.sh index 0ac06b162f5..7d9e41fbca2 100755 --- a/bin/fm-bootstrap.sh +++ b/bin/fm-bootstrap.sh @@ -10,6 +10,12 @@ # "BACKEND_INVALID: (known: )", # "STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - ", # "CREW_DISPATCH: invalid config/crew-dispatch.json - ", +# "MODEL_REGISTRY: invalid config/models.json - ", +# "MODEL_REGISTRY: ", +# "MODEL_PRICE: ", +# "MODEL_VERIFY: ", +# "ADMISSION_CONTROL: invalid config/crew-dispatch.json +# _scheduling.admission_control - ", # "FLEET_SYNC: : skipped|recovered|STUCK: ", # "PR_CHECK_MIGRATION: ", # "TANGLE: ", @@ -17,24 +23,27 @@ # "NUDGE_SECONDMATES: secondmate : send failed: ", # "BOOTSTRAP_INFO: nudged fm- with ''", # "SECONDMATE_LIVENESS: secondmate : skipped: |respawn failed after : ", +# "SECONDMATE_HANDOFF: secondmate : pending delivery: item(s)", # "FMX: X mode on ..." or "FMX: X mode off ...". -# When a RUNNING secondmate worktree is fast-forwarded to firstmate's -# own current default-branch commit (a purely LOCAL fast-forward, never -# an origin fetch) AND its loaded instruction surface (AGENTS.md, bin/, -# or .agents/skills/) actually changed, bootstrap immediately nudges it +# When a RUNNING local secondmate worktree is fast-forwarded to +# firstmate's own current default-branch commit, that update is a +# purely local fast-forward and never an origin fetch. Remote routes +# instead converge the persistent home to their configured remote code +# root. If either placement changes its loaded instruction surface +# (AGENTS.md, bin/, or .agents/skills/), bootstrap immediately nudges it # via FM_HOME= bin/fm-send.sh fm- so meta resolves the -# current backend target and the standard from-firstmate marker is -# applied. A successful send prints one BOOTSTRAP_INFO line with the -# exact target and message sent; a failed send leaves an idempotent -# retry marker under state/.secondmate-nudge-pending/ and prints an -# actionable NUDGE_SECONDMATES line. +# current route and the standard from-firstmate marker is applied. A +# successful send prints one BOOTSTRAP_INFO line with the exact target +# and message sent; a failed send leaves an idempotent retry marker +# under state/.secondmate-nudge-pending/ and prints an actionable +# NUDGE_SECONDMATES line. # Already-current or no-instruction-change homes are silently left alone. # The secondmate sweep also propagates declared inherited local material # into each validated live secondmate home. -# SECONDMATE_SYNC lines report actionable skipped local-HEAD syncs or -# inheritance failures for live secondmate homes, plus quarantine -# diagnostics for divergent shared captain-preference copies; -# no-op/current and successful updates stay quiet. +# SECONDMATE_SYNC lines report actionable skipped placement-specific +# syncs or inheritance failures for live secondmate homes, plus +# quarantine diagnostics for divergent shared captain-preference +# copies; no-op/current and successful updates stay quiet. # SECONDMATE_LIVENESS lines report only actionable failures from the # recovery-grade state owned by bin/fm-backend.sh's # fm_backend_agent_state: skipped distinguishes an existing ambiguous @@ -49,8 +58,11 @@ # "treehouse get --lease" support. # no-mistakes is also MISSING when its installed version is older than # 1.31.2. +# gh-axi is also MISSING when its installed version is older than +# 0.1.29, the first release whose bare --squash shorthand works for +# firstmate's non-interactive PR merge path. # tasks-axi and quota-axi are required bootstrap tools (same class as -# lavish-axi). tasks-axi is also version and feature gated (0.1.1+ +# lavish-axi). tasks-axi is also version and feature gated (0.2.2+ # with update --archive-body and mv [...]); an installed but # incompatible build reports MISSING like no-mistakes. A compatible # tasks-axi default backend is silent. quota-axi is required for the @@ -64,6 +76,13 @@ # guesses at malformed or unsafe existing files, and secondmate homes # await the primary-authoritative inherited value instead of creating # their own. +# The optional `_scheduling.admission_control` block inside the same +# config/crew-dispatch.json is validated against the fleet-admission +# schema owned by bin/fm-admission-lib.sh. A home with no such block, or +# one carrying only `_`-prefixed notes, is silent and unaffected; a +# malformed or unknown-field policy prints ADMISSION_CONTROL so the +# policy stops safely instead of being silently ignored. This is a +# read-only detect line and still prints in FM_BOOTSTRAP_DETECT_ONLY=1. # X mode is OPTIONAL and inert unless FM_HOME/.env has a non-empty # FMX_PAIRING_TOKEN. When opted in, bootstrap requires curl+jq, writes # the relay poll shim and 30s cadence config, and prints an FMX line. @@ -76,15 +95,16 @@ # refresh relays any completed fm-fleet-sync.sh output before the # aggregate timeout skip line with timeout and elapsed seconds. # Set FM_FLEET_PRUNE=0 to skip branch pruning during that refresh. -# Set FM_BOOTSTRAP_DETECT_ONLY=1 to skip the five MUTATING sweeps +# Set FM_BOOTSTRAP_DETECT_ONLY=1 to skip the six MUTATING sweeps # (PR-check migration, secondmate_sync, secondmate_liveness_sweep, -# x_mode_setup, fleet_sync) while still printing every read-only detect line +# secondmate_handoff_resume, x_mode_setup, fleet_sync) while still +# printing every read-only detect line # above; the TANGLE line switches to advisory-only wording with no # checkout command. Used by # fm-session-start.sh's read-only path when another live session holds # the fleet lock, so a second concurrent session never race-mutates -# PR-check artifacts, secondmate homes, X-mode artifacts, project -# clones, or repair instructions. +# PR-check artifacts, secondmate homes, pending handoff outboxes, +# X-mode artifacts, project clones, or repair instructions. # Unset/0 (the default) runs every sweep exactly as before - this flag # is purely additive. # fm-bootstrap.sh install ... @@ -108,12 +128,20 @@ DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" . "$SCRIPT_DIR/fm-ff-lib.sh" # shellcheck source=bin/fm-config-inherit-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-config-inherit-lib.sh" +# shellcheck source=bin/fm-secondmate-nudge-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-secondmate-nudge-lib.sh" # shellcheck source=bin/fm-startup-memory-budget-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-startup-memory-budget-lib.sh" # shellcheck source=bin/fm-x-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-x-lib.sh" # shellcheck source=bin/fm-backend.sh disable=SC1091 . "$SCRIPT_DIR/fm-backend.sh" +# shellcheck source=bin/fm-model-registry-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-model-registry-lib.sh" +# shellcheck source=bin/fm-admission-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-admission-lib.sh" +# shellcheck source=bin/fm-remote-readiness-lib.sh disable=SC1091 +. "$SCRIPT_DIR/fm-remote-readiness-lib.sh" fleet_sync_origin_backed_project_count() { local count proj @@ -201,8 +229,8 @@ fleet_sync() { secondmate_sync() { # shellcheck source=bin/fm-wake-lib.sh disable=SC1091 . "$SCRIPT_DIR/fm-wake-lib.sh" - # Local-HEAD secondmate sync: fast-forward every LIVE secondmate home - # to the primary checkout's current default-branch commit. Purely LOCAL - no + # Placement-specific secondmate sync: local homes fast-forward to the primary + # checkout's current default-branch commit. That path is purely LOCAL - no # fetch, no origin dependency: a linked-worktree home already holds the primary's # commit (fm-ff-lib.sh), while a standalone clone without it is skipped until # /updatefirstmate refreshes it from origin. Startup sends reread nudges only @@ -226,32 +254,17 @@ secondmate_sync() { fi FF_NUDGE_WINDOWS="" FF_SEEN_HOMES="" - SECOND_MATE_NUDGE_MESSAGE='firstmate was updated to the latest - please re-read your AGENTS.md to pick up the new instructions.' + SECOND_MATE_NUDGE_MESSAGE=$FM_SECOND_MATE_NUDGE_MESSAGE + REMOTE_SECOND_MATE_NUDGE_MESSAGE=$FM_REMOTE_SECOND_MATE_NUDGE_MESSAGE SECOND_MATE_NUDGE_PENDING_DIR="$STATE/.secondmate-nudge-pending" secondmate_nudge_marker_path() { - case "$1" in - *[!/A-Za-z0-9._-]*|""|*/*) return 1 ;; - esac - printf '%s/%s.pending' "$SECOND_MATE_NUDGE_PENDING_DIR" "$1" + fm_secondmate_nudge_marker_path "$STATE" "$1" } secondmate_write_nudge_marker() { - local id=$1 home=$2 commit=$3 instr=$4 selector marker tmp parent - selector="fm-$id" - marker=$(secondmate_nudge_marker_path "$id") || return 1 - parent=${marker%/*} - mkdir -p "$parent" || return 1 - tmp=$(mktemp "$parent/.nudge.XXXXXX" 2>/dev/null) || return 1 - { - printf 'id=%s\n' "$id" - printf 'selector=%s\n' "$selector" - printf 'home=%s\n' "$home" - printf 'commit=%s\n' "$commit" - printf 'instructions=%s\n' "$instr" - printf 'message=%s\n' "$SECOND_MATE_NUDGE_MESSAGE" - } > "$tmp" || { rm -f "$tmp"; return 1; } - mv -f "$tmp" "$marker" || { rm -f "$tmp"; return 1; } + local id=$1 home=$2 commit=$3 instr=$4 message=${5:-$SECOND_MATE_NUDGE_MESSAGE} remote=${6:-0} + fm_secondmate_nudge_write "$STATE" "$id" "$home" "$commit" "$instr" "$message" "$remote" } secondmate_send_nudge() { @@ -279,7 +292,7 @@ secondmate_sync() { } secondmate_retry_pending_nudges() { - local marker id selector home commit message expected_marker meta meta_home home_real head + local marker id selector home commit message remote expected_marker meta meta_home home_real head out [ -d "$SECOND_MATE_NUDGE_PENDING_DIR" ] || return 0 for marker in "$SECOND_MATE_NUDGE_PENDING_DIR"/*.pending; do [ -f "$marker" ] || continue @@ -296,14 +309,27 @@ secondmate_sync() { home=$(fm_meta_get "$marker" home) commit=$(fm_meta_get "$marker" commit) message=$(fm_meta_get "$marker" message) + remote=$(fm_meta_get "$marker" remote) + [ -n "$remote" ] || remote=0 [ "$selector" = "fm-$id" ] || { echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: retry marker selector mismatch" continue } - [ "$message" = "$SECOND_MATE_NUDGE_MESSAGE" ] || { - echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: retry marker message mismatch" - continue - } + case "$remote" in + 0) [ "$message" = "$SECOND_MATE_NUDGE_MESSAGE" ] || { + echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: retry marker message mismatch" + continue + } ;; + 1) [ "$message" = "$REMOTE_SECOND_MATE_NUDGE_MESSAGE" ] || { + echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: remote retry marker message mismatch" + continue + } ;; + *) + echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: retry marker placement is invalid" + continue + ;; + esac + [ "$remote" -ne 1 ] || continue meta="$STATE/$id.meta" [ -f "$meta" ] && [ "$(fm_meta_get "$meta" kind)" = secondmate ] || { echo "NUDGE_SECONDMATES: secondmate ${id:-unknown}: send failed: retry target has no live secondmate metadata" @@ -399,7 +425,7 @@ secondmate_sync() { fm_lock_release "$home_lock" || true continue } - if FM_CONFIG_INHERIT_REPORT="$report" \ + if FM_CONFIG_INHERIT_REPORT="$report" FM_CONFIG_INHERIT_LIVE=1 \ propagate_secondmate_inheritance "$FM_HOME" "$home_real" "$CONFIG" "$DATA"; then : else @@ -420,6 +446,66 @@ secondmate_sync() { rm -f "$report" fm_lock_release "$home_lock" || true done < <(live_secondmate_meta_records "$STATE" "$DATA/secondmates.md") + + # Remote routes converge through the generic transport. Their code root and + # inherited files are authoritative on that host; no local path probe or + # local fast-forward is attempted for them. + local remote_host sync_out inherit_out nudge_needed remote_marker remote_pending converged out remote_lock remote_generation + while IFS='|' read -r id _home _window meta; do + remote_host=$(fm_meta_get "$meta" remote_host) + [ -n "$remote_host" ] || continue + remote_lock=$(fm_remote_inherit_transaction_lock_path "$STATE" "$id" 2>/dev/null || true) + if [ -z "$remote_lock" ] || ! fm_lock_acquire_wait "$remote_lock"; then + echo "NUDGE_SECONDMATES: secondmate $id: send failed: cannot lock remote inheritance transaction" + continue + fi + if ! "$SCRIPT_DIR/fm-procevent-remote-reply.sh" arm "$id" >/dev/null 2>&1; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote reply source could not be registered" + fi + remote_generation=$(fm_remote_inherit_generation_next "$STATE" "$id" 2>/dev/null || true) + if [ -z "$remote_generation" ]; then + echo "SECONDMATE_SYNC: secondmate $id: skipped: remote inheritance generation could not be published" + fm_lock_release "$remote_lock" || true + continue + fi + remote_marker=$(secondmate_nudge_marker_path "$id" 2>/dev/null || true) + remote_pending=0 + if [ -f "$remote_marker" ] && [ "$(fm_meta_get "$remote_marker" remote)" = 1 ]; then remote_pending=1; fi + if ! secondmate_write_nudge_marker "$id" "$_home" "" remote \ + "$REMOTE_SECOND_MATE_NUDGE_MESSAGE" 1; then + echo "NUDGE_SECONDMATES: secondmate $id: send failed: cannot record remote retry marker" + fm_lock_release "$remote_lock" || true + continue + fi + nudge_needed=0 + converged=1 + if sync_out=$("$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-secondmate-control.sh sync "$id" < /dev/null 2>&1); then + case "$sync_out" in synced:*) nudge_needed=1 ;; esac + else + echo "SECONDMATE_SYNC: secondmate $id: skipped: remote tracked-file sync failed on $remote_host: $(first_line "$sync_out")" + converged=0 + fi + if inherit_out=$(FM_CONFIG_INHERIT_LIVE=1 \ + "$SCRIPT_DIR/fm-remote-inherit-push.sh" "$id" "$remote_generation" 2>&1); then + if printf '%s\n' "$inherit_out" | grep -Eq '^(pushed|removed):'; then nudge_needed=1; fi + else + echo "SECONDMATE_SYNC: secondmate $id: skipped: remote inheritance failed on $remote_host: $(first_line "$inherit_out")" + converged=0 + fi + [ "$remote_pending" -eq 0 ] || nudge_needed=1 + if [ "$converged" -eq 1 ] && [ "$nudge_needed" -eq 1 ]; then + if out=$(FM_HOME="$FM_HOME" FM_ROOT_OVERRIDE="$FM_ROOT" FM_STATE_OVERRIDE="$STATE" \ + "$SCRIPT_DIR/fm-send.sh" "fm-$id" "$REMOTE_SECOND_MATE_NUDGE_MESSAGE" 2>&1); then + rm -f "$remote_marker" + [ "${FM_BOOTSTRAP_VERBOSE_FACTS:-0}" != 1 ] || echo "BOOTSTRAP_INFO: nudged remote fm-$id after convergence" + else + echo "NUDGE_SECONDMATES: secondmate $id: send failed: $(first_line "$out")" + fi + elif [ "$converged" -eq 1 ]; then + rm -f "$remote_marker" + fi + fm_lock_release "$remote_lock" || true + done < <(live_secondmate_meta_records "$STATE" "$DATA/secondmates.md") return 0 } @@ -436,7 +522,7 @@ secondmate_liveness_sweep() { # primary-only no-op there. Mid-session liveness remains explicitly out of # scope and requires a separate periodic signal. [ -d "$STATE" ] || return 0 - local meta id window harness backend target agent_state out cause + local meta id window harness backend target agent_state out cause remote_host remote_rc readiness_reason route_out remote_backend SECONDMATE_RESPAWNED_IDS="" for meta in "$STATE"/*.meta; do [ -f "$meta" ] || continue @@ -445,6 +531,73 @@ secondmate_liveness_sweep() { window=$(fm_meta_get "$meta" window) [ -n "$window" ] || continue harness=$(fm_meta_get "$meta" harness) + remote_host=$(fm_meta_get "$meta" remote_host) + if [ -n "$remote_host" ]; then + remote_rc=0 + fm_remote_readiness_ensure "$SCRIPT_DIR" "$id" || remote_rc=$? + if [ "$remote_rc" -eq 255 ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote host unavailable or endpoint state unknown; route preserved on $remote_host" + continue + fi + if [ "$remote_rc" -ne 0 ]; then + readiness_reason=$(printf '%s\n' "$FM_REMOTE_READINESS_OUT" \ + | awk '/^check [^=]+=(fixable|human):|^action:|^error:/ { print; exit }') + [ -n "$readiness_reason" ] || readiness_reason=$(first_line "$FM_REMOTE_READINESS_OUT") + [ -n "$readiness_reason" ] || readiness_reason="unknown readiness failure" + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote readiness failed on $remote_host: $readiness_reason" + continue + fi + if out=$("$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-secondmate-control.sh state "$id" < /dev/null 2>/dev/null); then + remote_rc=0 + else + remote_rc=$? + fi + if [ "$remote_rc" -eq 255 ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote host unavailable or endpoint state unknown; route preserved on $remote_host" + continue + fi + if [ "$remote_rc" -ne 0 ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote endpoint probe unreadable on $remote_host" + continue + fi + agent_state=$(printf '%s\n' "$out" | tail -1) + case "$agent_state" in + alive) + if route_out=$("$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-secondmate-control.sh route "$id" < /dev/null 2>/dev/null); then + remote_rc=0 + else + remote_rc=$? + fi + if [ "$remote_rc" -eq 255 ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote host unavailable or endpoint route unknown; route preserved on $remote_host" + continue + fi + if [ "$remote_rc" -ne 0 ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: alive remote endpoint route is unreadable on $remote_host; inspect and migrate or retire it explicitly" + continue + fi + remote_backend=$(printf '%s\n' "$route_out" | sed -n 's/^backend=//p' | tail -1) + if [ "$remote_backend" != herdr ]; then + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: alive remote endpoint is recorded on backend '${remote_backend:-missing}'; migrate or retire it explicitly" + continue + fi + [ "${FM_BOOTSTRAP_VERBOSE_FACTS:-0}" != 1 ] || echo "BOOTSTRAP_INFO: remote secondmate $id already live (host=$remote_host)" + ;; + dead|missing) + cause="remote endpoint $agent_state on its configured host" + if out=$(FM_SPAWN_NO_GUARD=1 "$FM_ROOT/bin/fm-spawn.sh" "$id" --secondmate 2>&1); then + SECONDMATE_RESPAWNED_IDS="$SECONDMATE_RESPAWNED_IDS $id" + else + echo "SECONDMATE_LIVENESS: secondmate $id: respawn failed after $cause: $(first_line "$out")" + fi + ;; + ambiguous|unreadable|unverified) + echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote endpoint state is $agent_state on $remote_host" + ;; + *) echo "SECONDMATE_LIVENESS: secondmate $id: skipped: remote endpoint returned an invalid state" ;; + esac + continue + fi backend=$(fm_backend_of_meta "$meta") target=$(fm_backend_target_of_meta "$meta") [ -n "$target" ] || target="$window" @@ -494,6 +647,27 @@ secondmate_liveness_sweep() { return 0 } +secondmate_handoff_resume() { + [ -d "$DATA/handoff" ] || return 0 + "$SCRIPT_DIR/fm-backlog-handoff.sh" --resume-pending >/dev/null 2>&1 || true +} + +secondmate_handoff_detect() { + local outbox id count + [ -d "$DATA/handoff" ] || return 0 + for outbox in "$DATA/handoff"/*.outbox.md; do + [ -e "$outbox" ] || continue + id=$(basename "$outbox" .outbox.md) + case "$id" in ''|*[!A-Za-z0-9._-]*) id=unknown ;; esac + if [ ! -f "$outbox" ] || [ -L "$outbox" ]; then + echo "SECONDMATE_HANDOFF: secondmate $id: pending delivery: unsafe outbox" + continue + fi + count=$(awk '/^- \[[ x]\] / { count++ } END { print count + 0 }' "$outbox" 2>/dev/null || printf unknown) + echo "SECONDMATE_HANDOFF: secondmate $id: pending delivery: $count item(s)" + done +} + install_cmd() { case "$1" in tmux|node|git|gh|curl|jq|orca|zellij) echo "brew install $1 # or the platform's package manager" ;; @@ -536,6 +710,7 @@ if ! BACKEND_TOOLS=$(fm_backend_required_tools "$BACKEND"); then fi TOOLS="$BACKEND_TOOLS $COMMON_TOOLS" NO_MISTAKES_MIN=1.31.2 +GH_AXI_MIN=0.1.29 treehouse_supports_lease() { treehouse get --help 2>&1 | grep -Eq '(^|[^[:alnum:]_-])--lease([^[:alnum:]_-]|$)' @@ -721,6 +896,121 @@ EOF echo "FMX: X mode on - relay poll armed via state/x-watch.check.sh; 30s watcher cadence in config/x-mode.env" } +# Model registry checks, all detect-only and free: a schema validation, the +# referential-integrity check that binds config/crew-dispatch.json to +# config/models.json, and the local price-drift comparison. +# +# The integrity check is the entire reason the registry is a second file: it +# catches a bad model at CONFIG-EDIT time, before any worker is launched against +# it. A dispatch rule naming a model with no probe record, or one whose recorded +# status is rejected, fails here. +# +# The no-registry branch is what keeps the ruled "inert but never silent" posture +# honest: with no config/models.json the spawn-time refusal cannot run, so if the +# dispatch config routes to any provider-prefixed model this says so plainly +# rather than leaving the zero-budget rule quietly unenforced. +model_registry_validate() { + local reg dispatch err drift routed + reg="$CONFIG/models.json" + dispatch="$CONFIG/crew-dispatch.json" + + if [ ! -f "$reg" ]; then + [ -f "$dispatch" ] || return 0 + command -v jq >/dev/null 2>&1 || return 0 + routed=$(jq -r ' + def profiles($v): + if ($v | type) == "array" then $v elif ($v | type) == "object" then [$v] else [] end; + [ (((.rules // [])[]? | profiles(.use?)[]?), (profiles(.default // null)[]?)) + | .model? // empty ] + | map(select(type == "string" and (. != "default") and (test("/")))) + | unique | join(", ")' "$dispatch" 2>/dev/null || true) + if [ -n "$routed" ]; then + echo "MODEL_REGISTRY: no config/models.json, so the zero-budget rule is not enforced for routed provider models: $routed" + fi + return 0 + fi + + if ! command -v jq >/dev/null 2>&1; then + echo "MISSING: jq (install: $(install_cmd jq))" + return 0 + fi + if ! err=$(fm_model_registry_validate "$reg"); then + echo "MODEL_REGISTRY: invalid config/models.json - $err" + return 0 + fi + fm_model_registry_integrity "$dispatch" "$reg" || true + drift=$(fm_model_price_drift "$reg" || true) + [ -z "$drift" ] || printf '%s\n' "$drift" +} + +admission_control_validate() { + local file reason state + file=$(fm_admission_config_file "$CONFIG") + [ -f "$file" ] || return 0 + if ! command -v jq >/dev/null 2>&1; then + # The MISSING: jq line already carries the install-consent flow. + return 0 + fi + # crew_dispatch_validate owns the file-level JSON parse for this exact file; + # reporting it twice would give one broken file two diagnostics. + jq -e . "$file" >/dev/null 2>&1 || return 0 + reason=$(fm_admission_validate_reason "$file") || true + if [ -n "$reason" ]; then + echo "ADMISSION_CONTROL: invalid config/crew-dispatch.json _scheduling.admission_control - $reason" + return 0 + fi + if [ "${FM_BOOTSTRAP_VERBOSE_FACTS:-0}" = 1 ]; then + # A dispatch file with no admission block at all is the ordinary case and is + # not a fact worth reporting; only a policy that exists gets a verbose line. + state=$(fm_admission_state "$file") + case "$state" in + inert|active) echo "BOOTSTRAP_INFO: fleet admission control $state" ;; + esac + fi +} + +# Outcome records in the wake ledger that join no wake record. They read as +# ordinary supervision cost while measuring nothing, and the state they leave - +# queued=unknown - is also what a legitimately wiped state/ produces, so the +# only way the fleet ever notices is a count reported at session start. +# bin/fm-wake-ledger.sh owns the join and the number; this only reports it. +wake_ledger_reconcile() { + local unjoined + # A code root without the ledger script predates this check; that is a missing + # feature, not an unreadable ledger, so say nothing rather than misreport it. + [ -x "$SCRIPT_DIR/fm-wake-ledger.sh" ] || return 0 + # An absent ledger reports a real 0 and stays silent here. A failure means the + # ledger exists and could not be read, which must not read as "clean". + if ! unjoined=$("$SCRIPT_DIR/fm-wake-ledger.sh" reconcile --count 2>/dev/null); then + echo "WAKE_LEDGER: the wake ledger could not be read, so unjoined outcome records cannot be counted - treat supervision-cost figures as unverified until it is readable" + return 0 + fi + case "$unjoined" in + ''|*[!0-9]*) return 0 ;; + 0) return 0 ;; + esac + echo "WAKE_LEDGER: $unjoined outcome record(s) join no wake record - supervision-cost figures drawn from this ledger overcount until they are purged (bin/fm-wake-ledger.sh reconcile)" +} + +# The entitlement probe half of the observation floor. A MUTATING sweep: it makes +# live requests and writes state/model-health.json, so it runs only when this +# session actually holds the fleet lock, alongside the other mutating sweeps. +# +# Interval-gated by each model's observation level, so the steady-state cost is +# usually zero probes and one file read rather than ~4s per routed model on every +# session start. fm-model-verify.sh owns the probe mechanics, the hard timeout, +# and the closed stdin. +model_probe_sweep() { + local out + [ -f "$CONFIG/models.json" ] || return 0 + # stderr is captured alongside stdout: the script's failure diagnostics (a + # due-selection query that dies at runtime) arrive there, and a sweep whose + # failures are swallowed reads as a healthy sweep that probed nothing. + # Non-fatal either way - bootstrap detects and reports, never aborts. + out=$("$SCRIPT_DIR/fm-model-verify.sh" 2>&1 || true) + [ -z "$out" ] || printf '%s\n' "$out" +} + crew_dispatch_validate() { local file err file="$CONFIG/crew-dispatch.json" @@ -875,6 +1165,9 @@ fi if command -v no-mistakes >/dev/null 2>&1 && ! tool_version_at_least no-mistakes "$NO_MISTAKES_MIN"; then echo "MISSING: no-mistakes (install: $(install_cmd no-mistakes))" fi +if command -v gh-axi >/dev/null 2>&1 && ! tool_version_at_least gh-axi "$GH_AXI_MIN"; then + echo "MISSING: gh-axi (install: $(install_cmd gh-axi))" +fi if command -v quota-axi >/dev/null 2>&1 && ! fm_quota_axi_compatible; then echo "MISSING: quota-axi (install: $(install_cmd quota-axi))" fi @@ -900,6 +1193,9 @@ if [ "${FM_BOOTSTRAP_VERBOSE_FACTS:-0}" = 1 ] && [ -n "$crew" ] && [ "$crew" != echo "BOOTSTRAP_INFO: crew harness override active: $crew" fi crew_dispatch_validate +model_registry_validate +admission_control_validate +wake_ledger_reconcile if [ "${FM_BOOTSTRAP_VERBOSE_FACTS:-0}" = 1 ] \ && ! fm_backlog_backend_manual "$CONFIG" && fm_tasks_axi_compatible; then echo "BOOTSTRAP_INFO: tasks-axi available" @@ -907,7 +1203,10 @@ fi if [ "${FM_BOOTSTRAP_DETECT_ONLY:-0}" != 1 ]; then secondmate_liveness_sweep secondmate_sync + secondmate_handoff_resume x_mode_setup fleet_sync + model_probe_sweep fi +secondmate_handoff_detect exit 0 diff --git a/bin/fm-brief.sh b/bin/fm-brief.sh index fc289dd6e86..c6397c15914 100755 --- a/bin/fm-brief.sh +++ b/bin/fm-brief.sh @@ -6,7 +6,8 @@ # description, acceptance criteria, and context, and may adjust other sections # when the task genuinely deviates (e.g. working an existing external PR instead # of shipping a new one). -# Usage: fm-brief.sh [--scout] [--herdr-lab] +# Usage: fm-brief.sh --mode [--herdr-lab] [--slot-base [--contribution-target ]] +# fm-brief.sh --scout [--herdr-lab] [--slot-base ] # fm-brief.sh --secondmate {...|--no-projects} # --scout writes the scout contract instead: the deliverable is a report at # data//report.md (no branch, no push, no PR) and the worktree is scratch. @@ -26,24 +27,71 @@ # The flag must be explicit because {TASK} is filled after scaffolding and the # caller-supplied repo string cannot reliably identify this repo. Briefs made # without it carry a loud declaration so an omitted contract cannot be silent. -# For ship tasks, the definition of done is shaped by the project's delivery mode -# (data/projects.md via fm-project-mode.sh; see the project-management skill -# and AGENTS.md task lifecycle): -# no-mistakes implement -> /no-mistakes pipeline -> PR -> captain merge (default) -# direct-PR implement -> push + open PR via gh-axi (no pipeline) -> captain merge +# For ship tasks, --mode is REQUIRED and shapes the definition of done. Firstmate +# resolves it per task at intake (AGENTS.md section 7); data/projects.md holds the +# captain's standing posture as context, and this script never reads it: +# no-mistakes implement -> /no-mistakes pipeline -> PR -> configured merge authority +# direct-PR implement -> push + open PR via gh-axi (no pipeline) -> configured merge authority # local-only implement on branch, stop and report "ready in branch" (no push/PR); -# captain approves, firstmate merges to local main +# the configured merge authority approves, firstmate merges to local main +# no-mistakes-prod-only is a registry policy, not a task mode; resolve it to one of +# the three concrete modes at intake before calling this script. +# The generated ship brief records the chosen mode as a fixed machine-readable +# "Delivery contract: mode=" line. bin/fm-spawn.sh reads that line and refuses +# to launch a ship task whose explicit --mode disagrees, so an adjusted brief and the +# recorded task metadata cannot drift apart. +# --slot-base and --contribution-target are the task's two base references, resolved +# by bin/fm-spawn.sh and passed in here; bin/fm-task-base-lib.sh owns the contract. +# This script is TOLD them rather than deriving them, exactly like --mode: its +# argument is a caller-supplied string, not a checkout it could read. +# They are stated only when they DIFFER, so a project with one base keeps a +# byte-identical brief. When they differ the brief carries a machine-readable +# "Base contract: slot= contribution=" line that fm-spawn.sh +# checks against its own resolution, plus the read-here/write-there rule: read the +# slot base because it is the code the fleet runs and what the task's citations were +# taken against, and cut the branch from the contribution target so the PR carries no +# commit that target never had. A ship brief's branch step then names that commit +# explicitly instead of branching from wherever the worktree happens to sit. A scout +# cuts no branch and takes only --slot-base, which its report cites its findings +# against. Both are refused on a secondmate charter. # Ship briefs begin with a worktree-isolation assertion before the branch step. -# Scout tasks ignore mode - their deliverable is a report, not a merge. +# --mode is refused on scout and secondmate scaffolds: a scout's deliverable is a +# report rather than a merge, and a charter is not a delivery contract. +# There is no --yolo flag here. The worker never owns approval decisions, so yolo is +# a spawn-time and firstmate-side input only (AGENTS.md section 7). # Every scaffold's status protocol distinguishes the configured # declared-external-wait verb (FM_CLASSIFY_PAUSED_VERB, default "paused") from # "blocked:": pause for a known external wait expected to clear on its own, # blocked when firstmate must act. +# Every crewmate scaffold (ship and scout) also carries a "Who is speaking to +# you" section. Firstmate's own steers arrive marked (bin/fm-send.sh, carrier +# owned by bin/fm-operational-input.sh), so an unmarked message is a human at the +# keyboard who may believe the pane is firstmate; the worker must identify itself +# as a worker on this task before acting on one. The section names the single +# exclusion - a bare slash or codex `$` command, which a harness +# recognizes only at the start of the line and which therefore cannot be marked - +# so the rule is never quietly violated by the routine validation trigger. It +# also restates that escalation is the status file, because the opposite failure +# is silent: a gate parked with only "Captain, ..." in a pane nobody reads stalls +# the task with no wake and no visible symptom. +# Ship and scout tasks include the Claude context-pressure snapshot path and +# the host-computed 70%-used /compact trigger. Other harnesses see the section +# but do not fabricate a reading when the optional snapshot is absent. +# Secondmate charters carry the same 70%-used /compact trigger read from the +# bottom CTX status-line row, with no snapshot file. # Ship tasks include a project-memory section so durable project-intrinsic # learnings can be committed to AGENTS.md through the project's delivery path; # it carries the AGENTS.md authoring bar (widely useful knowledge only, pointers # over copied detail) and has the crewmate add the fm-ensure-agents-md.sh # self-governance section when a touched project AGENTS.md lacks it. +# Every scaffold carries a verification-discipline section: witness a negative +# control fail before trusting success reported only by absence, and wait on +# completion artifacts rather than process names. +# Ship scaffolds additionally carry branch conflict resolution: rebase and +# resolve branch/base conflicts autonomously at any file count when intent is +# clear, escalating only genuinely ambiguous intent to firstmate. In +# no-mistakes mode that delegated rebase is the stated exception to the +# do-not-commit-while-a-run-is-active rule. # Refuses to overwrite an existing brief. set -eu @@ -94,16 +142,89 @@ fi KIND=ship HERDR_LAB=0 NO_PROJECTS=0 +MODE= +MODE_SET=0 +SLOT_BASE= +CONTRIB_TARGET= POS=() +want_value= for a in "$@"; do + if [ -n "$want_value" ]; then + case "$a" in + --*) echo "error: --$want_value requires a value" >&2; exit 1 ;; + esac + case "$want_value" in + mode) MODE=$a; MODE_SET=1 ;; + slot-base) SLOT_BASE=$a ;; + contribution-target) CONTRIB_TARGET=$a ;; + *) echo "error: internal parser state for --$want_value" >&2; exit 1 ;; + esac + want_value= + continue + fi case "$a" in --scout) KIND=scout ;; --secondmate) KIND=secondmate ;; --herdr-lab) HERDR_LAB=1 ;; --no-projects) NO_PROJECTS=1 ;; + --mode) want_value=mode ;; + --mode=*) MODE=${a#--mode=}; MODE_SET=1 ;; + --slot-base) want_value=slot-base ;; + --slot-base=*) SLOT_BASE=${a#--slot-base=} ;; + --contribution-target) want_value=contribution-target ;; + --contribution-target=*) CONTRIB_TARGET=${a#--contribution-target=} ;; + # yolo never reaches the worker: it is firstmate's approval authority, not a + # brief input. Refuse it loudly so it is never silently dropped here and then + # believed to have been recorded. + --yolo|--yolo=*) echo "error: --yolo is not a brief input; pass it to bin/fm-spawn.sh, which records the task's approval posture" >&2; exit 1 ;; *) POS+=("$a") ;; esac done +[ -z "$want_value" ] || { echo "error: --$want_value requires a value" >&2; exit 1; } + +# Ship delivery mode is an explicit per-task decision (AGENTS.md section 7). A +# missing or invalid value stops the scaffold rather than silently defaulting. +if [ "$KIND" = ship ]; then + [ "$MODE_SET" -eq 1 ] || { + echo "error: ship briefs require --mode ; resolve it at intake from the captain's instruction and the project's registered posture in data/projects.md" >&2 + exit 1 + } + case "$MODE" in + no-mistakes|direct-PR|local-only) ;; + no-mistakes-prod-only) + echo "error: no-mistakes-prod-only is a registry policy, not a task mode; classify this task's surface and resolve it to no-mistakes or direct-PR at intake" >&2 + exit 1 ;; + *) echo "error: --mode must be one of no-mistakes, direct-PR, local-only (got '$MODE')" >&2; exit 1 ;; + esac +elif [ "$MODE_SET" -eq 1 ]; then + echo "error: --mode applies only to ship briefs; a scout delivers a report and a secondmate charter is not a delivery contract" >&2 + exit 1 +fi + +# A worker needs TWO base references whenever they differ (bin/fm-task-base-lib.sh +# owns the contract, bin/fm-spawn.sh resolves them). This script is handed the +# resolved pair rather than deriving it, exactly like --mode: its REPO argument is +# a caller-supplied name, not a checkout it could read. +if [ "$KIND" = secondmate ] && { [ -n "$SLOT_BASE" ] || [ -n "$CONTRIB_TARGET" ]; }; then + echo "error: --slot-base and --contribution-target apply only to crewmate ship or scout briefs; a secondmate charter cuts no contribution branch" >&2 + exit 1 +fi +if [ -n "$CONTRIB_TARGET" ] && [ -z "$SLOT_BASE" ]; then + echo "error: --contribution-target requires --slot-base; stating where to write without stating where to read is the confusion this contract exists to prevent" >&2 + exit 1 +fi +if [ "$KIND" = scout ] && [ -n "$CONTRIB_TARGET" ]; then + echo "error: --contribution-target applies only to ship briefs; a scout delivers a report and cuts no branch, so it has only a slot base to read and cite" >&2 + exit 1 +fi +case "$SLOT_BASE" in + ''|*[!0-9a-f]*) [ -z "$SLOT_BASE" ] || { echo "error: --slot-base must be a full commit SHA (got '$SLOT_BASE')" >&2; exit 1; } ;; +esac +case "$CONTRIB_TARGET" in + ''|unresolved) ;; + *[!0-9a-f]*) echo "error: --contribution-target must be a full commit SHA or the literal 'unresolved' (got '$CONTRIB_TARGET')" >&2; exit 1 ;; +esac + ID=${POS[0]} if [ "$KIND" = secondmate ] && [ "$HERDR_LAB" -eq 1 ]; then @@ -127,6 +248,36 @@ shell_quote() { } STATUS_FILE=$(shell_quote "$STATE/$ID.status") +IFS= read -r -d '' CONTEXT_SECTION <\` token after the marker; include that exact token in your parent status reply (or in the status pointer to a detailed doc) so the parent can correlate the answer. Optional helper: \`bin/fm-secondmate-report.sh\` can append a correlated status line for you, but a plain \`echo\` that includes the same \`corr=\` is equally valid - do not depend on the helper being present. @@ -247,6 +401,78 @@ EOF HERDR_SECTION=${HERDR_SECTION%$'\n'} fi +# Shared by the ship and scout scaffolds: a crewmate cannot otherwise tell +# firstmate's steers from a human typing into its pane, and both failure +# directions are real. Rule 4's status file is named explicitly because the +# opposite error - answering the captain in a pane nobody reads - is silent. +IFS= read -r -d '' WHO_IS_SPEAKING <` instead of assuming the file in front of you is what the fleet runs.' \ +'Never cut or rebase your branch onto `'"$SLOT_SHORT"'`: that silently carries every commit the fleet has landed but never contributed into your PR, which is exactly what this separation prevents.' \ +'Verify your branch before pushing: `git merge-base --is-ancestor '"$CONTRIB_SHORT"' HEAD` must succeed.' \ +'' \ +'If a file your task names does not exist at `'"$CONTRIB_SHORT"'` at all, then this task cannot be a contribution to that base.' \ +'Do NOT recreate the file and do not rebase onto the running trunk to make it appear: append `blocked: absent from the contribution base` and stop.') + BRANCH_FROM=" $CONTRIB_SHORT" + fi +fi +[ -z "$BASE_SECTION" ] || BASE_SECTION="$BASE_SECTION + +" + if [ "$KIND" = scout ]; then cat > "$BRIEF" <" line that bin/fm-spawn.sh checks against its own +# explicit --mode before launching. case "$MODE" in direct-PR) SETUP2="" RULE1='1. Never push to the default branch (push only your `fm/'"$ID"'` branch). Never merge a PR.' IFS= read -r -d '' DOD < [ "$verb" = "${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT}" ] } +# --- backend-pushed blocked-on-human stale wakes ---------------------------- +# +# A push-capable backend can report an agent-state edge to `blocked` - the +# harness stopped for a human prompt (a permission dialog, a trust prompt, an +# interactive menu). bin/fm-transition-lib.sh's policy table is the one owner of +# that classification, and the producer already applies EDGE-triggered dedupe on +# agent state: one wake per `->blocked` edge, cleared by the next `->working` +# edge. The wake therefore arrives pre-deduplicated, and the ordinary stale +# path's STATUS-LINE dedupe must not be layered on top of it. A task correctly +# parked on a captain decision has an unchanged terminal status line by +# definition, so keying on status text absorbs every block after the first and +# the worker waits on a prompt nobody sees. Keying a second time on agent state +# would duplicate the producer's dedupe in a second owner instead. +# +# The two functions below are the ONE owner of that wake's detail grammar: the +# producer builds the detail with stale_detail_blocked_on_human and every +# consumer recognizes it with stale_detail_is_blocked_on_human, so neither side +# spells the text itself and a reworded producer cannot silently stop being +# exempt. The invariant suffix carries the meaning; the backend and agent-state +# prefix is telemetry for the reader. +FM_CLASSIFY_BLOCKED_ON_HUMAN_DETAIL='waiting on human, escalated immediately, not via wedge timer' + +# Build the stale-wake detail for a backend-pushed blocked-on-human edge. +stale_detail_blocked_on_human() { # + printf '%s: agent %s - %s' "$1" "$2" "$FM_CLASSIFY_BLOCKED_ON_HUMAN_DETAIL" +} + +# 0 if a stale wake's parenthesized detail is a backend-pushed blocked-on-human +# edge. Suffix-anchored on the invariant above so an ordinary wedge detail, a +# window name, or an empty detail never matches. +stale_detail_is_blocked_on_human() { # + local detail=$1 + [ -n "$detail" ] || return 1 + case "$detail" in + *"$FM_CLASSIFY_BLOCKED_ON_HUMAN_DETAIL") return 0 ;; + esac + return 1 +} + +# --- pause kind: can this wait change without the captain? ------------------- +# +# status_is_paused answers "does this pane idle by design". It does NOT answer the +# question a re-surface cadence actually asks: can the thing being waited on change +# without the captain acting? An external wait can, so rechecking it is real work. +# A captain-gated wait cannot - it clears only when the captain acts, and the +# captain acting is already the away-mode exit signal, which runs the full return +# catch-up, so a timed recheck can never surface anything that exit does not. +# +# The backlog already records exactly this distinction per work item as hold_kind +# (captain|external|load|parked|future), so these read that existing vocabulary +# instead of parsing pause prose or inventing a parallel one. +# +# NOT a pure read: task_hold_kind shells out to the backlog reader in FM_HOME, the +# same tool and field bin/fm-decision-hold.sh verifies a captain hold with. Callers +# run it only once a pause has already aged past its window, never on every wake. +FM_CLASSIFY_CAPTAIN_HOLD_KIND_DEFAULT='captain' + +_fm_show_field() { # -> value, or empty + printf '%s\n' "$1" | sed -n "s/^[[:space:]]*$2: //p" | head -1 +} + +# Print the backlog hold kind recorded for , or `unknown` when it cannot +# be established: no reader, an unreadable or absent item, or an item that carries +# no active hold at all. `unknown` is deliberately never a kind a caller may treat +# as captain-gated, so an indeterminate wait keeps its ordinary handling. +task_hold_kind() { # [home] + local id=$1 home=${2:-${FM_HOME:-}} out kind + case "$id" in ''|*[!A-Za-z0-9._-]*) printf 'unknown'; return ;; esac + command -v tasks-axi >/dev/null 2>&1 || { printf 'unknown'; return; } + if [ -n "$home" ] && [ -d "$home" ]; then + out=$(cd "$home" && tasks-axi show "$id" --full 2>/dev/null) || out='' + else + out=$(tasks-axi show "$id" --full 2>/dev/null) || out='' + fi + [ -n "$out" ] || { printf 'unknown'; return; } + [ "$(_fm_show_field "$out" held)" = yes ] || { printf 'unknown'; return; } + kind=$(_fm_show_field "$out" hold_kind) + case "$kind" in + ''|*[!A-Za-z0-9._-]*) printf 'unknown' ;; + *) printf '%s' "$kind" ;; + esac +} + +# 0 if 's declared wait is gated on the captain rather than on something +# that can clear by itself. Consumers use it to decide CADENCE only: a captain-gated +# wait stays exactly as visible as before in the backlog digest, the fleet view, and +# the away-mode return catch-up - it simply stops being re-asked on a timer. +pause_is_captain_gated() { # [home] + [ "$(task_hold_kind "$@")" = "${FM_CLASSIFY_CAPTAIN_HOLD_KIND:-$FM_CLASSIFY_CAPTAIN_HOLD_KIND_DEFAULT}" ] +} + # --- durable keyed decisions ------------------------------------------------ # # The status stream is an append-only EVENT log. Reading it last-event-wins @@ -207,9 +302,15 @@ EOF # the file, no globals beyond the optional FM_CLASSIFY_RESOLVE_VERB override. This # is the durable open-set the fleet snapshot and any point-in-time consumer must use # instead of trusting the last status line. +# The scan_open_decisions wrapper below enumerates a whole directory rather than +# a single caller-chosen path, so a status file that is itself a symlink (e.g. +# escaping the state directory) is rejected outright with a plain [ -L ] check +# before any read - a cheap builtin, unlike fm_wake_latest_event's O_NOFOLLOW +# subprocess read, which exists for that function's much narrower payload-driven +# path resolution rather than this directory-local glob. status_open_decisions() { # local f=$1 line verb key note resolve held open='' stripped - [ -f "$f" ] || return 0 + [ -f "$f" ] && [ -r "$f" ] && [ ! -L "$f" ] || return 0 resolve=${FM_CLASSIFY_RESOLVE_VERB:-$FM_CLASSIFY_RESOLVE_VERB_DEFAULT} held=${FM_CLASSIFY_CAPTAIN_HELD_VERB:-$FM_CLASSIFY_CAPTAIN_HELD_VERB_DEFAULT} while IFS= read -r line || [ -n "$line" ]; do @@ -233,6 +334,30 @@ status_open_decisions() { # printf '%s' "$open" } +# Fleet-wide wrapper around status_open_decisions: scans every task's status +# log under and prefixes each still-open decision with its owning task +# id, so a per-wake or per-session surface can print the consolidated open set +# without re-walking the fold itself. A thin directory scan only - the fold +# above remains the ONE place the open/resolved semantics are decided. Prints +# one "\t\t\t" line per open decision, in glob (task id) +# order; prints nothing when none are open. +scan_open_decisions() { # + local state=$1 f task open line + for f in "$state"/*.status; do + [ -e "$f" ] || continue + task=$(basename "$f"); task="${task%.status}" + open=$(status_open_decisions "$f") || continue + [ -n "$open" ] || continue + while IFS= read -r line; do + [ -n "$line" ] || continue + printf '%s\t%s\n' "$task" "$line" + done < ... return 1 } -# Classify WHY an idle/stale crew MIGHT be safely absorbed instead of surfaced, -# from bin/fm-crew-state.sh's one authoritative current-state line -# ("state: · source: · "). Prints exactly one token: -# working - an actively-running no-mistakes step (running/fixing/ci) or a busy -# pane; the crew is legitimately mid-work on a static-looking pane -# (e.g. waiting on CI); +# 0 if is a SETTLED terminal state for crew : the crew's +# own work is over and the next move belongs above it, so an idle pane is the +# CORRECT condition rather than a wedge symptom. The state word comes from +# bin/fm-crew-state.sh, which reconciles the run-step and pane against the status +# log - so, unlike a read of the log's last line, a leftover terminal line under +# an ACTIVE run never reaches here (that crew reconciles as working and keeps its +# wedge timer). +# +# done - the run passed or its checks are green, or the log reports +# done and the pane is exactly idle. Nothing is left for the +# crew to do on its own. +# parked, blocked - the next move belongs above the crew, but ONLY while a +# durable open decision proves it. status_open_decisions is +# the one owner of that fold. A crew idling at a pipeline gate +# it is supposed to answer ITSELF opens no decision, so it +# keeps aging and still escalates as a possible wedge. +# +# `failed` is deliberately NOT settled: it also reconciles a CANCELLED run, the +# mid-supersession state in which a crew is expected to recover custody and +# resume, so an idle pane there is a genuine stall. `unknown` is never settled +# either - a dead endpoint or torn-down worktree must keep aging. +# +# A missing or unreadable status file yields no open decision, so parked/blocked +# stay unsettled and keep escalating. +crew_state_is_settled() { # [state-dir] + local id=$1 s=$2 dir=${3:-${STATE:-${FM_STATE_OVERRIDE:-}}} + case "$s" in + done) return 0 ;; + parked|blocked) + [ -n "$id" ] || return 1 + [ -n "$(status_open_decisions "$dir/$id.status")" ] + ;; + *) return 1 ;; + esac +} + +# --- process liveness: descendant CPU advancement --------------------------- +# +# Every absorb source below is SEMANTIC: the no-mistakes run step, the status +# log, the harness busy signal. All three correctly read "not working" the +# moment an agent backgrounds a long command and ends its turn - while the real +# work continues in a child process none of them can see. Measured 2026-08-03: +# one crew running the portable suite in the background produced seven +# consecutive false wedge escalations across 42 minutes, each demanding a deep +# inspection, and at least three other crews hit the same pattern while pipeline +# stages ran underneath them. +# +# This adds the missing PROCESS-level evidence ALONGSIDE those sources, never in +# place of them. Two rules keep it from becoming a blindfold: +# +# 1. Identity, never a bare pid. The kernel reissues pids, so a recorded pid +# routinely resolves to an unrelated live process; `kill -0`, a `/proc/` +# directory test and a bare `ps -p` all report that impostor as alive +# (data/learnings.md records the measured evidence). Every stored sample is +# bound to the fm_pid_identity of the agent it was taken from, and a sample +# whose anchor identity no longer matches is discarded, never compared. +# 2. ADVANCEMENT, never existence. A descendant that merely EXISTS is no +# evidence of work: a hung child would then mask a genuine wedge, trading a +# false alarm for the far more dangerous silence. Only cumulative CPU that +# GREW since the previous sample counts. +# +# The agent is resolved from kernel facts, never a vendor process name: the +# task's recorded worktree is the agent's working directory, and the agent is +# the LEADER of the foreground process group on that pane's terminal. +# +# Leader, not the whole group, and that choice is load-bearing. Whether a tool +# subprocess ends up in its parent's process group or its own is entirely a +# harness implementation detail - Claude Code detaches its Bash-tool children +# into a new session, a plain shell leaves a background job in the shell's own +# group and on the shell's terminal. Excluding the whole group would therefore +# make this signal silently measure nothing for any harness that spawns children +# the second way, which is far worse than the cost of the leader rule: where a +# harness does NOT get its own foreground group (a multi-process launcher such +# as pi-signed, or an agent started without `exec`), the second harness process +# is a descendant of the leader and its own CPU is counted. The threshold below +# and the caller's completed-turn bound are what keep that over-inclusion safe. +# +# The LEADER's own utime/stime is deliberately EXCLUDED. That is the agent +# itself working, which the semantic busy contract (bin/fm-busy-lib.sh) already +# owns, and counting it would let an idle agent's rendering jitter - or worse, a +# genuinely looping wedged agent - read as work. What is counted is everything +# below it: every live strict descendant's CPU, plus every already-reaped +# descendant's CPU through the leader's own cutime/cstime. That sum is monotonic +# across a child exiting (the child's total leaves the live set and lands in its +# parent's cutime), so one aggregate is enough and no per-child bookkeeping is +# needed - which is also what makes a long run of short-lived children, such as +# a test suite driving one script after another, register as advancing at all. +# +# A Linux-compatible /proc is required. Where it is absent this reports `none`, +# which is exactly today's no-evidence behaviour: the wake surfaces. + +# Kernel ticks of descendant CPU that must accrue between two samples before a +# crew counts as working. A tick is 10ms at the standard Linux USER_HZ of 100, +# so the default is one full CPU-second per sample: far above what a harness +# helper being spawned and reaped costs incidentally, and far below any real +# build, test run, or pipeline stage. At the watcher's default 15s poll that is +# roughly 7% of one core sustained. +FM_CHILD_CPU_MIN_TICKS_DEFAULT=100 +# A baseline older than this proves nothing about NOW, so it is refreshed and +# the probe reports no evidence for that poll rather than comparing against it. +FM_CHILD_CPU_MAX_SAMPLE_AGE_DEFAULT=120 +# The baseline is replaced only once it is at least this old. Both the no-verb +# signal path and the stale path can probe within one watcher cycle; without +# this the second probe would reset the baseline to a near-zero measurement +# window and report live work as static. +FM_CHILD_CPU_SAMPLE_INTERVAL_DEFAULT=5 + +# fm_pid_identity lives in bin/fm-wake-lib.sh, which creates the state directory +# when sourced. This library is also sourced by strictly read-only readers +# (bin/fm-crew-state.sh), so it is loaded on first probe instead of at source +# time. Every consumer that actually probes - the watcher and the away-mode +# daemon - has already loaded it, so this normally sources nothing. +_fm_child_cpu_need_identity() { + command -v fm_pid_identity >/dev/null 2>&1 && return 0 + # shellcheck source=bin/fm-wake-lib.sh + . "$_FM_CLASSIFY_LIB_DIR/fm-wake-lib.sh" 2>/dev/null || return 1 + command -v fm_pid_identity >/dev/null 2>&1 +} + +# Print "\t" for the agent that owns : the +# total CPU its descendants have consumed, live and already reaped, and the +# identity that total is bound to. Returns 1 when no agent can be resolved - +# no readable /proc, no live process whose working directory is that worktree, +# no controlling terminal, no foreground group, no live foreground group leader, +# or two different panes claiming the same worktree (refused rather than +# guessed). +_fm_child_cpu_measure() { # + local wt=$1 proc d pid line rest cur c i + local tty='' tpgid='' total=0 identity + local -a ppid_of pgrp_of own_of reaped_of kids_of visited queue + [ -n "$wt" ] && [ -d "$wt" ] || return 1 + proc=${FM_PROC_ROOT_OVERRIDE:-/proc} + [ -d "$proc" ] || return 1 + for d in "$proc"/[0-9]*; do + pid=${d#"$proc"/} + # stderr is redirected BEFORE stdin on purpose: redirections are applied + # left to right, so this also swallows the shell's own "no such file" + # complaint for a process that exited mid-scan, or for the literal glob left + # behind when the process table is empty. + read -r line 2>/dev/null < "$d/stat" || continue + # The comm field is parenthesised and may itself contain ") ", so strip + # greedily to the LAST one; every field after it is numeric. What remains + # starts at proc stat field 3, so field N sits at position N-2: ppid 4->2, + # pgrp 5->3, tty_nr 7->5, tpgid 8->6, utime 14->12, stime 15->13, + # cutime 16->14, cstime 17->15. + rest=${line##*') '} + # shellcheck disable=SC2086 # deliberate split of the fixed-shape numeric stat tail + set -- $rest + [ "$#" -ge 15 ] || continue + ppid_of[pid]=$2 + pgrp_of[pid]=$3 + own_of[pid]=$(( ${12} + ${13} )) + reaped_of[pid]=$(( ${14} + ${15} )) + # `-ef` compares the RESOLVED directory by device and inode, so it needs no + # readlink fork and is immune to a symlinked or bind-mounted worktree path. + [ "$d/cwd" -ef "$wt" ] 2>/dev/null || continue + # tty_nr 0 means no controlling terminal and tpgid -1 means no foreground + # group; neither can anchor an agent. + [ "$5" != 0 ] && [ "$6" != -1 ] || continue + if [ -z "$tty" ]; then + tty=$5 + tpgid=$6 + elif [ "$tty" != "$5" ] || [ "$tpgid" != "$6" ]; then + return 1 + fi + done + [ -n "$tpgid" ] || return 1 + # The foreground group's LEADER is the agent, and it must still be alive: a + # group whose leader has gone describes an agent already replaced. + [ -n "${pgrp_of[tpgid]+set}" ] && [ "${pgrp_of[tpgid]}" = "$tpgid" ] || return 1 + for pid in "${!ppid_of[@]}"; do + kids_of[ppid_of[pid]]="${kids_of[ppid_of[pid]]:-} $pid" + done + # The leader's own cutime/cstime: the CPU of every descendant it has already + # reaped. Its own utime/stime is excluded on purpose (see above). + total=${reaped_of[tpgid]} + visited[tpgid]=1 + queue=("$tpgid") + i=0 + while [ "$i" -lt "${#queue[@]}" ]; do + cur=${queue[i]} + i=$(( i + 1 )) + for c in ${kids_of[cur]:-}; do + [ -z "${visited[c]+set}" ] || continue + visited[c]=1 + total=$(( total + own_of[c] + reaped_of[c] )) + queue+=("$c") + done + done + identity=$(fm_pid_identity "$tpgid") || return 1 + printf '%s\t%s' "$total" "$identity" +} + +# Print the process-liveness verdict for crew : +# advancing - its agent's descendants burned at least FM_CHILD_CPU_MIN_TICKS +# of CPU since the previous sample, so work is happening NOW in a +# child process even though the pane and every semantic source +# look idle; +# static - an agent was resolved but its descendants did not advance (they +# are hung, finished, or absent), or there is no usable baseline +# to compare against yet; +# none - no agent could be resolved at all: no /proc, no worktree +# recorded, no live agent on it, or an unreadable process table. +# NOT a pure read: it maintains the state/.childcpu sample the next call +# measures against. Safe to call more than once per cycle - the baseline is +# replaced only once it is FM_CHILD_CPU_SAMPLE_INTERVAL old, so a second probe +# re-scores the SAME baseline instead of resetting the measurement window. +fm_child_cpu_state() { # + local state=$1 id=$2 meta wt file now measured identity ticks verdict + local prev_ts='' prev_ticks='' prev_identity='' age min interval max_age + [ -n "$state" ] && [ -n "$id" ] || { printf 'none'; return; } + meta="$state/$id.meta" + [ -f "$meta" ] || { printf 'none'; return; } + wt=$(grep '^worktree=' "$meta" 2>/dev/null | tail -1 | cut -d= -f2- || true) + [ -n "$wt" ] || { printf 'none'; return; } + _fm_child_cpu_need_identity || { printf 'none'; return; } + measured=$(_fm_child_cpu_measure "$wt") || { printf 'none'; return; } + ticks=${measured%%$'\t'*} + identity=${measured#*$'\t'} + file="$state/$id.childcpu" + now=$(date +%s) + min=${FM_CHILD_CPU_MIN_TICKS:-$FM_CHILD_CPU_MIN_TICKS_DEFAULT} + interval=${FM_CHILD_CPU_SAMPLE_INTERVAL:-$FM_CHILD_CPU_SAMPLE_INTERVAL_DEFAULT} + max_age=${FM_CHILD_CPU_MAX_SAMPLE_AGE:-$FM_CHILD_CPU_MAX_SAMPLE_AGE_DEFAULT} + # Identity last, so a stray separator inside it can never shift the numbers. + if [ -f "$file" ]; then + IFS=$'\t' read -r prev_ts prev_ticks prev_identity < "$file" 2>/dev/null || true + fi + # A missing, truncated, or corrupt baseline is no baseline: measure afresh. + case "$prev_ts" in ''|*[!0-9]*) prev_ts='' ;; esac + case "$prev_ticks" in ''|*[!0-9]*) prev_ts='' ;; esac + verdict=static + if [ -n "$prev_ts" ] && [ "$prev_identity" = "$identity" ]; then + age=$(( now - prev_ts )) + if [ "$age" -ge 0 ] && [ "$age" -le "$max_age" ] \ + && [ $(( ticks - prev_ticks )) -ge "$min" ]; then + verdict=advancing + fi + fi + if [ -z "$prev_ts" ] || [ "$prev_identity" != "$identity" ] \ + || [ $(( now - prev_ts )) -ge "$interval" ]; then + printf '%s\t%s\t%s\n' "$now" "$ticks" "$identity" > "$file" 2>/dev/null || true + fi + printf '%s' "$verdict" +} + +# 0 if crew shows advancing descendant CPU. The positive half of the +# process-liveness signal; see fm_child_cpu_state for the exact verdict. +crew_child_cpu_advancing() { # [state-dir] + local id=$1 state=${2:-${STATE:-${FM_STATE_OVERRIDE:-}}} + [ "$(fm_child_cpu_state "$state" "$id")" = advancing ] +} + +# The ONE read of bin/fm-crew-state.sh's authoritative current-state line +# ("state: · source: · ") that both classifications below +# share. Prints " ", because the two callers need +# different halves of the same read and that read may make a bounded +# no-mistakes call - splitting it into two reads would double that cost for +# every definite verdict. FM_CREW_STATE_BIN lets tests stub it. +_fm_crew_read_class() { # + local id=$1 line state src + [ -n "$id" ] || { printf 'definite unreadable'; return; } + line=$("$FM_CREW_STATE_BIN" "$id" 2>/dev/null) || true + state=unreadable + case "$line" in + state:*) + state=${line#state: }; state=${state%% *} + if [ "$state" = paused ]; then printf 'paused %s' "$state"; return; fi + if [ "$state" = working ]; then + src=${line#*source: }; src=${src%% *} + case "$src" in run-step|pane) printf 'working %s' "$state"; return ;; esac + fi + ;; + esac + case "$state" in + working|unknown|unreadable) printf 'inconclusive %s' "$state" ;; + *) printf 'definite %s' "$state" ;; + esac +} + +# Classify bin/fm-crew-state.sh's authoritative current-state line without +# consulting process liveness. Prints working, paused, definite, or inconclusive. +# FM_CREW_STATE_BIN lets tests stub the semantic verdict. +crew_semantic_class() { # + local read + read=$(_fm_crew_read_class "$1") + printf '%s' "${read%% *}" +} + +# Classify WHY an idle/stale crew MIGHT be safely absorbed instead of surfaced. +# Prints exactly one token: +# working - an actively-running no-mistakes step (running/fixing/ci), a busy +# pane, or advancing descendant CPU; the crew is legitimately +# mid-work on a static-looking pane (e.g. waiting on CI, or +# supervising a command it backgrounded before ending its turn); # paused - the crew's authoritative current state is a declared external-wait # pause (paused:), which is EXPECTED to idle; -# none - neither, so the wake must surface (a stopped/finished/parked/failed/ -# torn-down/unknown crew, or an unreadable verdict). -# One fm-crew-state.sh read serves BOTH absorb reasons at once. Reading the state -# authoritatively (not the status log) is what keeps run-step precedence: a crew -# that appended paused: but then STARTED a run reports working, never paused. -# NOT a pure read: fm-crew-state.sh may make a bounded no-mistakes call, so callers -# run it only on no-verb signal and first-sighting stale paths, never every wake. -# FM_CREW_STATE_BIN lets tests stub the verdict. -crew_absorb_class() { # - local id=$1 line state src +# settled - the crew's reconciled state is terminal and the idle pane is the +# expected finished/waiting condition (crew_state_is_settled above); +# none - none of those, so the wake must surface (a failed or cancelled run, +# a torn-down or unknown crew whose descendants are hung, dead, or +# absent, a run parked at a gate the crew owns, or an unreadable +# verdict). +# The two extra sources are consulted on exactly the semantic verdicts they +# answer for and never both: process liveness only after an INCONCLUSIVE read, +# the settled test only after a DEFINITE one. So the semantic sources keep their +# precedence in every direction - a crew that appended paused: but then STARTED a +# run reports working, never paused, a crew whose log still shows a +# pre-validation done: reports working, not settled, and a definite verdict is +# never overridden by whatever its process tree happens to still be doing. +# NOT a pure read: the shared current-state read may make a bounded no-mistakes +# call and the probe maintains its own sample, so callers run it only on no-verb +# signal and first-sighting stale paths, never every wake. +crew_absorb_class() { # [state-dir] + local id=$1 state_dir=${2:-${STATE:-${FM_STATE_OVERRIDE:-}}} read semantic state [ -n "$id" ] || { printf 'none'; return; } - line=$("$FM_CREW_STATE_BIN" "$id" 2>/dev/null) || true - case "$line" in state:*) ;; *) printf 'none'; return ;; esac - state=${line#state: }; state=${state%% *} - if [ "$state" = paused ]; then printf 'paused'; return; fi - if [ "$state" = working ]; then - src=${line#*source: }; src=${src%% *} - case "$src" in run-step|pane) printf 'working'; return ;; esac + read=$(_fm_crew_read_class "$id") + semantic=${read%% *} + state=${read#* } + case "$semantic" in + working|paused) printf '%s' "$semantic"; return ;; + definite) + crew_state_is_settled "$id" "$state" ${state_dir:+"$state_dir"} \ + && { printf 'settled'; return; } + printf 'none' + return + ;; + esac + if [ "$(fm_child_cpu_state "$state_dir" "$id")" = advancing ]; then + printf 'working' + return fi printf 'none' } diff --git a/bin/fm-composer-lib.sh b/bin/fm-composer-lib.sh index 6e2509ec2c1..723220e941e 100644 --- a/bin/fm-composer-lib.sh +++ b/bin/fm-composer-lib.sh @@ -45,9 +45,11 @@ # fm_composer_strip_ghost for the real-typed-content extraction, strips the box # borders, trims, and hands the result plus a flag to # fm_composer_classify_content for the shared -# empty|pending|unknown verdict. orca/cmux read a plain (unstyled) screen so -# they have no ghost styling to strip and rely on the idle-placeholder match -# below. Re-sourcing is a cheap idempotent redefinition, so this file needs no +# empty|pending|unknown verdict. An adapter's own ASCII trim is deliberately not +# enough: the shared verdict re-normalizes non-ASCII blank padding first, which +# is what fm_composer_normalize_blanks exists for. orca/cmux read a plain +# (unstyled) screen so they have no ghost styling to strip and rely on the +# idle-placeholder match below. Re-sourcing is a cheap idempotent redefinition, so this file needs no # include guard (matching bin/fm-tmux-lib.sh). # fm_composer_strip_ansi: drop every CSI escape sequence, leaving plain text. @@ -160,6 +162,40 @@ fm_composer_strip_ghost() { ' } +# fm_composer_normalize_blanks: map the non-ASCII blanks a TUI can emit as +# composer PADDING onto an ASCII space, and drop the zero-width ones, so the +# `[[:space:]]` trims below - and the ones every adapter already ran before +# calling - can see an otherwise-blank row as blank. Takes the text as an +# argument and prints the normalized text. +# +# WHY THIS EXISTS (task composer-nbsp-fix): real Claude Code 2.1.220 pads its +# EMPTY composer row with U+00A0 NBSP, so the captured row is exactly +# `❯` + `\xc2\xa0`. bash's `[[:space:]]` does not match U+00A0, so no trim could +# remove it, the leading-glyph strip below left a lone NBSP behind, and this +# owner concluded "real, unsubmitted content remains" -> `pending` on a +# genuinely idle pane, stably, on every poll. Away-mode escalation injection +# (bin/fm-supervise-daemon.sh) and verified submit both proceed only on an +# affirmative `empty`, so that one byte pair wedged escalation delivery for +# ~9.5 hours at a stretch. The defect was reader-independent: both the herdr +# ANSI reader and tmux `capture-pane -e` surface the NBSP faithfully because it +# originates in claude's own output. +# +# WHY IT IS SAFE: every character mapped here RENDERS AS BLANK, and nothing else +# is touched, so this can only ever make an OTHERWISE-BLANK row read as blank. A +# row holding any visible glyph keeps that glyph byte for byte and can never +# become `empty`; NBSP-joined real text stays `pending`. Normalizing in this one +# fleet-wide owner covers both ANSI readers and all four adapters at once, which +# is why no adapter carries its own copy. +fm_composer_normalize_blanks() { # -> normalized text on stdout + local s=$1 + s=${s//$'\xc2\xa0'/ } # U+00A0 no-break space (claude's composer padding) + s=${s//$'\xe2\x80\x87'/ } # U+2007 figure space + s=${s//$'\xe2\x80\xaf'/ } # U+202F narrow no-break space + s=${s//$'\xe2\x80\x8b'/} # U+200B zero-width space + s=${s//$'\xef\xbb\xbf'/} # U+FEFF zero-width no-break space / BOM + printf '%s' "$s" +} + # fm_composer_classify_content: the single shared composer-content verdict. # 1 when came from a genuine agent-composer container (a # bordered composer box, or a structurally-identified bare AGENT @@ -183,6 +219,16 @@ fm_composer_idle_matches() { fm_composer_classify_content() { # [idle_re] [idle_case] [plain_content] local bordered=$1 content=$2 idle_re=${3:-} idle_case=${4:-sensitive} plain_content plain_content=${5:-$content} + # Non-ASCII blanks first: a harness pads an EMPTY composer row with them and + # neither the caller's trims nor the ones here can see them (see + # fm_composer_normalize_blanks). Re-trim afterwards so the rest of this + # function keeps the trimmed inputs it documents. + content=$(fm_composer_normalize_blanks "$content") + plain_content=$(fm_composer_normalize_blanks "$plain_content") + content="${content#"${content%%[![:space:]]*}"}" + content="${content%"${content##*[![:space:]]}"}" + plain_content="${plain_content#"${plain_content%%[![:space:]]*}"}" + plain_content="${plain_content%"${plain_content##*[![:space:]]}"}" if [ "$bordered" != 1 ] && [ -z "$content" ] && [ -n "$plain_content" ]; then case "$plain_content" in '❯'|'›') printf 'empty'; return 0 ;; diff --git a/bin/fm-config-inherit-lib.sh b/bin/fm-config-inherit-lib.sh index bffbd5234d7..9957ec66e11 100644 --- a/bin/fm-config-inherit-lib.sh +++ b/bin/fm-config-inherit-lib.sh @@ -3,15 +3,24 @@ # set of LOCAL (gitignored) config items down into each secondmate home's # config/, so a secondmate's OWN crewmates inherit the primary's settings # (e.g. primary config/crew-dispatch.json makes a secondmate use the same dispatch -# profile rules, primary config/crew-harness=codex makes a secondmate's crewmates +# profile rules, primary config/models.json subjects that home's own crewmates to +# the same zero-budget allowlist, primary config/crew-harness=codex makes a secondmate's crewmates # spawn on codex too, primary config/backlog-backend=manual makes that home # hand-edit backlog files too, primary config/backend pins that home's local # runtime-backend default for future spawns, primary config/startup-memory-budget # bounds that home's startup-memory curation, and primary -# config/herdr-presentation-spaces enables the same default-off Herdr presentation -# projection). It also pushes the one primary-authoritative shared -# captain-preference file, data/captain-shared.md, into each secondmate home's -# data/ as a read-only copy. +# config/herdr-presentation-spaces carries the same Herdr presentation-projection +# choice - that item is default-ON, so an absent primary file and an absent +# destination file both mean on and the generic absence mirror below already +# converges a secondmate to the primary's default rather than turning it off; +# only an explicit primary "off" propagates an opt-out, and primary +# config/trace-context is copied at the launch convergence point as part of the +# default-off W3C trace-context setup, while live convergence leaves it unchanged. +# The primary passes its frozen home-session decision into a newly launched +# Secondmate; see docs/trace-context.md. +# It also pushes +# the one primary-authoritative shared captain-preference file, +# data/captain-shared.md, into each secondmate home's data/ as a read-only copy. # # Usage: . bin/fm-config-inherit-lib.sh (no FM_* setup required) # @@ -34,6 +43,15 @@ # secondmates, and a secondmate never spawns secondmates, so it must not flow # downstream. # +# That single declaration is also the ONE owner of the inherited-material +# allowlist for remote routes: bin/fm-remote-inherit-push.sh (sender) and +# bin/fm-remote-inherit.sh (receiver, executing inside the remote home) both +# derive their item set from fm_config_inherit_items rather than restating it, +# so a new inheritable item cannot be accepted by one side and refused by the +# other. A local and remote code root that disagree about this list must be +# reconciled by the ordinary remote sync/update path before the transfer +# succeeds; there is no separate allowlist version negotiation. +# # shellcheck source=bin/fm-startup-memory-budget-lib.sh . "$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/fm-startup-memory-budget-lib.sh" @@ -46,7 +64,40 @@ FM_SHARED_CAPTAIN_MODE="444" # The declared inheritable set (space-separated, config-dir-relative item paths). # Extend here to inherit more of the primary's local config; override via the # environment only in tests. Items must not contain whitespace. -FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget}" +# models.json travels WITH crew-dispatch.json and must not be separated from it: +# the registry is what the dispatch config's model names are validated against, so +# inheriting the rules without the registry would both leave a secondmate's own +# crewmates outside zero-budget enforcement and make every inherited model read as +# unregistered in that home. +FM_INHERITABLE_CONFIG="${FM_INHERITABLE_CONFIG:-crew-dispatch.json models.json crew-harness backlog-backend backend herdr-presentation-spaces startup-memory-budget trace-context}" + +# Items whose value is a home-SESSION enablement decision rather than durable +# local configuration. They are inherited at the launch convergence point, where +# the primary also hands the new process its frozen on/off decision, and left +# untouched by live convergence into an already-running home, whose decision is +# already frozen for its current session (bin/fm-trace-context-lib.sh). +FM_SESSION_SCOPED_INHERITABLE_CONFIG="trace-context" + +# True when is session-scoped in the sense above. +fm_config_inherit_item_session_scoped() { # + local item=$1 candidate + for candidate in $FM_SESSION_SCOPED_INHERITABLE_CONFIG; do + [ "$candidate" = "$item" ] && return 0 + done + return 1 +} + +# The complete declared inherited-material set as home-relative paths, one per +# line, in propagation order: every FM_INHERITABLE_CONFIG item under config/, +# then the one shared data file. This is what remote senders and receivers +# derive from, so both ends of a transfer agree by construction. +fm_config_inherit_items() { + local item + for item in $FM_INHERITABLE_CONFIG; do + printf 'config/%s\n' "$item" + done + printf '%s\n' "$FM_SHARED_CAPTAIN_REL" +} fm_inherit_file_mode() { if [ "$(uname)" = Darwin ]; then @@ -403,6 +454,10 @@ propagate_inheritable_config() { case "$item" in ''|/*|.|..|../*|*/../*|*/..) return 1 ;; esac + if [ "${FM_CONFIG_INHERIT_LIVE:-0}" = 1 ] && fm_config_inherit_item_session_scoped "$item"; then + record_inheritable_config_result "$item" unchanged "session-scoped" + continue + fi src="$src_config/$item" dest="$dest_config/$item" # This one scalar config is consumed as a local safety boundary, so reject diff --git a/bin/fm-config-push.sh b/bin/fm-config-push.sh index b760666dd1b..f57eb1dcb89 100755 --- a/bin/fm-config-push.sh +++ b/bin/fm-config-push.sh @@ -9,11 +9,10 @@ # propagation machinery as bootstrap, but deliberately does not # fast-forward tracked files. # After a successful per-home propagation that changes any allowlisted config/* -# item, writes a generation-specific literal-content reread instruction and -# sends its pointer to that live secondmate via fm-config-inherit-lib.sh -# (fm_config_send_reread_nudge). -# Unchanged config and data/captain-shared.md-only updates send no reread -# message unless a previous send failure is pending for that home. +# item, local routes receive the generation-specific literal-content pointer from +# fm-config-inherit-lib.sh. Remote routes receive one durable marked reread nudge +# through their SSH route. Unchanged config and data/captain-shared.md-only +# updates send no reread unless a previous send failure is pending for that home. # Warnings-only skips exit 0; real propagation or reread-send errors exit non-zero. set -u @@ -26,8 +25,8 @@ live secondmate home. This is local-material-only: - does not fast-forward tracked files - - after successful config/* changes, writes a generation-specific - literal-content reread instruction and sends its pointer to that live secondmate + - after successful config/* changes, sends a local literal-content pointer or + one durable marked remote reread nudge (no message when config is unchanged unless a previous send failure is pending) - reports each live home and each inheritable item as pushed, unchanged, skipped, or error @@ -71,10 +70,14 @@ SECONDMATES_MD="$DATA/secondmates.md" # shellcheck source=bin/fm-ff-lib.sh . "$SCRIPT_DIR/fm-ff-lib.sh" +# shellcheck source=bin/fm-backend.sh +. "$SCRIPT_DIR/fm-backend.sh" # shellcheck source=bin/fm-wake-lib.sh . "$SCRIPT_DIR/fm-wake-lib.sh" # shellcheck source=bin/fm-config-inherit-lib.sh . "$SCRIPT_DIR/fm-config-inherit-lib.sh" +# shellcheck source=bin/fm-secondmate-nudge-lib.sh +. "$SCRIPT_DIR/fm-secondmate-nudge-lib.sh" print_item_report() { local report=$1 item status reason @@ -116,6 +119,57 @@ while IFS='|' read -r id home _window meta; do printf 'secondmate %s: skipped - no home= in %s and no registry home\n' "$id" "$meta" continue fi + remote_host=$(fm_meta_get "$meta" remote_host) + if [ -n "$remote_host" ]; then + printf 'secondmate %s (%s:%s):\n' "$id" "$remote_host" "$home" + remote_lock=$(fm_remote_inherit_transaction_lock_path "$STATE" "$id" 2>/dev/null || true) + if [ -z "$remote_lock" ] || ! fm_lock_acquire_wait "$remote_lock"; then + echo " config-reread: transaction lock failed" + errors=1 + continue + fi + remote_generation=$(fm_remote_inherit_generation_next "$STATE" "$id" 2>/dev/null || true) + if [ -z "$remote_generation" ]; then + echo " config-reread: generation publication failed" + errors=1 + fm_lock_release "$remote_lock" || true + continue + fi + remote_marker=$(fm_secondmate_nudge_marker_path "$STATE" "$id" 2>/dev/null || true) + remote_pending=0 + if [ -f "$remote_marker" ] && [ "$(fm_meta_get "$remote_marker" remote)" = 1 ]; then remote_pending=1; fi + if ! fm_secondmate_nudge_write "$STATE" "$id" "$home" "" remote \ + "$FM_REMOTE_SECOND_MATE_NUDGE_MESSAGE" 1; then + echo " config-reread: retry marker failed" + errors=1 + fm_lock_release "$remote_lock" || true + continue + fi + if remote_out=$(FM_CONFIG_INHERIT_LIVE=1 \ + "$SCRIPT_DIR/fm-remote-inherit-push.sh" "$id" "$remote_generation" 2>&1); then + printf '%s\n' "$remote_out" | sed 's/^/ /' + remote_nudge=0 + if printf '%s\n' "$remote_out" | grep -Eq '^(pushed|removed):'; then remote_nudge=1; fi + [ "$remote_pending" -eq 0 ] || remote_nudge=1 + if [ "$remote_nudge" -eq 1 ]; then + if FM_HOME="$FM_HOME" FM_ROOT_OVERRIDE="$FM_ROOT" FM_STATE_OVERRIDE="$STATE" \ + "$SCRIPT_DIR/fm-send.sh" "fm-$id" "$FM_REMOTE_SECOND_MATE_NUDGE_MESSAGE" >/dev/null 2>&1; then + rm -f -- "$remote_marker" + echo " config-reread: sent" + else + echo " config-reread: send failed; retry retained" + errors=1 + fi + else + rm -f -- "$remote_marker" + fi + else + [ -z "$remote_out" ] || printf '%s\n' "$remote_out" | sed 's/^/ /' + errors=1 + fi + fm_lock_release "$remote_lock" || true + continue + fi if ! validate_secondmate_home "$id" "$home"; then printf 'secondmate %s (%s): skipped - unsafe home: %s\n' "$id" "$home" "$VALIDATION_ERROR" continue @@ -167,7 +221,8 @@ while IFS='|' read -r id home _window meta; do continue } reports="$reports $report" - if FM_CONFIG_INHERIT_REPORT="$report" propagate_secondmate_inheritance "$FM_HOME" "$home_real" "$CONFIG" "$DATA"; then + if FM_CONFIG_INHERIT_REPORT="$report" FM_CONFIG_INHERIT_LIVE=1 \ + propagate_secondmate_inheritance "$FM_HOME" "$home_real" "$CONFIG" "$DATA"; then : else errors=1 diff --git a/bin/fm-context-statusline.sh b/bin/fm-context-statusline.sh new file mode 100755 index 00000000000..3830f372ad3 --- /dev/null +++ b/bin/fm-context-statusline.sh @@ -0,0 +1,130 @@ +#!/usr/bin/env bash +# Render Claude Code's host-computed context-window pressure for a Firstmate +# session without estimating from transcript text. +# Usage: fm-context-statusline.sh [--record ] +# +# Claude Code invokes a statusLine command with its JSON payload on stdin. +# This command reads context_window.remaining_percentage, used_percentage, +# total_tokens, and current_usage. It prints one compact display line and, when +# --record is present, atomically writes the reading plus the derived +# 70%-compaction advisory as JSON for the worker named by a generated brief. +# Only used_percentage and remaining_percentage are required: they alone drive +# the display and the 70% trigger. total_tokens and current_usage are optional; +# when either is missing or invalid, the percentages and trigger still render, +# each missing optional field is named in the display, and the snapshot keeps +# only the fields actually present plus a missing_optional_fields list, so +# degradation is visible and no value is ever invented. +# Input without both valid percentages prints nothing, removes any requested +# stale snapshot, and exits zero so telemetry can never disrupt the host +# session. +# Snapshot persistence is best-effort: the parent directory is recreated when +# missing, and a failed write never suppresses the display line nor removes +# the last valid snapshot. Percentages display truncated to one decimal so a +# reading below 70 never shows as 70 without the compaction advisory. +set -u + +usage() { + awk ' + NR == 1 { next } + /^#/ { sub(/^# ?/, ""); print; next } + { exit } + ' "$0" +} + +record= +case "${1:-}" in + '') ;; + -h|--help) usage; exit 0 ;; + --record) + [ "$#" -eq 2 ] || { usage >&2; exit 2; } + record=$2 + case "$record" in + /*) ;; + *) echo "error: --record requires an absolute path" >&2; exit 2 ;; + esac + ;; + *) usage >&2; exit 2 ;; +esac + +IFS= read -r -d '' node_program <<'NODE' || true +const fs = require("fs"); +const path = require("path"); + +const record = process.argv[1] || ""; + +function clearStaleRecord() { + if (!record) return; + try { + fs.unlinkSync(record); + } catch (_) {} +} + +function finiteNumber(value) { + return typeof value === "number" && Number.isFinite(value); +} + +function formatPercentage(value) { + let text = value.toFixed(1); + if (Number(text) > value) text = (Number(text) - 0.1).toFixed(1); + return text.replace(/\.0$/, ""); +} + +function formatTokens(value) { + if (value >= 1000000) return `${(value / 1000000).toFixed(value % 1000000 === 0 ? 0 : 1)}m`; + if (value >= 1000) return `${(value / 1000).toFixed(value % 1000 === 0 ? 0 : 1)}k`; + return String(value); +} + +try { + const payload = JSON.parse(fs.readFileSync(0, "utf8")); + const contextWindow = payload && payload.context_window; + if (!contextWindow || typeof contextWindow !== "object" || Array.isArray(contextWindow)) throw new Error("missing context_window"); + + const used = contextWindow.used_percentage; + const remaining = contextWindow.remaining_percentage; + const total = contextWindow.total_tokens; + const currentUsage = contextWindow.current_usage; + if (!finiteNumber(used) || used < 0 || used > 100) throw new Error("invalid used_percentage"); + if (!finiteNumber(remaining) || remaining < 0 || remaining > 100) throw new Error("invalid remaining_percentage"); + const hasTotal = finiteNumber(total) && total > 0; + const hasCurrentUsage = Boolean(currentUsage) && typeof currentUsage === "object" && !Array.isArray(currentUsage); + const missingOptional = []; + if (!hasTotal) missingOptional.push("total_tokens"); + if (!hasCurrentUsage) missingOptional.push("current_usage"); + + const compactRecommended = used >= 70; + if (record) { + try { + const parent = path.dirname(record); + fs.mkdirSync(parent, { recursive: true }); + const temporary = path.join(parent, `.${path.basename(record)}.${process.pid}.tmp`); + const recordedWindow = { ...contextWindow }; + if (!hasTotal) delete recordedWindow.total_tokens; + if (!hasCurrentUsage) delete recordedWindow.current_usage; + const snapshot = { + context_window: recordedWindow, + compact_at_used_percentage: 70, + compact_recommended: compactRecommended, + }; + if (missingOptional.length > 0) snapshot.missing_optional_fields = missingOptional; + fs.writeFileSync(temporary, `${JSON.stringify(snapshot, null, 2)}\n`, { mode: 0o600 }); + try { + fs.renameSync(temporary, record); + } catch (error) { + try { fs.unlinkSync(temporary); } catch (_) {} + throw error; + } + } catch (_) {} + } + + let line = `CTX ${formatPercentage(used)}% used / ${formatPercentage(remaining)}% left`; + if (hasTotal) line += ` (${formatTokens(total)})`; + if (missingOptional.length > 0) line += ` [missing: ${missingOptional.join(", ")}]`; + if (compactRecommended) line += " | COMPACT NOW: /compact"; + process.stdout.write(line); +} catch (_) { + clearStaleRecord(); +} +NODE + +node -e "$node_program" "$record" diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 30fc7b7236e..4617eebb32e 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -35,6 +35,10 @@ # checks" from "checks green, waiting on merge" (see nm_ci_checks_state) - # a ci-step log-tail check overrides working -> done once checks read # green, so a green PR is never silently read as still-validating. +# A terminal checks-passed is a REPORTED claim and is corroborated against +# the run's own ci log before it is repeated: a claim the evidence does not +# record reports blocked, because a head no check run examined is not work +# ready for review. See nm_ci_checks_state for the measured defect. # 3. Reconcile the status log: if its last line says needs-decision/blocked but # the run-step shows the run moved on, the log is deterministically stale and # is flagged superseded. A genuinely parked run plus a needs-decision log @@ -64,6 +68,8 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" . "$SCRIPT_DIR/fm-classify-lib.sh" # shellcheck source=bin/fm-busy-lib.sh . "$SCRIPT_DIR/fm-busy-lib.sh" +# shellcheck source=bin/fm-nm-run-lib.sh +. "$SCRIPT_DIR/fm-nm-run-lib.sh" ID=${1:-} [ -n "$ID" ] || { echo "usage: fm-crew-state.sh " >&2; exit 2; } @@ -167,41 +173,20 @@ crew_busy_verdict() { # } # --- no-mistakes run lookup (authoritative when a run matches this branch) -- +# trim, strip_quotes, the bounded nm_run call, nm_field's TOON parse, and the +# branch+head attribution rule below are thin wrappers over the ONE owner in +# bin/fm-nm-run-lib.sh, shared with fm-teardown.sh's pre-teardown run abort. -trim() { - local s=${1:-} - s="${s#"${s%%[![:space:]]*}"}" - s="${s%"${s##*[![:space:]]}"}" - printf '%s' "$s" -} -strip_quotes() { - local s - s=$(trim "${1:-}") - case "$s" in - \"*\") s=${s#\"}; s=${s%\"} ;; - esac - trim "$s" -} - -# Bounded no-mistakes call in the worktree; stdout only, never fails the script. -HAVE_TIMEOUT=none -if command -v timeout >/dev/null 2>&1; then HAVE_TIMEOUT=timeout -elif command -v gtimeout >/dev/null 2>&1; then HAVE_TIMEOUT=gtimeout -elif command -v perl >/dev/null 2>&1; then HAVE_TIMEOUT=perl -fi +trim() { fm_nm_trim "$@"; } +strip_quotes() { fm_nm_strip_quotes "$@"; } nm_run() { # - case "$HAVE_TIMEOUT" in - timeout) ( cd "$WT" && timeout "$NM_TIMEOUT" no-mistakes "$@" ) 2>/dev/null || true ;; - gtimeout) ( cd "$WT" && gtimeout "$NM_TIMEOUT" no-mistakes "$@" ) 2>/dev/null || true ;; - perl) ( cd "$WT" && perl -e 'my $t = shift; my $pid = fork; die "fork failed" unless defined $pid; if (!$pid) { setpgrp(0, 0); exec @ARGV } local $SIG{ALRM} = sub { kill "TERM", -$pid; select undef, undef, undef, 0.2; kill "KILL", -$pid; exit 124 }; alarm $t; waitpid $pid, 0; exit($? >> 8)' "$NM_TIMEOUT" no-mistakes "$@" ) 2>/dev/null || true ;; - *) true ;; - esac + fm_nm_run "$WT" "$NM_TIMEOUT" "$@" } # Scalar value of a TOON key in the captured run output ($RUN_OUT). RUN_OUT="" nm_field() { # - printf '%s\n' "$RUN_OUT" | sed -n "s/^[[:space:]]*$1:[[:space:]]*\(.*\)/\1/p" | head -1 + fm_nm_field "$RUN_OUT" "$1" } # Finding count from a findings[N]{...} table header; empty when none. nm_findings_count() { @@ -308,6 +293,14 @@ nm_effective_ci_step_status() { # for the MOST RECENT recognized marker (the log is append-only/chronological, # so the last match is current): green with nothing red after it means CI is # green right now, still only waiting on merge/close. +# +# "no CI checks reported" is NOT green, and reading it as green is the defect +# measured on 2026-08-02: a cross-repo fork pull request holds its workflows at +# action_required until a maintainer approves them, so zero checks ever run and +# the pipeline reports that absence as a terminal success. Nothing was red +# because nothing executed. An absent verifier is a distinct state from a +# passing one and never maps to green here, so a head no verifier examined +# reaches firstmate as not-ready rather than as work ready for review. nm_ci_checks_state() { local run_id log_tail marker run_id=$(strip_quotes "$(nm_field id)") @@ -318,11 +311,15 @@ nm_ci_checks_state() { | grep -E 'CI checks passed|no CI checks reported - still monitoring|no CI checks reported yet|checks failed|issues detected|CI checks running|base branch advanced.*re-arming CI monitor timeout' \ | tail -1) case "$marker" in - *"checks passed"*|*"no CI checks reported - still monitoring"*) printf 'green' ;; - *"no CI checks reported yet"*|*"checks failed"*|*"issues detected"*|*"CI checks running"*|*"base branch advanced"*"re-arming CI monitor timeout"*) printf 'not-ready' ;; + *"checks passed"*) printf 'green' ;; + *"no CI checks reported - still monitoring"*|*"no CI checks reported yet"*|*"checks failed"*|*"issues detected"*|*"CI checks running"*|*"base branch advanced"*"re-arming CI monitor timeout"*) printf 'not-ready' ;; *) printf 'unknown' ;; esac } + +nm_ci_state_is_green() { + [ "${1:-}" = green ] +} # Coarse fallback for cross-branch attribution. `no-mistakes axi status` (bare) # reports the active-or-most-recent run for the CURRENT branch when one # exists, else falls back to some other branch's run purely as informational @@ -385,40 +382,19 @@ nm_runs_status_for_branch() { # CREW_BRANCH=$(git -C "$WT" symbolic-ref --quiet --short HEAD 2>/dev/null || true) # 0 if the active axi-status run's head field matches this worktree's code -# identity. Branch match is a precondition (caller). Rules: -# - missing/empty head field: cannot bind; reject the run -# - equal commits (short or full SHA): match -# - worktree HEAD is an ancestor of run head: match (pipeline fix commits on -# the same history advanced the run tip) -# - run head is a strict ancestor of worktree HEAD: no match (local work -# advanced outside the run) -# - diverged / run head not in this worktree: no match (rewritten branch tip) +# identity. Branch match is a precondition (caller). Rule owned by +# fm_nm_head_matches_worktree in bin/fm-nm-run-lib.sh. nm_run_head_matches_worktree() { - local run_head local_full run_full + local run_head run_head=$(strip_quotes "$(nm_field head)") - [ -n "$run_head" ] || return 1 - local_full=$(git -C "$WT" rev-parse HEAD 2>/dev/null) || return 1 - run_full=$(git -C "$WT" rev-parse --verify "${run_head}^{commit}" 2>/dev/null) || return 1 - [ "$run_full" = "$local_full" ] && return 0 - if git -C "$WT" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null; then - return 0 - fi - return 1 + fm_nm_head_matches_worktree "$WT" "$run_head" } # Coarse runs-list rows are " ...". 0 if the short # sha for this branch row matches the worktree head under the same rules as # nm_run_head_matches_worktree (equal, or local is ancestor of run tip). nm_coarse_head_matches_worktree() { # - local run_head=$1 local_full run_full - [ -n "$run_head" ] || return 1 - local_full=$(git -C "$WT" rev-parse HEAD 2>/dev/null) || return 1 - run_full=$(git -C "$WT" rev-parse --verify "${run_head}^{commit}" 2>/dev/null) || return 1 - [ "$run_full" = "$local_full" ] && return 0 - if git -C "$WT" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null; then - return 0 - fi - return 1 + fm_nm_head_matches_worktree "$WT" "$1" } HAVE_RUN=0 @@ -488,7 +464,29 @@ if [ "$HAVE_RUN" = 1 ]; then if [ -n "$outcome" ]; then case "$outcome" in passed) RUN_STATE="done"; RUN_DETAIL="run passed: PR merged/closed" ;; - checks-passed) RUN_STATE="done"; RUN_DETAIL="checks green: PR ready for review" ;; + # checks-passed is the pipeline's REPORTED terminal claim, and on + # 2026-08-02 it was reported for a head whose check-run set was empty. + # Corroborate it against the run's own ci log before repeating it: a + # claim of green that the run's own evidence does not record is not a + # green head, and a task whose checks never ran needs firstmate rather + # than a place in the ready-for-review queue. + checks-passed) + CI_LOG_STATE=$(nm_ci_checks_state) + if nm_ci_state_is_green "$CI_LOG_STATE"; then + RUN_STATE="done"; RUN_DETAIL="checks green: PR ready for review" + else + case "$CI_LOG_STATE" in + not-ready) + RUN_STATE=blocked + RUN_DETAIL="run reported a passing result its own CI log does not record: nothing verified this head" + ;; + unknown|"") + RUN_STATE=unknown + RUN_DETAIL="run reported checks-passed, but its CI log is unavailable: claim could not be corroborated" + ;; + esac + fi + ;; failed) RUN_STATE=failed; RUN_DETAIL="run failed" ;; cancelled) RUN_STATE=failed; RUN_DETAIL="run cancelled" ;; *) RUN_STATE=unknown; RUN_DETAIL="outcome: $outcome" ;; @@ -523,7 +521,7 @@ if [ "$HAVE_RUN" = 1 ]; then case "$CI_STEP_STATUS" in running) CI_LOG_STATE=$(nm_ci_checks_state) - if [ "$CI_LOG_STATE" = green ]; then + if nm_ci_state_is_green "$CI_LOG_STATE"; then RUN_STATE="done" RUN_DETAIL="checks green: PR ready for review (still monitoring for merge/close)" fi @@ -538,18 +536,23 @@ if [ "$HAVE_RUN" = 1 ]; then if [ "$RUN_STATE" = working ] && log_reports_ci_ready; then if [ "$RUN_SOURCE" = coarse ]; then - emit "done" status-log "$(status_line_note "$LOG_LINE")${SEP}run still monitoring PR" - fi - [ -n "$CI_STEP_STATUS" ] || CI_STEP_STATUS=$(nm_effective_ci_step_status) - if [ "$RUN_STATUS" = fixing ]; then - CI_LOG_STATE=not-ready - elif [ "$CI_STEP_STATUS" = running ] && [ -z "$CI_LOG_STATE" ]; then - CI_LOG_STATE=$(nm_ci_checks_state) - elif [ "$CI_STEP_STATUS" = fixing ]; then - CI_LOG_STATE=not-ready - fi - if [ "$CI_LOG_STATE" != not-ready ]; then - emit "done" status-log "$(status_line_note "$LOG_LINE")${SEP}run still monitoring PR" + RUN_STATE=unknown + RUN_DETAIL="status log reported readiness, but coarse run data cannot corroborate the claim" + else + [ -n "$CI_STEP_STATUS" ] || CI_STEP_STATUS=$(nm_effective_ci_step_status) + if [ "$RUN_STATUS" = fixing ]; then + CI_LOG_STATE=not-ready + elif [ "$CI_STEP_STATUS" = running ] && [ -z "$CI_LOG_STATE" ]; then + CI_LOG_STATE=$(nm_ci_checks_state) + elif [ "$CI_STEP_STATUS" = fixing ]; then + CI_LOG_STATE=not-ready + fi + if nm_ci_state_is_green "$CI_LOG_STATE"; then + emit "done" status-log "$(status_line_note "$LOG_LINE")${SEP}run still monitoring PR" + elif [ -z "$CI_LOG_STATE" ] || [ "$CI_LOG_STATE" = unknown ]; then + RUN_STATE=unknown + RUN_DETAIL="status log reported readiness, but CI evidence is unavailable: claim could not be corroborated" + fi fi fi diff --git a/bin/fm-ff-lib.sh b/bin/fm-ff-lib.sh index bce9993ae78..438f10f0b10 100644 --- a/bin/fm-ff-lib.sh +++ b/bin/fm-ff-lib.sh @@ -414,6 +414,7 @@ sweep_live_secondmate_metas() { local state=$1 base_mode=$2 nudge_requires_instr=${3:-no} registry=${4:-$FM_HOME/data/secondmates.md} id home window meta [ -d "$state" ] || return 0 while IFS='|' read -r id home window meta; do + if grep -q '^remote_host=.' "$meta" 2>/dev/null; then continue; fi process_secondmate "$id" "$home" "$window" "$base_mode" "$nudge_requires_instr" done < <(live_secondmate_meta_records "$state" "$registry") } diff --git a/bin/fm-fleet-snapshot.sh b/bin/fm-fleet-snapshot.sh index f8598d9ea19..88e604abe31 100755 --- a/bin/fm-fleet-snapshot.sh +++ b/bin/fm-fleet-snapshot.sh @@ -55,6 +55,13 @@ # # Compatibility: JSON is the primary machine-readable surface. # Human views must render this output instead of parsing state files again. +# +# Exit status is part of the contract. +# 0 means the snapshot is complete, and an EMPTY fleet is a complete snapshot +# rather than a failure. +# Any read, projection, or assembly failure exits nonzero, names what failed on +# stderr, and prints no snapshot, so a caller that branches on exit status can +# never read a failed read as a healthy empty fleet. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -142,6 +149,8 @@ usage: fm-fleet-snapshot.sh --json Print a read-only structured snapshot of the firstmate fleet. JSON is the stable machine-readable output contract. +An empty fleet is a complete snapshot and exits 0; a failed read exits nonzero, +says what failed, and prints no snapshot. --secondmate-home-summary emits the bounded structured summary used after a validated registered-home handoff. It is local-only, skips nested secondmate @@ -174,6 +183,42 @@ esac command -v jq >/dev/null 2>&1 || { echo "fm-fleet-snapshot: jq not found" >&2; exit 1; } +# --- bulk JSON transport ---------------------------------------------------- +# +# Linux caps a SINGLE argv entry at MAX_ARG_STRLEN (131072 bytes), independent +# of the much larger ARG_MAX total, so `jq --argjson big "$json"` dies with +# "Argument list too long" as soon as one value crosses that cap - which a +# fleet-sized backlog, a long status stream, or a secondmate roll-up already +# does. Bulk values therefore travel on stdin, which has no per-value cap. +# +# stdin is used rather than a temp file or a process substitution because this +# command is read-only and must stay that way: there is no file to create, +# secure, or leave behind when an interrupted read abandons it, no temp-directory +# dependency, and unlike --slurpfile the filter receives the value itself instead +# of a one-element array it has to unwrap. +# +# Only bulk values move; small scalars stay on argv as --arg/--argjson, where +# they read more clearly at the call site. + +# json_envelope [ ]... +# Assemble the stdin envelope from values that are ALREADY valid JSON, then pipe +# it straight into `jq -n`, whose filter reads it back with `input`. Any `def` +# definitions still come first, exactly as in a normal jq program. +# +# Shell variables and function arguments are ordinary memory rather than execve +# arguments, and printf is a builtin, so assembly never puts bulk data on argv +# either. A malformed envelope - the shape an empty or failed intermediate +# produces - makes jq fail rather than yield a plausible-looking partial record. +json_envelope() { + local out='{' sep='' + while [ "$#" -ge 2 ]; do + out=$out$sep'"'$1'":'$2 + sep=',' + shift 2 + done + printf '%s}' "$out" +} + bool_json() { if [ "$1" = 1 ]; then printf 'true'; else printf 'false'; fi } @@ -399,10 +444,18 @@ backlog_json() { # [] - defaults to this home's $BACKLOG task_json_lines() { local meta id kind harness mode yolo project worktree home projects backend target status_log report_path + local remote_host remote_root remote_state remote_rc remote_home_present local pr pr_source event_json current_json endpoint_exists agent_alive meta_json status_json report_json worktree_json home_json - local last_event_raw current_state current_source pending_decision blocked_event report_present=0 pr_from_status + local current_state current_source pending_decision blocked_event report_present=0 pr_from_status local open_decisions_tsv open_decisions_json + # pipefail is scoped to this subshell so a row that cannot be serialized fails + # the whole read. Without it the row generator's failure is swallowed by the + # trailing `jq -s`, which happily slurps the surviving rows and exits 0 - so a + # task would disappear from a snapshot that still reported success, and a + # supervisor reviewing the fleet would see a task that does not exist. + ( + set -o pipefail for meta in "$STATE"/*.meta; do [ -e "$meta" ] || continue id=$(basename "$meta" .meta) @@ -415,8 +468,17 @@ task_json_lines() { worktree=$(meta_value "$meta" worktree) home=$(meta_value "$meta" home) projects=$(meta_value "$meta" projects) - backend=$(fm_backend_of_meta "$meta") - target=$(fm_backend_target_of_meta "$meta") + remote_host=$(meta_value "$meta" remote_host) + remote_root=$(meta_value "$meta" remote_root) + remote_home_present=null + if [ -n "$remote_host" ]; then + backend=$(meta_value "$meta" remote_backend) + [ -n "$backend" ] || backend=unknown + target=$(meta_value "$meta" remote_target) + else + backend=$(fm_backend_of_meta "$meta") + target=$(fm_backend_target_of_meta "$meta") + fi status_log="$STATE/$id.status" report_path="$DATA/$id/report.md" pr=$(meta_value "$meta" pr) @@ -432,7 +494,6 @@ task_json_lines() { current_json=$(crew_state_json "$id") event_json=$(status_event_json "$status_log") - last_event_raw=$(printf '%s' "$event_json" | jq -r '.last_event.raw // ""') current_state=$(printf '%s' "$current_json" | jq -r '.state // ""') current_source=$(printf '%s' "$current_json" | jq -r '.source // ""') @@ -468,16 +529,38 @@ task_json_lines() { blocked_event=$(printf '%s' "$open_decisions_json" | jq 'if any(.[]; .verb == "blocked") then 1 else 0 end') endpoint_exists=null - if [ -n "$target" ]; then - if fm_backend_target_exists "$backend" "$target" "fm-$id" >/dev/null 2>&1; then - endpoint_exists=true + agent_alive=not_checked + if [ -n "$remote_host" ]; then + if remote_state=$(run_timed "$FM_SNAPSHOT_SECONDMATE_TIMEOUT" \ + "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-secondmate-control.sh state "$id" < /dev/null 2>/dev/null); then + remote_rc=0 else - endpoint_exists=false + remote_rc=$? + fi + if [ "$remote_rc" -eq 0 ]; then + remote_home_present=true + remote_state=$(printf '%s\n' "$remote_state" | tail -1) + case "$remote_state" in + alive) endpoint_exists=true; agent_alive=alive ;; + dead) endpoint_exists=true; agent_alive=dead ;; + missing) endpoint_exists=false; agent_alive=dead ;; + *) endpoint_exists=null; agent_alive=unknown ;; + esac + else + endpoint_exists=null + agent_alive=unknown + fi + else + if [ -n "$target" ]; then + if fm_backend_target_exists "$backend" "$target" "fm-$id" >/dev/null 2>&1; then + endpoint_exists=true + else + endpoint_exists=false + fi + fi + if [ "$kind" = secondmate ] && [ -n "$target" ]; then + agent_alive=$(fm_backend_agent_alive "$backend" "$target" 2>/dev/null || printf unknown) fi - fi - agent_alive=not_checked - if [ "$kind" = secondmate ] && [ -n "$target" ]; then - agent_alive=$(fm_backend_agent_alive "$backend" "$target" 2>/dev/null || printf unknown) fi [ -f "$report_path" ] && report_present=1 || report_present=0 @@ -485,9 +568,25 @@ task_json_lines() { status_json=$event_json report_json=$(path_present_json "$report_path") if [ -n "$worktree" ]; then worktree_json=$(path_present_json "$worktree"); else worktree_json=$(jq -n '{path:null,present:false}'); fi - if [ -n "$home" ]; then home_json=$(path_present_json "$home"); else home_json=$(jq -n '{path:null,present:false}'); fi + if [ -n "$home" ] && [ -n "$remote_host" ]; then + home_json=$(jq -n --arg path "$home" --argjson present "$remote_home_present" '{path:$path,present:$present}') + elif [ -n "$home" ]; then + home_json=$(path_present_json "$home") + else + home_json=$(jq -n '{path:null,present:false}') + fi - jq -n \ + # The status stream, its keyed open-decision fold, and the reconciled + # current state are all unbounded, so they travel on stdin. + json_envelope \ + current_state "$current_json" \ + meta_path "$meta_json" \ + status_log "$status_json" \ + report "$report_json" \ + worktree_path "$worktree_json" \ + home_path "$home_json" \ + open_decisions "$open_decisions_json" \ + | jq -n \ --arg id "$id" \ --arg kind "$kind" \ --arg harness "$harness" \ @@ -499,23 +598,25 @@ task_json_lines() { --arg projects "$projects" \ --arg backend "$backend" \ --arg target "$target" \ + --arg remote_host "$remote_host" \ + --arg remote_root "$remote_root" \ --arg pr "$pr" \ --arg pr_source "$pr_source" \ --arg agent_alive "$agent_alive" \ --arg observed_at "$SNAPSHOT_NOW" \ - --arg last_event_raw "$last_event_raw" \ - --argjson current_state "$current_json" \ - --argjson meta_path "$meta_json" \ - --argjson status_log "$status_json" \ - --argjson report "$report_json" \ - --argjson worktree_path "$worktree_json" \ - --argjson home_path "$home_json" \ --argjson endpoint_exists "$endpoint_exists" \ - --argjson open_decisions "$open_decisions_json" \ --argjson pending_decision "$(bool_json "$pending_decision")" \ --argjson blocked_event "$(bool_json "$blocked_event")" \ --argjson report_present "$(bool_json "$report_present")" \ - '{ + 'input as $in + | $in.current_state as $current_state + | $in.meta_path as $meta_path + | $in.status_log as $status_log + | $in.report as $report + | $in.worktree_path as $worktree_path + | $in.home_path as $home_path + | $in.open_decisions as $open_decisions + | { id:$id, kind:$kind, harness:($harness // ""), @@ -523,6 +624,7 @@ task_json_lines() { yolo:($yolo // ""), project:($project // ""), backend:$backend, + remote:(if $remote_host == "" then null else {host:$remote_host,root:$remote_root} end), paths:{ meta:$meta_path, status_log:$status_log, @@ -543,7 +645,7 @@ task_json_lines() { blocked_event:$blocked_event, open_decisions:$open_decisions, scout_report_present:$report_present, - last_event_text:$last_event_raw + last_event_text:($status_log.last_event.raw // "") }, actions:( if $kind == "secondmate" then @@ -555,8 +657,9 @@ task_json_lines() { steer:"bin/fm-send.sh fm-\($id) \u0027\u0027", return_channel_note:null} end) - }' + }' || exit 1 done | jq -s 'sort_by(.id)' + ) } # Main-home current-inventory validity: same orphan / unstructured-current checks @@ -564,10 +667,12 @@ task_json_lines() { # Meta inventory remains the sole source of live workers; this object only # discloses backlog↔task inconsistency for renderers (Bearings omitted/gates). main_inventory_json() { # - jq -n \ - --argjson backlog "$1" \ - --argjson tasks "$2" ' - ([ $backlog.records[]? + json_envelope backlog "$1" tasks "$2" \ + | jq -n ' + input as $in + | $in.backlog as $backlog + | $in.tasks as $tasks + | ([ $backlog.records[]? | select((.state == "in_flight" or .state == "queued") and (.structured | not)) ]) as $unstructured_current | ([ $backlog.records[]? | select(.state == "in_flight" and .structured and .requires_child_metadata) ]) as $owned_in_flight @@ -592,19 +697,21 @@ main_inventory_json() { # # This mode never reads parent events or terminal text and never aggregates # nested secondmates. secondmate_home_summary_json() { # - jq -n \ + json_envelope backlog "$1" tasks "$2" \ + | jq -n \ --arg generated "$SNAPSHOT_NOW" \ --arg home "$FM_HOME" \ --argjson child_n "$FM_SNAPSHOT_SECONDMATE_CHILDREN" \ --argjson queued_n "$FM_SNAPSHOT_SECONDMATE_QUEUED" \ --argjson decisions_n "$FM_SNAPSHOT_SECONDMATE_DECISIONS" \ - --argjson landed_n "$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" \ - --argjson backlog "$1" \ - --argjson tasks "$2" ' + --argjson landed_n "$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" ' def trunc($n): tostring | gsub("\\s+"; " ") | if length > $n then .[:$n] + "…" else . end; - ([ $backlog.records[]? + input as $in + | $in.backlog as $backlog + | $in.tasks as $tasks + | ([ $backlog.records[]? | select((.state == "in_flight" or .state == "queued") and (.structured | not)) ]) as $unstructured_current | ([ $backlog.records[]? | select(.state == "in_flight" and .structured) ]) as $owned_in_flight | ([ $backlog.records[]? @@ -839,9 +946,13 @@ BASH | select(startswith("- ")) | (capture("^- (?[^[:space:]]+)")?) as $id | select($id != null) - | (capture("^.*\\(home:[[:space:]]*(?[^;)]*);[[:space:]]*scope:[[:space:]]*.*;[[:space:]]*projects:[[:space:]]*[^;)]*;[[:space:]]*added[[:space:]]+[0-9]{4}-[0-9]{2}-[0-9]{2}\\)[[:space:]]*$")?) as $home - | {id:$id.id,home:($home.home // null),registered:true, - registry_error:(if $home == null or ($home.home | length) == 0 then "registry entry has no home" else null end)} ] + | ([capture("^.*\\(host:[[:space:]]*(?[^;)]*);[[:space:]]*root:[[:space:]]*(?[^;)]*);[[:space:]]*home:[[:space:]]*(?[^;)]*);[[:space:]]*scope:[[:space:]]*.*;[[:space:]]*projects:[[:space:]]*[^;)]*;[[:space:]]*added[[:space:]]+[0-9]{4}-[0-9]{2}-[0-9]{2}\\)[[:space:]]*$")?][0] // null) as $remote + | ([capture("^.*\\(home:[[:space:]]*(?[^;)]*);[[:space:]]*scope:[[:space:]]*.*;[[:space:]]*projects:[[:space:]]*[^;)]*;[[:space:]]*added[[:space:]]+[0-9]{4}-[0-9]{2}-[0-9]{2}\\)[[:space:]]*$")?][0] // null) as $local + | ($local // $remote) as $route + | (($local == null) and ($remote != null)) as $is_remote + | {id:$id.id,home:($route.home // null),host:(if $is_remote then $remote.host else null end),root:(if $is_remote then $remote.root else null end), + remote:$is_remote,registered:true, + registry_error:(if $route == null or ($route.home | length) == 0 then "registry entry has no home" else null end)} ] | group_by(.id) | map(if length > 1 then .[0] + {registry_error:"duplicate secondmate id in registry"} else .[0] end) JQ @@ -963,10 +1074,16 @@ BASH } terminal_evidence_json() { # - local task=$1 note=$2 evidence_contradicts=$3 backend target exists expected out rc clean bytes lines seen=false contradiction=false reason='' + local task=$1 note=$2 evidence_contradicts=$3 backend target exists expected out rc clean bytes lines seen=false contradiction=false reason='' remote_host backend=$(printf '%s' "$task" | jq -r '.backend // ""') target=$(printf '%s' "$task" | jq -r '.endpoint.target // ""') exists=$(printf '%s' "$task" | jq -r '.endpoint.exists // "unknown"') + remote_host=$(printf '%s' "$task" | jq -r '.remote.host // ""') + if [ -n "$remote_host" ]; then + jq -n --arg observed "$SNAPSHOT_NOW" --arg reason "remote terminal evidence is not collected by the primary" \ + '{provenance:"remote-direct-report-terminal",trust:"untrusted-supplement",captured:false,observed_at:$observed,freshness:"not-collected",reason:$reason,lines:0,bytes:0,event_note_seen:false,contradiction:false}' + return 0 + fi expected=$(printf '%s' "$task" | jq -r '"fm-" + (.id // "")') if [ -z "$target" ] || [ "$exists" = false ]; then [ "$exists" = false ] && reason="recorded endpoint is absent" || reason="no recorded endpoint" @@ -1011,7 +1128,8 @@ terminal_evidence_json() { # - jq -n --argjson summary "$1" --argjson activities "$2" --argjson decisions "$3" ' + json_envelope summary "$1" activities "$2" decisions "$3" \ + | jq -n ' def keyed: . != null and . != "" and . != "default"; def result($e; $matches; $complete; $surface): $e + { @@ -1022,7 +1140,11 @@ parent_evidence_reconciliation_json() { # local tasks=$1 registry union rows total_registered total shown truncated - local row id home registered registry_error task status_file event_raw event_note event_epoch event_age + local row id home host remote registered registry_error task status_file event_raw event_note event_epoch event_age local activity_scan activities decisions reconciliation provenance freshness reason summary summary_rc summary_bytes summary_valid summary_reason summary_invalidity state current_reason terminal terminal_contradiction contradiction - local records='[]' seen_homes='' + local records='' records_json seen_homes='' registry=$(registry_secondmates_json) || return 1 - union=$(jq -n --argjson registry "$registry" --argjson tasks "$tasks" ' - ($registry.records // []) as $registered + union=$(json_envelope registry "$registry" tasks "$tasks" \ + | jq -n ' + input as $in + | $in.registry as $registry + | $in.tasks as $tasks + | ($registry.records // []) as $registered | (($registered | map(.id)) // []) as $registered_ids | ([ $registered[] as $r | $r + {parent_task:([$tasks[] | select(.id == $r.id)][0] // null)} ] @@ -1091,9 +1217,9 @@ secondmate_current_json() { # parent_task:$t} ]) | sort_by(.id) | {registry:$registry,records:.}') || return 1 - total_registered=$(printf '%s' "$union" | jq '[.records[] | select(.registered)] | length') - total=$(printf '%s' "$union" | jq '.records | length') - rows=$(printf '%s' "$union" | jq -c --argjson cap "$FM_SNAPSHOT_SECONDMATES" '(if $cap == 0 then .records else .records[:$cap] end)[]') + total_registered=$(printf '%s' "$union" | jq '[.records[] | select(.registered)] | length') || return 1 + total=$(printf '%s' "$union" | jq '.records | length') || return 1 + rows=$(printf '%s' "$union" | jq -c --argjson cap "$FM_SNAPSHOT_SECONDMATES" '(if $cap == 0 then .records else .records[:$cap] end)[]') || return 1 shown=$(printf '%s\n' "$rows" | grep -c . || true) truncated=$((total - shown)) @@ -1101,6 +1227,8 @@ secondmate_current_json() { # [ -n "$row" ] || continue id=$(printf '%s' "$row" | jq -r '.id') home=$(printf '%s' "$row" | jq -r '.home // ""') + host=$(printf '%s' "$row" | jq -r '.host // ""') + remote=$(printf '%s' "$row" | jq -r '.remote // false') registered=$(printf '%s' "$row" | jq -r '.registered') registry_error=$(printf '%s' "$row" | jq -r '.registry_error // ""') task=$(printf '%s' "$row" | jq -c '.parent_task // {}') @@ -1128,40 +1256,53 @@ secondmate_current_json() { # esac fi if [ -z "$reason" ]; then - if ! validate_secondmate_home "$id" "$home" 2>/dev/null; then + if [ "$remote" = true ]; then + [ -n "$host" ] || reason="invalid remote route: missing SSH host" + case " $seen_homes " in + *" $host:$home "*) reason="invalid home: duplicate resolved remote route" ;; + *) seen_homes="$seen_homes $host:$home" ;; + esac + elif ! validate_secondmate_home "$id" "$home" 2>/dev/null; then reason="invalid home: $VALIDATION_ERROR" else home=$VALIDATED_HOME case " $seen_homes " in - *" $home "*) reason="invalid home: duplicate resolved home route" ;; - *) seen_homes="$seen_homes $home" ;; + *" local:$home "*) reason="invalid home: duplicate resolved home route" ;; + *) seen_homes="$seen_homes local:$home" ;; esac fi fi if [ -z "$reason" ]; then - summary=$(run_timed "$FM_SNAPSHOT_SECONDMATE_TIMEOUT" env \ - FM_ROOT_OVERRIDE="$FM_ROOT" \ - FM_HOME="$home" \ - FM_STATE_OVERRIDE="$home/state" \ - FM_DATA_OVERRIDE="$home/data" \ - FM_CONFIG_OVERRIDE="$home/config" \ - FM_PROJECTS_OVERRIDE="$home/projects" \ - FM_SNAPSHOT_NOW="$SNAPSHOT_NOW" \ - FM_SNAPSHOT_NOW_EPOCH="$SNAPSHOT_EPOCH" \ - FM_SNAPSHOT_SECONDMATE_CHILDREN="$FM_SNAPSHOT_SECONDMATE_CHILDREN" \ - FM_SNAPSHOT_SECONDMATE_QUEUED="$FM_SNAPSHOT_SECONDMATE_QUEUED" \ - FM_SNAPSHOT_SECONDMATE_DECISIONS="$FM_SNAPSHOT_SECONDMATE_DECISIONS" \ - FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME="$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" \ - "$SCRIPT_DIR/fm-fleet-snapshot.sh" --secondmate-home-summary 2>/dev/null) - summary_rc=$? + if [ "$remote" = true ]; then + summary=$(run_timed "$FM_SNAPSHOT_SECONDMATE_TIMEOUT" \ + "$SCRIPT_DIR/fm-on.sh" "$id" fm-fleet-snapshot.sh --secondmate-home-summary < /dev/null 2>/dev/null) + summary_rc=$? + else + summary=$(run_timed "$FM_SNAPSHOT_SECONDMATE_TIMEOUT" env \ + FM_ROOT_OVERRIDE="$FM_ROOT" \ + FM_HOME="$home" \ + FM_STATE_OVERRIDE="$home/state" \ + FM_DATA_OVERRIDE="$home/data" \ + FM_CONFIG_OVERRIDE="$home/config" \ + FM_PROJECTS_OVERRIDE="$home/projects" \ + FM_SNAPSHOT_NOW="$SNAPSHOT_NOW" \ + FM_SNAPSHOT_NOW_EPOCH="$SNAPSHOT_EPOCH" \ + FM_SNAPSHOT_SECONDMATE_CHILDREN="$FM_SNAPSHOT_SECONDMATE_CHILDREN" \ + FM_SNAPSHOT_SECONDMATE_QUEUED="$FM_SNAPSHOT_SECONDMATE_QUEUED" \ + FM_SNAPSHOT_SECONDMATE_DECISIONS="$FM_SNAPSHOT_SECONDMATE_DECISIONS" \ + FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME="$FM_SNAPSHOT_SECONDMATE_LANDED_PER_HOME" \ + "$SCRIPT_DIR/fm-fleet-snapshot.sh" --secondmate-home-summary 2>/dev/null) + summary_rc=$? + fi if [ "$summary_rc" -ne 0 ]; then [ "$summary_rc" -eq 124 ] && reason="structured home snapshot timed out" || reason="structured home snapshot failed" else summary_bytes=$(printf '%s' "$summary" | LC_ALL=C wc -c | tr -d ' ') if [ "$summary_bytes" -gt "$FM_SNAPSHOT_SECONDMATE_MAX_BYTES" ]; then reason="structured home snapshot exceeded byte limit" - elif ! printf '%s' "$summary" | jq -e --arg home "$home" --arg generated "$SNAPSHOT_NOW" ' - .schema == "fm-secondmate-home-summary.v1" and .home == $home and .generated == $generated + elif ! printf '%s' "$summary" | jq -e --arg home "$home" --arg generated "$SNAPSHOT_NOW" --argjson remote "$remote" ' + .schema == "fm-secondmate-home-summary.v1" and .home == $home + and (($remote == true) or .generated == $generated) and (.valid | type) == "boolean" and (.state | type) == "string" and (.invalidity | type) == "object" and (.invalidity.ids | type) == "array" and (.active_children | type) == "array" and (.decisions_open | type) == "array" @@ -1200,13 +1341,26 @@ secondmate_current_json() { # '{provenance:"parent-direct-report-terminal",trust:"untrusted-supplement",captured:false,observed_at:$observed,freshness:"not-collected",reason:"no useful contradiction check",lines:0,bytes:0,event_note_seen:false,contradiction:false}') fi if printf '%s' "$terminal" | jq -e '.contradiction == true' >/dev/null; then contradiction=true; fi - record=$(jq -n \ - --arg id "$id" --arg home "$home" --arg state "$state" --arg current_reason "$current_reason" --arg observed "$SNAPSHOT_NOW" \ - --argjson registered "$registered" --argjson summary "$summary" --argjson summary_valid "$summary_valid" --argjson decisions "$decisions" \ - --argjson activities "$activities" --argjson activity_scan "$activity_scan" \ - --argjson reconciliation "$reconciliation" --argjson terminal "$terminal" --argjson contradiction "$contradiction" \ + record=$(json_envelope \ + summary "$summary" \ + decisions "$decisions" \ + activities "$activities" \ + activity_scan "$activity_scan" \ + reconciliation "$reconciliation" \ + terminal "$terminal" \ + | jq -n -c \ + --arg id "$id" --arg home "$home" --arg host "$host" --argjson remote "$remote" --arg state "$state" --arg current_reason "$current_reason" --arg observed "$SNAPSHOT_NOW" \ + --argjson registered "$registered" --argjson summary_valid "$summary_valid" \ + --argjson contradiction "$contradiction" \ --arg event_raw "$event_raw" --arg event_note "$event_note" --argjson event_age "$event_age" ' - {id:$id,home:$home,registered:$registered, + input as $in + | $in.summary as $summary + | $in.decisions as $decisions + | $in.activities as $activities + | $in.activity_scan as $activity_scan + | $in.reconciliation as $reconciliation + | $in.terminal as $terminal + | {id:$id,home:$home,host:($host | if . == "" then null else . end),remote:$remote,registered:$registered, current:{state:$state,reason:($current_reason | if . == "" then null else . end)},invalidity:$summary.invalidity, provenance:{selected:"structured-home",structured_home:$home,summary_valid:$summary_valid, trust:(if $summary_valid then "complete" else "partial-structured" end),parent_event_role:"historical-only"}, @@ -1215,7 +1369,7 @@ secondmate_current_json() { # decisions_open:$summary.decisions_open,holds:$summary.holds,queued:$summary.queued, landed:$summary.landed,endpoints:$summary.endpoints,counts:$summary.counts,omitted:$summary.omitted, parent_event:{raw:$event_raw,note:$event_note,age_seconds:$event_age,open_activities:$activities,open_decisions:$decisions,activity_scan:$activity_scan,reconciliation:$reconciliation}, - terminal_evidence:$terminal,contradiction:$contradiction}') + terminal_evidence:$terminal,contradiction:$contradiction}') || return 1 else if [ -n "$event_raw" ]; then provenance='parent-event-fallback' @@ -1230,36 +1384,54 @@ secondmate_current_json() { # terminal=$(jq -n --arg observed "$SNAPSHOT_NOW" \ '{provenance:"parent-direct-report-terminal",trust:"untrusted-supplement",captured:false,observed_at:$observed,freshness:"not-collected",reason:"no parent event to compare",lines:0,bytes:0,event_note_seen:false,contradiction:false}') fi - record=$(jq -n \ - --arg id "$id" --arg home "$home" --arg reason "$reason" --arg observed "$SNAPSHOT_NOW" \ + record=$(json_envelope \ + decisions "$decisions" \ + activities "$activities" \ + activity_scan "$activity_scan" \ + terminal "$terminal" \ + | jq -n -c \ + --arg id "$id" --arg home "$home" --arg host "$host" --argjson remote "$remote" --arg reason "$reason" --arg observed "$SNAPSHOT_NOW" \ --arg provenance "$provenance" --arg freshness "$freshness" --arg event_raw "$event_raw" --arg event_note "$event_note" \ - --argjson registered "$registered" --argjson event_age "$event_age" --argjson activities "$activities" --argjson activity_scan "$activity_scan" \ - --argjson decisions "$decisions" --argjson terminal "$terminal" ' - {id:$id,home:($home | if . == "" then null else . end),registered:$registered, + --argjson registered "$registered" --argjson event_age "$event_age" ' + input as $in + | $in.decisions as $decisions + | $in.activities as $activities + | $in.activity_scan as $activity_scan + | $in.terminal as $terminal + | {id:$id,home:($home | if . == "" then null else . end),host:($host | if . == "" then null else . end),remote:$remote,registered:$registered, current:{state:"unknown",reason:$reason},invalidity:null, provenance:{selected:$provenance,structured_home:($home | if . == "" then null else . end),parent_event_role:"fallback-only-not-current"}, freshness:{status:$freshness,observed_at:$observed,age_seconds:$event_age}, active_children:[],decisions_open:[],holds:[],queued:[],landed:[],endpoints:[],counts:{active_children:0,decisions_open:0,holds:0,queued:0,landed:0,endpoints:0},omitted:[], parent_event:{raw:$event_raw,note:$event_note,age_seconds:$event_age,open_activities:$activities,open_decisions:$decisions,activity_scan:$activity_scan}, - terminal_evidence:$terminal,contradiction:false}') + terminal_evidence:$terminal,contradiction:false}') || return 1 fi - records=$(jq -n --argjson records "$records" --argjson record "$record" '$records + [$record]') + # Each record is one compact line, collected in the shell and slurped once + # below. Re-parsing the whole accumulated array per row would put it back on + # argv and cost O(n^2) parses for no benefit. + records=$records$record$'\n' done < - jq -n --argjson current "$1" ' - {records:[ $current.records[] + json_envelope current "$1" \ + | jq -n ' + input as $in + | $in.current as $current + | {records:[ $current.records[] | select(.provenance.selected == "structured-home") as $mate | $mate.landed[] | . + {home:$mate.home,home_id:$mate.id}], @@ -1281,13 +1453,19 @@ scout_report_lines() { jq -n '[]' return 0 fi - LC_ALL=C find "$DATA" -mindepth 2 -maxdepth 2 -type f -name report.md -print \ - | sort \ - | while IFS= read -r report; do - id=$(basename "$(dirname "$report")") - jq -n --arg id "$id" --arg path "$report" '{id:$id,path:$path}' - done \ - | jq -s 'sort_by(.id)' + # pipefail again: an unreadable data directory, or a row that cannot be + # serialized, must fail the scan rather than reporting the reports it did + # happen to reach as the complete set. + ( + set -o pipefail + LC_ALL=C find "$DATA" -mindepth 2 -maxdepth 2 -type f -name report.md -print \ + | sort \ + | while IFS= read -r report; do + id=$(basename "$(dirname "$report")") + jq -n --arg id "$id" --arg path "$report" '{id:$id,path:$path}' || exit 1 + done \ + | jq -s 'sort_by(.id)' + ) } BACKLOG_JSON=$(backlog_json) || { echo "fm-fleet-snapshot: backlog read failed" >&2; exit 1; } @@ -1299,7 +1477,8 @@ if [ "$OUTPUT_MODE" = secondmate-home-summary ]; then exit 0 fi -SCOUT_REPORTS_JSON=$(scout_report_lines) +SCOUT_REPORTS_JSON=$(scout_report_lines) \ + || { echo "fm-fleet-snapshot: scout report scan failed" >&2; exit 1; } MAIN_INVENTORY_JSON=$(main_inventory_json "$BACKLOG_JSON" "$TASKS_JSON") \ || { echo "fm-fleet-snapshot: main inventory summary failed" >&2; exit 1; } SECONDMATE_CURRENT_JSON=$(secondmate_current_json "$TASKS_JSON") \ @@ -1307,7 +1486,14 @@ SECONDMATE_CURRENT_JSON=$(secondmate_current_json "$TASKS_JSON") \ SECONDMATE_LANDED_JSON=$(secondmate_landed_from_current_json "$SECONDMATE_CURRENT_JSON") \ || { echo "fm-fleet-snapshot: secondmate landed projection failed" >&2; exit 1; } -jq -n \ +json_envelope \ + backlog "$BACKLOG_JSON" \ + tasks "$TASKS_JSON" \ + main_inventory "$MAIN_INVENTORY_JSON" \ + scout_reports "$SCOUT_REPORTS_JSON" \ + secondmate_current "$SECONDMATE_CURRENT_JSON" \ + secondmate_landed "$SECONDMATE_LANDED_JSON" \ + | jq -n \ --arg generated "$SNAPSHOT_NOW" \ --arg fm_home "$FM_HOME" \ --arg fm_root "$FM_ROOT" \ @@ -1315,13 +1501,14 @@ jq -n \ --arg data "$DATA" \ --arg config "$CONFIG" \ --arg projects "$PROJECTS" \ - --argjson backlog "$BACKLOG_JSON" \ - --argjson tasks "$TASKS_JSON" \ - --argjson main_inventory "$MAIN_INVENTORY_JSON" \ - --argjson scout_reports "$SCOUT_REPORTS_JSON" \ - --argjson secondmate_current "$SECONDMATE_CURRENT_JSON" \ - --argjson secondmate_landed "$SECONDMATE_LANDED_JSON" \ - 'def backlog_by_id($id): ($backlog.records[]? | select(.structured == true and .id == $id) | .) // null; + 'input as $in + | $in.backlog as $backlog + | $in.tasks as $tasks + | $in.main_inventory as $main_inventory + | $in.scout_reports as $scout_reports + | $in.secondmate_current as $secondmate_current + | $in.secondmate_landed as $secondmate_landed + | def backlog_by_id($id): ($backlog.records[]? | select(.structured == true and .id == $id) | .) // null; def task_by_id($id): ($tasks[]? | select(.id == $id) | .) // null; def report_kind($id): (task_by_id($id).kind // backlog_by_id($id).kind // "scout"); { @@ -1338,4 +1525,4 @@ jq -n \ secondmate_guidance:{ note:"For kind=secondmate, bearings selects validated structured state from that registered home; parent events and bounded terminal evidence are fallback-only supplements and never current-state authority." } - }' + }' || { echo "fm-fleet-snapshot: snapshot assembly failed" >&2; exit 1; } diff --git a/bin/fm-fleet-view.sh b/bin/fm-fleet-view.sh index 909c792b294..d30647b0852 100755 --- a/bin/fm-fleet-view.sh +++ b/bin/fm-fleet-view.sh @@ -4,6 +4,11 @@ # This command intentionally does not parse fleet state itself. # It shells out to fm-fleet-snapshot.sh --json and renders that stable # structured contract for humans. +# +# It inherits that command's exit-status contract: an empty fleet renders and +# exits 0, while a snapshot that failed or produced nothing exits nonzero and +# says so, because supervision reviews the fleet from this view and must never +# read a failed read as a healthy fleet. set -u SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -26,7 +31,20 @@ esac command -v jq >/dev/null 2>&1 || { echo "fm-fleet-view: jq not found" >&2; exit 1; } -SNAPSHOT=$("$SCRIPT_DIR/fm-fleet-snapshot.sh" --json) || exit $? +# A failed snapshot must never render as a healthy fleet. An empty FLEET is +# valid and still produces a full snapshot object, so empty OUTPUT can only mean +# the snapshot did not complete - and reporting that as success would let fleet +# supervision degrade silently as the fleet and backlog grow. +SNAPSHOT=$("$SCRIPT_DIR/fm-fleet-snapshot.sh" --json) +SNAPSHOT_RC=$? +if [ "$SNAPSHOT_RC" -ne 0 ]; then + echo "fm-fleet-view: fleet snapshot failed (exit $SNAPSHOT_RC); no fleet state was read" >&2 + exit "$SNAPSHOT_RC" +fi +if [ -z "$SNAPSHOT" ]; then + echo "fm-fleet-view: fleet snapshot produced no output; no fleet state was read" >&2 + exit 1 +fi printf '%s\n' "$SNAPSHOT" | jq -r ' def dash($v): if $v == null or $v == "" then "-" else $v end; @@ -93,4 +111,4 @@ printf '%s\n' "$SNAPSHOT" | jq -r ' "", "## Secondmates", .secondmate_guidance.note -' +' || { echo "fm-fleet-view: rendering the fleet snapshot failed" >&2; exit 1; } diff --git a/bin/fm-guard.sh b/bin/fm-guard.sh index 5698d376ab0..24151de92eb 100755 --- a/bin/fm-guard.sh +++ b/bin/fm-guard.sh @@ -6,12 +6,18 @@ # non-default branch, because that means firstmate-on-itself work landed in the # primary instead of an isolated worktree. # Then, if a task is in flight (a state/.meta exists) or X-mode relay -# polling is active (state/x-watch.check.sh exists) and no identity-matched -# watcher has a liveness beacon (state/.last-watcher-beat, touched every poll -# cycle) fresh within FM_GUARD_GRACE seconds, prints a loud, clearly delimited -# banner so the agent cannot skim past it in the tool output of whatever it was -# doing - the one channel every harness has. The full banner is emitted once per -# distinct staleness episode in this FM_HOME (keyed to beacon mtime or absence); +# polling is active (state/x-watch.check.sh exists) and supervision is not +# healthy, prints a loud, clearly delimited banner so the agent cannot skim past +# it in the tool output of whatever it was doing - the one channel every harness +# has. Supervision health is MODEL-AWARE (fm_watcher_supervision_verdict in +# bin/fm-wake-lib.sh): under the Claude Stop auto-arm model the watcher runs only +# between turns, so mid-turn a fresh beacon with no live watcher is healthy and +# only a stale beacon (beyond FM_GUARD_GRACE) is a genuine lapse; under every +# persistent-watcher harness a live identity-matched watcher with a fresh beacon +# is required. The banner names the true failing condition (a missing live +# watcher process vs a genuinely stale beacon). The full banner is emitted once +# per distinct down-episode in this FM_HOME (keyed to the failing condition, not +# the beacon mtime, which a healthy between-turns watcher advances every poll); # later guarded commands in the same episode print a one-line reminder instead. # Episode state lives only under state/.guard-watcher-stale-banner (volatile, # bounded). Independent alarms (queued wakes, worktree tangle) are never @@ -43,18 +49,14 @@ STALE_BANNER_MARKER="$STATE/.guard-watcher-stale-banner" # shellcheck source=bin/fm-supervision-lib.sh . "$SCRIPT_DIR/fm-supervision-lib.sh" -# Deterministic episode key from beacon state: same continuous stale beacon -# (or continuous absence) shares a key; a recovered-then-restale beacon gets a -# new mtime and therefore a new episode. +# Deterministic episode key from the qualitative down-state (the failing +# condition), NOT the beacon mtime: under the auto-arm model a healthy +# between-turns watcher advances that mtime every poll, which made the "same +# episode" key change every turn and re-print the full banner. Keying on the +# failing condition keeps one continuous down-episode stable, while positive +# recovery clears the marker (below) and re-arms the next episode. fm_guard_stale_episode_key() { - local state=$1 beat m - beat="$state/.last-watcher-beat" - if [ -e "$beat" ]; then - m=$(fm_sup_stat_mtime "$beat") - printf 'beat:%s\n' "${m:-unknown}" - else - printf 'beat:absent\n' - fi + printf '%s\n' "$1" } # Claim the full banner for this episode. Exit 0 = print full banner (this call @@ -150,10 +152,9 @@ in_flight=$FM_SUP_IN_FLIGHT sources=$FM_SUP_SOURCES needed=$FM_SUP_NEEDED beacon_desc=$FM_SUP_BEACON_DESC -watcher_healthy=false -if fm_watcher_healthy "$STATE" "$WATCH" "$GRACE" "$FM_HOME"; then - watcher_healthy=true -fi +fm_watcher_supervision_verdict "$STATE" "$WATCH" "$GRACE" "$FM_HOME" +watcher_healthy=$FM_WATCHER_VERDICT_OK +watcher_down_reason=$FM_WATCHER_VERDICT_REASON if [ "$needed" = false ]; then # Leave the unhealthy state (nothing riding on the watcher): clear so a later # work or X-mode need + stale combination is a fresh episode even if the @@ -168,7 +169,7 @@ fi # bordered banner FIRST so it reads as an alarm, not a buried stderr line. Later # calls in the same episode get a one-line reminder only. if [ "$watcher_healthy" = false ]; then - episode_key=$(fm_guard_stale_episode_key "$STATE") + episode_key=$(fm_guard_stale_episode_key "$watcher_down_reason") episode_key=${episode_key%$'\n'} print_full_banner=0 if [ "$READ_ONLY" -eq 1 ]; then @@ -193,12 +194,17 @@ if [ "$watcher_healthy" = false ]; then { printf '●%s\n' "$rule" printf '● WATCHER DOWN - SUPERVISION IS OFF\n' + if [ "$watcher_down_reason" = no-watcher ]; then + watcher_cause=$(printf 'no live watcher process holds this home lock (last beat: %s)' "$beacon_desc") + else + watcher_cause=$(printf 'no watcher has a fresh beacon (last beat: %s, grace %ss)' "$beacon_desc" "$GRACE") + fi if [ "$in_flight" -gt 0 ]; then - printf '● %s task(s) in flight, but no watcher has a fresh beacon (last beat: %s, grace %ss).\n' "$in_flight" "$beacon_desc" "$GRACE" + printf '● %s task(s) in flight, but %s.\n' "$in_flight" "$watcher_cause" elif [ "$sources" -gt 0 ]; then - printf '● %s process-event source(s) registered, but no watcher has a fresh beacon (last beat: %s, grace %ss).\n' "$sources" "$beacon_desc" "$GRACE" + printf '● %s process-event source(s) registered, but %s.\n' "$sources" "$watcher_cause" else - printf '● X-mode relay polling needs supervision, but no watcher has a fresh beacon (last beat: %s, grace %ss).\n' "$beacon_desc" "$GRACE" + printf '● X-mode relay polling needs supervision, but %s.\n' "$watcher_cause" fi if [ "$READ_ONLY" -eq 1 ]; then printf '● This read-only session should report the lapse, not repair it.\n' diff --git a/bin/fm-harness.sh b/bin/fm-harness.sh index 824b95804de..110f7c64a4e 100755 --- a/bin/fm-harness.sh +++ b/bin/fm-harness.sh @@ -53,9 +53,19 @@ detect_own() { *codex*) echo codex; return ;; *opencode*) echo opencode; return ;; *grok*) echo grok; return ;; - kimi) echo kimi; return ;; + # Substring, like the four harnesses above: a basename match is already + # narrow, and an exact arm here made this the ONLY consumer that missed a + # variant name. bin/fm-session-lock-lib.sh and bin/backends/tmux.sh both + # matched kimi as a substring, so a `kimi-nightly` process was recognized + # by those two while self-detection returned `unknown`. + *kimi*) echo kimi; return ;; + # The pi family stays EXACT on purpose: a bare `pi` substring would also + # match pip, pipenv, and any path component containing "pi". pi-launcher + # and Pi are the launcher wrapper and the npm shim basename; both were + # already recognized by bin/backends/tmux.sh's pane classifier but not + # here, so a firstmate running under either self-detected as `unknown`. pi-signed) echo pi; return ;; - pi) echo pi; return ;; + pi|pi-launcher|Pi) echo pi; return ;; node*|python*) # Bare interpreter: match the harness name in its script path. args=$(ps -o args= -p "$pid" 2>/dev/null) diff --git a/bin/fm-home-seed.sh b/bin/fm-home-seed.sh index 4b5b9815bcc..46227b76482 100755 --- a/bin/fm-home-seed.sh +++ b/bin/fm-home-seed.sh @@ -37,58 +37,21 @@ FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" PROJECTS="${FM_PROJECTS_OVERRIDE:-$FM_HOME/projects}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" REG="$DATA/secondmates.md" SUB_HOME_MARKER=".fm-secondmate-home" # shellcheck source=bin/fm-secondmate-registry-lib.sh . "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" +# shellcheck source=bin/fm-secondmate-charter-lib.sh +. "$SCRIPT_DIR/fm-secondmate-charter-lib.sh" +# shellcheck source=bin/fm-wake-lib.sh +. "$SCRIPT_DIR/fm-wake-lib.sh" usage() { echo "usage: fm-home-seed.sh {...|--no-projects}" >&2 echo " fm-home-seed.sh validate" >&2 } -normalize_registry_text() { - awk ' - { - gsub(/[;()]/, " ") - gsub(/[[:space:]]+/, " ") - sub(/^ /, "") - sub(/ $/, "") - if ($0 != "") { - out = out (out == "" ? "" : " ") $0 - } - } - END { print out } - ' -} - -brief_section_text() { - local brief=$1 heading=$2 - awk -v heading="# $heading" ' - $0 == heading { in_section=1; next } - in_section && /^# / { exit } - in_section { print } - ' "$brief" -} - -registry_summary_for_brief() { - local brief=$1 - if [ -n "${FM_SECONDMATE_CHARTER:-}" ]; then - printf '%s\n' "$FM_SECONDMATE_CHARTER" | normalize_registry_text - else - brief_section_text "$brief" "Charter" | normalize_registry_text - fi -} - -registry_scope_for_brief() { - local brief=$1 - if [ -n "${FM_SECONDMATE_SCOPE:-}" ]; then - printf '%s\n' "$FM_SECONDMATE_SCOPE" | normalize_registry_text - else - brief_section_text "$brief" "Routing scope" | normalize_registry_text - fi -} - validate_registry_home_text() { local home=$1 case "$home" in @@ -517,6 +480,20 @@ EOF SEED_ROLLBACK_ACTIVE=0 SEED_COMMITTED=0 +SEED_REGISTRY_LOCK= +SEED_REGISTRY_LOCK_HELD=0 + +seed_registry_lock_release() { + if [ "$SEED_REGISTRY_LOCK_HELD" -eq 1 ]; then + fm_lock_release "$SEED_REGISTRY_LOCK" + SEED_REGISTRY_LOCK_HELD=0 + fi +} + +seed_exit_cleanup() { + seed_rollback + seed_registry_lock_release +} SEED_HOME= SEED_HOME_ACQUIRED=0 SEED_HOME_CREATED=0 @@ -822,6 +799,12 @@ seed_home() { [ $# -gt 0 ] || { echo "error: secondmate needs at least one project, or --no-projects for a project-less home" >&2; return 1; } fi + mkdir -p "$STATE" || return 1 + SEED_REGISTRY_LOCK=$(secondmate_registry_lock_path "$STATE") + fm_lock_acquire_wait "$SEED_REGISTRY_LOCK" || return 1 + SEED_REGISTRY_LOCK_HELD=1 + trap seed_exit_cleanup EXIT + validate_registry for project in "$@"; do validate_seed_project "$project" @@ -844,7 +827,6 @@ seed_home() { SEED_SUB_REG_EXISTED=0 SEED_CHARTER_EXISTED=0 SEED_MARKER_EXISTED=0 - trap seed_rollback EXIT if [ -f "$REG" ]; then SEED_PARENT_REG_EXISTED=1 cp "$REG" "$SEED_BACKUP_DIR/parent-secondmates.md" @@ -939,6 +921,7 @@ seed_home() { write_registry "$id" "$home" "$projects_csv" "$SEED_PARENT_BRIEF" validate_registry SEED_COMMITTED=1 + seed_registry_lock_release trap - EXIT rm -rf -- "$SEED_BACKUP_DIR" printf 'home=%s\n' "$home" diff --git a/bin/fm-landed-lib.sh b/bin/fm-landed-lib.sh new file mode 100644 index 00000000000..0442d8dff27 --- /dev/null +++ b/bin/fm-landed-lib.sh @@ -0,0 +1,157 @@ +#!/usr/bin/env bash +# fm-landed-lib.sh - the single owner of "has this content already landed?" +# +# WHY CONTENT, NOT COMMIT REACHABILITY +# `rev-list --count ..HEAD` answers "is this COMMIT reachable from ", +# which a squash merge, a rebase, or a local replay all break: the content is in +# the trunk under a different commit id, so the count stays non-zero forever. +# The question that actually protects work is "does already contain what +# HEAD introduces". A 3-way merge answers it: when HEAD adds nothing does +# not already have, the merged tree IS 's tree. Trunk commits past the +# merge-base do not count as "added", so an out-of-date slot still reads landed. +# +# WHY THE LANDING TARGET IS NOT ALWAYS refs/remotes/origin/ +# A remote-tracking ref is only the landing target for a fleet that pushes to +# that remote and refetches it. Two shapes in real use break that assumption: +# - `origin` FETCHES from an upstream but PUSHES to a fork, so origin/ +# tracks upstream while the fork trunk (the local branch) is where work +# lands. +# - A project that lands locally by design never advances origin/ at +# all, so it drifts further behind every day. +# Callers therefore ask about a REF THEY CHOOSE. This library does not guess a +# landing target; fm_landed_default_branch_name resolves only the NAME, and each +# caller applies its own policy for which refs carrying that name to test. +# +# WHY THE FORK TRUNK NEEDS A REF OF ITS OWN +# In the fetch/push split above, NEITHER conventional ref is the fork trunk once +# the fork advances at the forge: origin/ is upstream, and the local +# branch is the fork trunk only while something keeps fast-forwarding it. When +# that lags, work provably merged into the fork reads as unlanded and holds its +# slot forever. The default fetch refspec cannot reach the fork - it points at +# the fetch URL - so the trunk is fetched from the PUSH url into a ref of its +# own under refs/fm-landing/. That ref is a landing target only because this +# fleet demonstrably pushes there; it is never inferred from a remote's name. +# fm_landed_refresh_push_target does the network read and is called only by a +# caller that already refreshes remotes, so the purely local guard path keeps +# whatever the last refresh left behind and never grows a network dependency. +# +# EXIT STATUS IS THREE-VALUED, ON PURPOSE +# fm_landed_tree_contains distinguishes "proven contained" from "proven not +# contained" from "could not tell". Collapsing the last two loses the difference +# between work that is demonstrably unlanded and a repository this could not +# read, and callers guarding real work need to refuse loudly on both while +# reporting them differently. +# +# Every read uses --no-optional-locks so inspecting a worktree that another lane +# owns never takes or rewrites its index lock. +set -u + +# 0 when already contains everything HEAD introduces, 1 when it provably +# does not, 2 when that could not be determined (unreadable ref, unreadable +# HEAD, or a merge conflict, which means the trees genuinely diverge but is +# reported as inconclusive rather than as a clean "not contained"). +fm_landed_tree_contains() { # + local dir=$1 ref=$2 ref_tree merged status + ref_tree=$(git --no-optional-locks -C "$dir" rev-parse --quiet --verify "$ref^{tree}" 2>/dev/null) || return 2 + [ -n "$ref_tree" ] || return 2 + git --no-optional-locks -C "$dir" rev-parse --quiet --verify HEAD >/dev/null 2>&1 || return 2 + merged=$(git --no-optional-locks -C "$dir" merge-tree --write-tree "$ref" HEAD 2>/dev/null) + status=$? + # merge-tree exits non-zero on conflict AND on usage/read failure. Both are + # inconclusive here: a conflict proves the trees diverge, but this predicate + # only ever claims "demonstrably contained", never "demonstrably unlanded". + [ "$status" -eq 0 ] || return 2 + merged=$(printf '%s\n' "$merged" | head -1) + [ -n "$merged" ] || return 2 + [ "$merged" = "$ref_tree" ] && return 0 + return 1 +} + +# The default branch NAME (no refs/ prefix), preferring what origin/HEAD records +# and otherwise the first conventional name that exists as either a local branch +# or a remote-tracking branch. Non-zero when no name resolves, which every +# caller must treat as unverifiable rather than as "nothing to protect". +fm_landed_default_branch_name() { # + local dir=$1 ref branch + ref=$(git --no-optional-locks -C "$dir" symbolic-ref --quiet --short refs/remotes/origin/HEAD 2>/dev/null || true) + if [ -n "$ref" ]; then + printf '%s\n' "${ref#origin/}" + return 0 + fi + for branch in main master; do + if git --no-optional-locks -C "$dir" rev-parse --verify --quiet "refs/heads/$branch" >/dev/null 2>&1 \ + || git --no-optional-locks -C "$dir" rev-parse --verify --quiet "refs/remotes/origin/$branch" >/dev/null 2>&1; then + printf '%s\n' "$branch" + return 0 + fi + done + return 1 +} + +# Normalize a remote URL for comparison: trailing slash and .git suffix only. +# Deliberately not a URL parser - scheme/host differences are real differences. +fm_landed_normalize_url() { # + local url=${1%/} + printf '%s\n' "${url%.git}" +} + +# The url this repo PUSHES to when that differs from where `origin` FETCHES, +# which is the fetch/push split described above and the layout firstmate itself +# uses. Returns 1 when there is no origin or the two urls agree, which is every +# ordinary single-remote repository and the conventional `upstream`-remote fork +# layout - in both, origin/ already tracks what this repo pushes. +fm_landed_push_url() { # + local dir=$1 fetch push + fetch=$(git --no-optional-locks -C "$dir" remote get-url origin 2>/dev/null) || return 1 + push=$(git --no-optional-locks -C "$dir" remote get-url --push origin 2>/dev/null) || return 1 + [ "$(fm_landed_normalize_url "$fetch")" != "$(fm_landed_normalize_url "$push")" ] || return 1 + printf '%s\n' "$push" +} + +# The private ref holding the push remote's trunk, or 1 when this repo +# has no distinct push url. Naming it does not make it exist; a caller that +# needs it populated calls fm_landed_refresh_push_target first, and a caller +# that only reads local state tests it like any other candidate. +fm_landed_push_target_ref() { # + fm_landed_push_url "$1" >/dev/null || return 1 + printf 'refs/fm-landing/origin/%s\n' "$2" +} + +# Refresh the push remote's trunk into that private ref. TOUCHES THE NETWORK, so +# only a caller that already refreshes remotes should call it. +# 0 when there was nothing to refresh (no distinct push url) or the refresh +# succeeded; 1 when a distinct push url exists but its trunk could not be read. +# A caller must treat that 1 as unverifiable and refuse: the landing target is +# precisely the ref it could not read, so nothing it can still reach is evidence +# that the work landed. +fm_landed_refresh_push_target() { # + local dir=$1 name=$2 url ref + url=$(fm_landed_push_url "$dir") || return 0 + ref=$(fm_landed_push_target_ref "$dir" "$name") || return 0 + git -C "$dir" fetch --quiet "$url" "+refs/heads/$name:$ref" >/dev/null 2>&1 || return 1 + return 0 +} + +# Every existing ref carrying that could be a landing target, most-local +# first, one per line. The local branch leads because the two fleet shapes above +# both land there; the push remote's trunk follows, because on a fetch/push +# split it is where work actually lands and the local branch only mirrors it +# when something keeps that up to date; origin/ is last. A caller that +# needs the remote to win can reorder, and a caller that needs ALL of them +# tested reads the whole list. Empty output (non-zero) means the name resolved +# but no ref carrying it exists. +fm_landed_candidate_refs() { # + local dir=$1 name=$2 found=1 ref push_ref + local -a candidates=("refs/heads/$name") + if push_ref=$(fm_landed_push_target_ref "$dir" "$name"); then + candidates+=("$push_ref") + fi + candidates+=("refs/remotes/origin/$name") + for ref in "${candidates[@]}"; do + if git --no-optional-locks -C "$dir" rev-parse --verify --quiet "$ref" >/dev/null 2>&1; then + printf '%s\n' "$ref" + found=0 + fi + done + return "$found" +} diff --git a/bin/fm-launch-lib.sh b/bin/fm-launch-lib.sh new file mode 100644 index 00000000000..4306cebfeb4 --- /dev/null +++ b/bin/fm-launch-lib.sh @@ -0,0 +1,302 @@ +#!/usr/bin/env bash +# fm-launch-lib.sh - the single owner of firstmate's verified launch commands. +# +# Every firstmate-launched agent session composes its command from exactly these +# three functions. There is no second copy anywhere, and a caller must never +# hand-write a launch string: the drift that causes is not hypothetical. A +# downstream registry once hand-copied claude's command as +# `claude --dangerously-skip-permissions`, dropping the ghost-text suppression +# variable documented in launch_template() below - the exact omission that makes +# firstmate read predicted-prompt text as real typed input when it captures a +# pane. One owner, or that happens again. +# +# Sourced by bin/fm-spawn.sh (crewmate, scout, and secondmate sessions). +# +# launch_template [] the verified launch command, with +# placeholders the caller substitutes +# model_flag_for_harness resolved --model flag, or empty +# effort_flag_for_harness resolved effort flag, or empty +# +# The knowledge half of each adapter (busy-state source, exit command, dialogs, +# quirks) lives in the harness-adapters skill, not here. +# +# shell_quote lives here because both flag resolvers depend on it; sourcing this +# library is what makes it available to bin/fm-spawn.sh. + +shell_quote() { + printf "'" + printf '%s' "$1" | sed "s/'/'\\\\''/g" + printf "'" +} + +# The verified launch command per adapter, as a template. A non-zero return has +# two distinct causes, and a caller reporting the refusal must tell them apart: +# 1. the harness has no verified adapter at all - the unverified-adapter guard +# every caller relies on, whose remedy is a raw launch command; +# 2. the harness is verified but this kind is deliberately unsupported for it - +# today only kimi with kind=primary, whose remedy is a different harness, +# NOT the raw-launch escape hatch. +# Never add a permissive default arm to either case. bin/fm-spawn.sh:466 and :470 +# name only cause 1 because fm-spawn never passes kind=primary; a consumer that +# does pass it owes the user the cause-2 wording. +# +# kind selects the session shape: +# ship|scout a crewmate working one task in an isolated worktree +# secondmate a firstmate PRIMARY launched in a provisioned secondmate home +# primary a firstmate PRIMARY launched in this home by the fleet launcher +# +# ship, scout, and secondmate all receive a launch brief, so their templates end +# in the encoded brief argument. A primary has no task, no worktree, no brief, +# and no status file, so it launches bare and is greeted by the session-start +# adapters already installed in the home (for pi and opencode those are the +# project-local extensions the harness auto-discovers once trusted, which is why +# a primary needs no explicit extension flag). A primary template therefore +# carries only its flag placeholders plus whatever briefless-launch flag that +# adapter was verified to need, and an unset flag leaves one trailing space; that +# is cosmetic in a shell command, and consumers may trim it. +# +# Placeholders every caller substitutes before launch: +# __MODELFLAG__ model_flag_for_harness output, or empty (see below) +# __EFFORTFLAG__ effort_flag_for_harness output, or empty (see below) +# __KIMIBIN__ shell-quoted absolute path to the resolved kimi binary +# Placeholders only a task-scoped (ship|scout|secondmate) launch substitutes: +# __BRIEF__ absolute path to data//brief.md +# __TURNEND__ absolute path to state/.turn-ended (for harnesses whose +# turn-end signal rides the launch command, e.g. codex -c notify=[...]) +# __PIEXT__ absolute path to state/.pi-ext.ts (pi turn-end extension, +# written by fm-spawn.sh; outside the worktree to avoid pi's trust gate) +# __PITURNEND__ absolute path to .pi/extensions/fm-primary-turnend-guard.ts in a pi secondmate home +# __PIWATCH__ absolute path to .pi/extensions/fm-primary-pi-watch.ts in a pi secondmate home +# __OPINPUT__ absolute path to the canonical operational-input encoder +# +# __KIMIBIN__ is resolved by bin/fm-spawn.sh alone, deliberately: the fleet +# launcher reaches Kimi through the pi harness rather than a native kimi binary, +# so there is no second caller to drift from. +# Revisit that only if a native kimi launch ever becomes a launcher entry. +# That same reasoning is why kimi has no primary arm: only bin/fm-spawn.sh can +# substitute __KIMIBIN__, and it only ever launches crewmates, so a primary kimi +# template could not be substituted by the caller that would ask for it. +# +# No primary template below is the crewmate command with its brief argument +# subtracted; each arm cites the specific in-repo evidence that fixes its flags, +# and tests/fm-launch-lib.test.sh pins each one against those same citations. +# The evidence is not uniform, and each arm says which kind it rests on: opencode +# and grok are pinned to an empirical briefless PRIMARY launch in a live e2e test, +# pi to the documented bare launch, and claude and codex to the secondmate +# precedent - this file's own `secondmate` kind is a firstmate PRIMARY (see the +# kind list above), and its shipped crewmate templates - the claude arm and the +# codex `kind = secondmate` arm in the second case block below - launch that +# interactive primary with exactly the autonomy flags those two arms carry. +# +# CONSUMER OBLIGATION (binding, not advisory). The rule, which governs whatever +# the templates below happen to say: EVERY primary template here starts a session +# that runs without permission prompts, and a consumer composing a primary launch +# MUST surface that to the captain at launch time. One short line at launch or in +# the menu row is enough. The captain is entitled to know the posture of the +# session their front door starts. There are no exempt adapters. The rule binds on +# the posture, not on the presence of a particular flag, so a consumer cannot +# satisfy the letter of this note while silently shipping a no-prompt session. +# +# All six reach that posture, four by an explicit bypass and the pi family +# structurally: +# claude --dangerously-skip-permissions +# codex --dangerously-bypass-approvals-and-sandbox +# opencode OPENCODE_CONFIG_CONTENT='{"permission":{"*":"allow"}}', which +# pre-allows every permission before the TUI starts +# grok --always-approve, which .agents/skills/harness-adapters/SKILL.md:314 +# records as auto-approving every tool execution, verified to run +# fully unattended and equivalent to --permission-mode bypassPermissions +# pi (and pi-signed, which shares its arm) no flag, because none exists +# to pass: SKILL.md:277 records that pi has no permission system at +# all, so a pi session is autonomous by construction. Its template - +# the bare selected binary plus the FM_PI_HARNESS identity marker, +# which changes no permission posture - is complete, NOT missing an +# autonomy flag its siblings carry - do not add one. +# Pi's first-run project trust dialog (SKILL.md:286) is folder trust, not a +# permission prompt, and is a separate concern that neither satisfies nor softens +# this obligation. +launch_template() { + local harness=$1 kind=${2:-ship} + # shellcheck disable=SC2016 # single quotes are deliberate: $(cat ...) expands in the crewmate pane, not here + case "$kind" in + primary) + case "$harness" in + # The ghost-text suppression prefix is firstmate-required on every claude + # launch, primary included (see the crewmate arm below for why). + # --dangerously-skip-permissions is settled knowledge, decided and recorded + # rather than inherited by accident. The evidence is the secondmate + # precedent: the claude arm in the crewmate case block below serves every + # crewmate kind INCLUDING secondmate, and this `secondmate` kind is itself a + # firstmate PRIMARY launched in a provisioned home, so the repo already + # ships an interactive firstmate primary carrying this flag. The captain's + # own attended session already runs as `claude --dangerously-skip-permissions`, + # so keeping it preserves the status quo instead of creating new exposure, + # and dropping it would break the supervision contract: a firstmate stalled + # on a permission prompt cannot run bin/fm-wake-drain.sh to drain its wake + # queue or bin/fm-watch-arm.sh to arm its own watcher. Note that README.md:90 + # documents the primary launch as bare `claude` and the only in-repo claude + # launch carrying the flag directly is the headless print-mode session at + # tests/fm-claude-stop-autoarm-live-e2e.test.sh:115; the secondmate + # precedent, not those, is what fixes this arm. See the CONSUMER OBLIGATION + # in the header: a consumer must tell the captain this session has no + # permission prompts. + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false claude --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__' ;; + # --dangerously-bypass-approvals-and-sandbox rests on the same secondmate + # precedent: the codex `kind = secondmate` arm in the crewmate case block + # below is the SECONDMATE template, and it launches an + # interactive firstmate primary with exactly this flag. The same status-quo + # and supervision-contract reasoning as the claude arm applies. It is worth + # recording what is NOT the evidence here: + # tests/fm-codex-continuity-live-e2e.test.sh:40 runs `codex exec`, headless, + # so it says nothing about the interactive primary TUI shape. The header's + # CONSUMER OBLIGATION covers this arm too. + codex) printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox' ;; + # --prompt carries the crewmate's brief, so it has no place in a briefless + # primary; --auto is the empirically verified briefless form (a primary + # opencode TUI is launched that way in + # tests/fm-opencode-primary-live-e2e.test.sh:256 and :310). The + # OPENCODE_CONFIG_CONTENT JSON pre-allows every permission, so the header's + # CONSUMER OBLIGATION covers this arm too. + opencode) printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' opencode __MODELFLAG__--auto' ;; + # Bare `pi` is the documented primary launch (README.md:102); the project + # trust prompt approved once per clone is what makes the tracked + # .pi/extensions/*.ts auto-load (README.md:108), so a primary needs no + # explicit -e flag. tests/fm-pi-primary-live-e2e.test.sh:266 adds + # --approve --no-session --no-context-files --no-extensions with explicit + # -e paths, but those are that test's isolation scaffolding - it runs + # against a throwaway clone - not the verified primary form, so they are + # deliberately not copied here. This template carries no autonomy flag + # because pi has none to carry: SKILL.md:277 records that pi has no + # permission system, so the session is autonomous by construction. The + # header's CONSUMER OBLIGATION therefore covers this arm like every other - + # a pi primary runs without permission prompts too, it just gets there + # structurally rather than by a bypass flag. Nothing is missing here. + # The FM_PI_HARNESS identity marker rides every Pi-family launch, primary + # included (README.md:104 documents the signed primary as + # `FM_PI_HARNESS=pi-signed pi-signed`): the selected $harness is both the + # invoked binary and the marker, so a signed primary's environment cannot + # relabel a plain Pi session. + pi|pi-signed) printf '%s%s' "FM_PI_HARNESS=$harness $harness" ' __MODELFLAG____EFFORTFLAG__' ;; + # --trust is supervision-safety knowledge, not one-time setup trivia: + # without folder trust the primary turn-end guard FAILS OPEN + # (.agents/skills/harness-adapters/SKILL.md:355), and because trust is + # granted once per clone a fresh clone is exactly when its absence bites + # (README.md:107, docs/turnend-guard.md:71). The empirical primary launch + # is tests/fm-grok-continuity-live-e2e.test.sh:76, + # `grok --trust --always-approve --reasoning-effort low`, where + # --reasoning-effort is what __EFFORTFLAG__ resolves to; README.md:96 + # documents the same `grok --trust`. --always-approve auto-approves every + # tool execution (.agents/skills/harness-adapters/SKILL.md:314), so the + # header's CONSUMER OBLIGATION covers this arm: a consumer must tell the + # captain this session has no permission prompts. + grok) printf '%s' 'grok --trust --always-approve __MODELFLAG____EFFORTFLAG__' ;; + # kimi refuses rather than emitting an unsubstitutable command: README.md:61 + # lists only Claude Code, Grok, Pi, pi-signed, Codex, and OpenCode as verified + # primary harnesses (docs/configuration.md:191 defers that narrower set to README), + # and only bin/fm-spawn.sh can resolve __KIMIBIN__ (see the header above). + # A non-zero return is the same refusal an unverified adapter gets. + kimi) return 1 ;; + *) return 1 ;; + esac + return 0 + ;; + esac + # shellcheck disable=SC2016 # single quotes are deliberate: $(cat ...) expands in the crewmate pane, not here + case "$harness" in + # CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false disables claude's interactive + # predicted-next-prompt ghost text, which renders as dim/faint text inside an + # otherwise-empty composer and would otherwise read like real typed input when + # firstmate captures the pane (see the harness-adapters skill). It is a per-launch env + # prefix scoped to this firstmate-launched agent; it never touches the captain's + # global config. The CLI's --prompt-suggestions flag is print/SDK-mode only and + # does NOT suppress the interactive ghost text (verified empirically), so the env + # var is the correct control. The dim-aware composer reader in fm-tmux-lib.sh is + # the defense-in-depth backstop for any pane this flag cannot reach. + claude) printf '%s' 'CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false claude --dangerously-skip-permissions __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + codex) + if [ "$kind" = secondmate ]; then + printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + else + printf '%s' 'codex __MODELFLAG____EFFORTFLAG__--dangerously-bypass-approvals-and-sandbox -c "notify=[\"bash\",\"-c\",\"touch __TURNEND__\"]" "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + fi + ;; + opencode) printf '%s' 'OPENCODE_CONFIG_CONTENT='\''{"permission":{"*":"allow"}}'\'' opencode __MODELFLAG__--prompt "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # pi-signed is a distinct executable identity that shares pi's verified flag + # surface, never an alias: the selected $harness is both the invoked binary + # and the FM_PI_HARNESS identity marker, so a signed primary's environment + # cannot relabel a plain Pi worker (or vice versa). The marker is part of + # the verified command, so it lives here, not in any caller. + pi|pi-signed) + if [ "$kind" = secondmate ]; then + printf '%s%s' "FM_PI_HARNESS=$harness $harness" ' __MODELFLAG____EFFORTFLAG__-e __PITURNEND__ -e __PIWATCH__ "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + else + printf '%s%s' "FM_PI_HARNESS=$harness $harness" ' __MODELFLAG____EFFORTFLAG__-e __PIEXT__ "$(__OPINPUT__ encode launch-brief < __BRIEF__)"' + fi + ;; + # grok (Grok Build TUI): a positional prompt starts the supervised interactive + # session. --always-approve auto-approves every tool execution (verified: the + # crewmate runs fully autonomously, no permission gate), which an unattended + # crewmate needs; it is the targeted equivalent of claude's + # --dangerously-skip-permissions. grok's turn-end signal does NOT ride the + # launch command - it is a Stop-event hook installed by fm-spawn.sh (global hook + + # per-task pointer), so the template is identical for ship/scout/secondmate. + grok) printf '%s' 'grok --always-approve __MODELFLAG____EFFORTFLAG__"$(__OPINPUT__ encode launch-brief < __BRIEF__)"' ;; + # Kimi Code rejects a positional prompt, so it launches bare and receives + # only an absolute brief pointer after fm-spawn.sh's TUI readiness gate. + # Its turn-end signal is a globally configured Stop hook plus a guarded + # per-task worktree token, so no launch placeholder belongs here. + kimi) printf '%s' '__KIMIBIN__ __MODELFLAG__--auto' ;; + *) return 1 ;; + esac +} + +model_flag_for_harness() { + local harness=$1 model=$2 + [ -n "$model" ] && [ "$model" != default ] || return 0 + case "$harness" in + claude|codex|opencode|pi|pi-signed|grok|kimi) + printf -- '--model %s ' "$(shell_quote "$model")" + ;; + esac +} + +effort_flag_for_harness() { + local harness=$1 effort=$2 + [ -n "$effort" ] && [ "$effort" != default ] || return 0 + case "$harness" in + claude) + case "$effort" in + low|medium|high|xhigh|max) printf -- '--effort %s ' "$(shell_quote "$effort")" ;; + esac + ;; + codex) + # The installed codex config schema uses model_reasoning_effort, and the + # bundled model catalog advertises low|medium|high|xhigh. Omit max rather + # than passing an unsupported value. + case "$effort" in + low|medium|high|xhigh) printf -- '-c %s ' "$(shell_quote "model_reasoning_effort=\"$effort\"")" ;; + esac + ;; + grok) + # grok exposes both --effort and --reasoning-effort; firstmate's profile + # axis is the reasoning knob. As of grok 0.2.99, --reasoning-effort accepts + # only low|medium|high and rejects both xhigh and max, so omit those rather + # than passing a known-bad value. + case "$effort" in + low|medium|high) printf -- '--reasoning-effort %s ' "$(shell_quote "$effort")" ;; + esac + ;; + pi|pi-signed) + # Pi 0.80.6 accepts the full shared effort vocabulary, including max, through + # its --thinking flag. + case "$effort" in + low|medium|high|xhigh|max) printf -- '--thinking %s ' "$(shell_quote "$effort")" ;; + esac + ;; + # opencode's interactive `opencode --prompt` launch has a verified --model + # flag but no verified effort flag. Its `opencode run --variant` flag belongs + # to a different, non-interactive launch mode, so fm-spawn does not pass it. + # kimi likewise has no reasoning-effort flag; the requested axis stays in + # task metadata but never reaches the launch command. + esac +} diff --git a/bin/fm-launch.sh b/bin/fm-launch.sh new file mode 100755 index 00000000000..643eb40ddf1 --- /dev/null +++ b/bin/fm-launch.sh @@ -0,0 +1,773 @@ +#!/usr/bin/env bash +# fm-launch.sh - the captain's front door: pick a harness, start ONE firstmate +# primary session in this home, and attach to it. +# +# Usage: +# bin/fm-launch.sh render the menu, select, launch, attach +# bin/fm-launch.sh --print-menu render the menu and exit 0 (no side effects) +# bin/fm-launch.sh --verbose add the launch mechanics on stderr +# bin/fm-launch.sh --help +# +# This launches a PRIMARY (a firstmate session: no task, no worktree, no brief), +# never a crewmate. bin/fm-spawn.sh remains the only way to start a crewmate, +# scout, or secondmate. Both compose their command from bin/fm-launch-lib.sh, so +# there is exactly one copy of every verified launch command. +# +# THE CONSUMER OBLIGATION, DISCHARGED HERE. bin/fm-launch-lib.sh's header binds +# every consumer that composes a `primary` launch to tell the captain, at launch +# time, that the session runs WITHOUT permission prompts. This launcher +# discharges that with FM_LAUNCH_AUTONOMY_NOTICE below, printed in the menu +# header on every render - before the choice, not after it, so the captain knows +# the posture of the session while they are still choosing it. The obligation +# binds on the POSTURE, not on any particular flag, so this line must survive +# even if every template's flags change. Do not make it conditional, do not move +# it behind --verbose, and do not drop it because one adapter reaches the posture +# structurally rather than through a bypass flag. +# +# WHAT THIS SCRIPT DOES NOT DO, deliberately: +# - No network, ever, before the menu. Both availability probes are local file +# reads (`command -v`, and one read of pi's auth record), which is what holds +# the sub-150ms first-paint target. No catalog fetch, no quota probe, no +# `pi --list-models` (a subprocess plus JSON parse, ~1s when measured). +# - No `herdr status` before the menu. The Herdr gate is needed to LAUNCH, not +# to CHOOSE, so it runs after selection and keeps a socket round trip off the +# critical path. +# - No side effects before selection: `q` and Ctrl-C are always clean. +# - No bin/fm-guard.sh call. That guard reports supervision mechanics to a +# RUNNING firstmate; this runs before one exists, and its warnings are +# exactly the mechanics the front door keeps off the captain's screen. +# +# Files, all local to the effective FM_HOME: +# config/launch-presets.json optional menu presets; absent means the built-in +# five below. Schema: docs/configuration.md. +# state/.launch-last the last-used entry id, written atomically +# (temp + mv) so a killed launcher can never +# corrupt the Enter default. +# state/.launch.lock the home-scoped launch lock (bin/fm-wake-lib.sh's +# primitives, the same ones bin/fm-spawn.sh uses +# for its per-task spawn lock), held from the +# reattach check through the launch-line send. See +# the launch-lock note above launch_lock_release(). +# +# Test seams (documented so the suite does not reach into internals): +# FM_LAUNCH_PRESETS override the presets path +# FM_LAUNCH_PI_AUTH override pi's auth record path +# FM_LAUNCH_NO_ATTACH=1 do everything except the final attach +# FM_LAUNCH_READY_ATTEMPTS bounded shell-prompt poll attempts (default 10) +# FM_LAUNCH_READY_SLEEP seconds between those attempts (default 0.3) +set -eu + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" + +PRESETS="${FM_LAUNCH_PRESETS:-$CONFIG/launch-presets.json}" +PI_AUTH="${FM_LAUNCH_PI_AUTH:-$HOME/.pi/agent/auth.json}" +LAST_USED="$STATE/.launch-last" + +# The herdr TAB label for this home's one primary session. The WORKSPACE label +# is already per-home (fm_backend_herdr_workspace_label: "firstmate", or +# "2ndmate-"), so this constant is unambiguous inside it and can never +# collide with a crewmate tab, which is always labeled fm-. +PRIMARY_TAB_LABEL="firstmate" + +# See THE CONSUMER OBLIGATION above before changing or moving this line. +FM_LAUNCH_AUTONOMY_NOTICE="Sessions start without permission prompts." + +VERBOSE=0 +PRINT_MENU=0 + +usage() { + cat <<'USAGE' +fm-launch.sh - the captain's front door: pick a harness, start ONE firstmate +primary session in this home, and attach to it. + +Usage: + bin/fm-launch.sh render the menu, select, launch, attach + bin/fm-launch.sh --print-menu render the menu and exit 0 (no side effects) + bin/fm-launch.sh --verbose add the launch mechanics on stderr + bin/fm-launch.sh --help + +Every session this starts runs WITHOUT permission prompts; the menu says so +before you choose. Menu presets and setup: docs/launcher.md. +USAGE +} + +while [ $# -gt 0 ]; do + case "$1" in + --verbose|-v) VERBOSE=1 ;; + --print-menu) PRINT_MENU=1 ;; + -h|--help) usage; exit 0 ;; + *) echo "fm-launch.sh: unknown option '$1' (try --help)" >&2; exit 2 ;; + esac + shift +done + +vlog() { [ "$VERBOSE" -eq 1 ] && printf 'fm-launch: %s\n' "$*" >&2 || true; } + +# --- presentation ----------------------------------------------------------- +# +# Every refusal is at most a few lines: what happened, and the one thing that +# fixes it. Never a stack trace, never a wall of diagnostics - --verbose exists +# for when the captain wants one. + +TTY=0 +[ -t 0 ] && [ -t 1 ] && TTY=1 +DIM=""; RESET="" +if [ "$TTY" -eq 1 ] && [ -z "${NO_COLOR:-}" ]; then + DIM=$'\033[2m'; RESET=$'\033[0m' +fi + +# refuse ...: print an indented refusal and exit non-zero. +refuse() { + local line + printf '\n' + for line in "$@"; do printf ' %s\n' "$line"; done + exit 1 +} + +# refuse_at_door ...: a refusal reached before the menu was ever drawn, so +# it carries the header the menu would have carried. +refuse_at_door() { + printf '\n Firstmate\n' + refuse "$@" +} + +# --- presets ---------------------------------------------------------------- +# +# An entry is a record of five fields - id, label, harness, model, effort - +# joined by the ASCII unit separator. NOT by a tab: a tab is IFS whitespace, so +# splitting on one silently COLLAPSES empty fields, and an entry missing its id +# would arrive here looking like a well-formed entry with everything shifted one +# place left. The unit separator is not IFS whitespace, so an empty field stays +# an empty field and the validation below can see it. An entry carries NO +# command. The command comes from launch_template() at launch time, +# which is the whole point: a hand-written launch string has already drifted once +# in a downstream registry (see bin/fm-launch-lib.sh's header). +# +# "default" for model or effort means "let the adapter choose", exactly as +# model_flag_for_harness/effort_flag_for_harness already define it. +# +# The built-in five ship with model="default" on the NATIVE entries on purpose: +# this repo is a shared template, and a pinned model id rots for every home whose +# account cannot reach it. The two Pi-routed entries DO pin a provider-qualified +# model because that string is what selects the provider at all - without it the +# entry is not "Grok via pi", it is just pi. +FM_LAUNCH_US=$'\037' +FM_LAUNCH_BUILTIN_PRESETS=( + $'claude\037Claude\037claude\037default\037high' + $'chatgpt-sol\037ChatGPT Sol\037pi\037openai-codex/gpt-5.6-sol\037high' + $'grok\037Grok\037pi\037xai/grok-4\037high' + $'codex\037Codex\037codex\037default\037high' + $'opencode\037OpenCode\037opencode\037default\037default' +) + +ENTRIES=() + +load_presets() { + local line + if [ ! -f "$PRESETS" ]; then + ENTRIES=("${FM_LAUNCH_BUILTIN_PRESETS[@]}") + return 0 + fi + command -v jq >/dev/null 2>&1 \ + || refuse_at_door "Cannot read the menu presets at $PRESETS." \ + "Install jq, or remove that file to use the built-in menu." + local parsed + parsed=$(jq -r ' + if (.entries | type) != "array" then error("entries must be an array") else . end + | .entries[] + | [(.id // ""), (.label // ""), (.harness // ""), (.model // "default"), (.effort // "default")] + | map(tostring) + | join("\u001f")' "$PRESETS" 2>/dev/null) \ + || refuse_at_door "The menu presets at $PRESETS are not valid." \ + "Fix the file, or remove it to use the built-in menu." + while IFS= read -r line; do + [ -n "$line" ] || continue + # jq renders a missing key as an empty field, so an entry short of its three + # required keys arrives here looking merely blank. Refuse it rather than + # rendering an unnamed row that cannot launch. + entry_split "$line" + [ -n "$E_ID" ] && [ -n "$E_LABEL" ] && [ -n "$E_HARNESS" ] \ + || refuse_at_door "An entry in $PRESETS is missing its id, label, or harness." \ + "Give every entry all three, or remove the file to use the built-in menu." + ENTRIES+=("$line") + done <: unpack one record into E_ID/E_LABEL/E_HARNESS/E_MODEL/ +# E_EFFORT. +# +# Every hot-path helper in this file communicates through globals rather than +# stdout, and none of them may be called through $(...). That is a speed +# contract, not a style preference: first paint is budgeted under 150 ms, a fork +# costs 2-4 ms, and the menu touches these helpers roughly seven times per entry. +# Routing them through command substitution measured 152 ms against a 2 ms bash +# floor - the forks WERE the entire budget. +entry_split() { # + local rec=$1 IFS=$FM_LAUNCH_US noglob=0 + # Splitting is not the whole story: each split word would still undergo + # pathname expansion, so a field carrying * ? or [ could match files in the + # launcher's cwd and shift every later field. Glob is off for the split and + # restored to whatever it was - via set -f/+f, never a subshell, per the + # no-fork contract above. + case $- in *f*) noglob=1 ;; esac + set -f + # shellcheck disable=SC2086 # deliberate word split on the unit separator via IFS + set -- $rec + [ "$noglob" -eq 1 ] || set +f + E_ID=${1:-}; E_LABEL=${2:-}; E_HARNESS=${3:-}; E_MODEL=${4:-}; E_EFFORT=${5:-} +} + +# --- availability probe ----------------------------------------------------- +# +# Probed, never declared. An entry is available only when it can actually start +# today, and an entry that cannot stays VISIBLE and dim rather than disappearing, +# so the menu never changes shape under the captain's muscle memory. +# +# Two local reads, no network: +# native - the harness binary resolves on PATH. This launcher is the front +# door, so its PATH is the captain's login PATH. +# pi-routed - `pi` resolves AND the entry's provider (the part of the model +# before "/") appears as a key in pi's auth record. +# +# pi's auth record is a flat JSON object keyed by provider id, whose values are +# OAuth objects with fixed field names (type, access, refresh, expires, +# accountId). Matching "" followed by a colon therefore identifies a +# top-level provider key and cannot collide with a nested field name. Doing it +# with a shell match rather than a jq subprocess is what keeps first paint fast; +# an absent or unreadable record simply means "no providers", which renders the +# entry unavailable with its sign-in line - the honest, fail-closed direction. +# harness_on_path : memoized `command -v`, because a PATH MISS is the +# single most expensive thing the menu does and presets may name one harness +# more than once (the two Pi-routed entries both need `pi`). +# +# Measured on a WSL home whose PATH carries 29 Windows mounts among 40 entries: +# a HIT costs under 2 ms because the lookup stops at the first match, while a +# MISS scans every remaining directory and costs about 55 ms - the built-in +# menu's two uninstalled native entries alone account for ~113 ms of a ~125 ms +# render. Dropping the Windows mounts from PATH and changing nothing else takes +# the same render to ~7 ms, so the sub-150 ms first-paint target is held by +# construction and what remains is 9p directory access, not this file's work. +# Nothing in this file can shorten that without lying about what is installed, +# so the ENFORCED ceiling in tests/fm-launch.test.sh is 500 ms. +HARNESS_PATH_MEMO="" + +harness_on_path() { # + local h=$1 hit + case "$HARNESS_PATH_MEMO" in + *":$h=1:"*) return 0 ;; + *":$h=0:"*) return 1 ;; + esac + if command -v "$h" >/dev/null 2>&1; then hit=1; else hit=0; fi + HARNESS_PATH_MEMO="$HARNESS_PATH_MEMO:$h=$hit:" + [ "$hit" -eq 1 ] +} + +PI_AUTH_BLOB="" +PI_AUTH_READ=0 + +pi_auth_load() { + [ "$PI_AUTH_READ" -eq 0 ] || return 0 + PI_AUTH_READ=1 + PI_AUTH_BLOB="" + # $( + local provider=$1 re + pi_auth_load + [ -n "$PI_AUTH_BLOB" ] || return 1 + # Held in a variable so the provider name interpolates while the rest stays + # regex: bash 3.2 treats a quoted pattern operand as a literal string. + re="\"$provider\"[[:space:]]*:" + [[ $PI_AUTH_BLOB =~ $re ]] +} + +# pi_provider_of : sets E_PROVIDER to the provider a pi model string +# selects, or empty when the model carries none (pi then uses its own configured +# default provider). +E_PROVIDER="" +pi_provider_of() { # + local model=$1 + E_PROVIDER="" + case "$model" in + default|'') return 0 ;; + */*) E_PROVIDER=${model%%/*} ;; + esac +} + +# probe_entry : sets P_STATUS (ok|unavailable) and P_NOTE - the ONE +# actionable line for an unavailable entry, or the route hint ("via pi") for an +# available Pi-routed one. Requires entry_split to have run for this record. +# +# The two refusal causes launch_template() distinguishes are BOTH reported here, +# with the wording each one owes the captain (bin/fm-launch-lib.sh header): +# a harness with no verified adapter at all is a preset error, while a verified +# harness that simply has no primary shape is fixed by choosing another entry - +# never by reaching for a raw launch command. Deriving cause 2 as "works as a +# crewmate, refuses as a primary" keeps this free of a hardcoded harness name. +P_STATUS="" +P_NOTE="" +probe_entry() { + local route="" + P_STATUS=unavailable + + if ! launch_template "$E_HARNESS" primary >/dev/null 2>&1; then + if launch_template "$E_HARNESS" ship >/dev/null 2>&1; then + P_NOTE="no primary session - choose another entry" + else + P_NOTE="unknown runtime $E_HARNESS - fix the menu presets" + fi + return 0 + fi + + # Both members of the pi family are Pi-routed: pi-signed shares pi's auth + # record, so it gets the same provider probe and route hint. It does NOT + # share pi's executable - pi-signed is an explicitly selected identity that + # never falls back to pi (bin/fm-spawn.sh), so the PATH probe below still + # resolves the entry's own harness name. + E_PROVIDER="" + case "$E_HARNESS" in + pi|pi-signed) route="via pi"; pi_provider_of "$E_MODEL" ;; + esac + + if ! harness_on_path "$E_HARNESS"; then + P_NOTE="not installed - install $E_HARNESS" + return 0 + fi + + if [ -n "$E_PROVIDER" ] && ! pi_provider_authed "$E_PROVIDER"; then + P_NOTE="$route - sign in: pi /login" + return 0 + fi + + P_STATUS=ok + P_NOTE=$route +} + +# --- menu render ------------------------------------------------------------ + +# entry_detail: sets E_DETAIL, the model/effort column, from the current +# entry_split fields. Renders only what is actually pinned - the literal word +# "default" is noise, not information - and shows a plain dash when the adapter +# decides everything. +E_DETAIL="" +entry_detail() { + E_DETAIL="" + [ "$E_MODEL" = default ] || E_DETAIL=${E_MODEL##*/} + if [ "$E_EFFORT" != default ]; then + [ -n "$E_DETAIL" ] && E_DETAIL="$E_DETAIL · $E_EFFORT" || E_DETAIL=$E_EFFORT + fi + [ -n "$E_DETAIL" ] || E_DETAIL="-" +} + +# pad : set to left-aligned to COLUMNS. +# printf's %-Ns pads by BYTES, which misaligns every row whose detail carries a +# multi-byte separator, so the width is computed from bash's character count. +pad() { # + local fill="" short=$(( $3 - ${#2} )) + [ "$short" -gt 0 ] && printf -v fill '%*s' "$short" '' + printf -v "$1" '%s%s' "$2" "$fill" +} + +STATUSES=() +NOTES=() + +probe_all() { + local rec + STATUSES=(); NOTES=() + for rec in "${ENTRIES[@]}"; do + entry_split "$rec" + probe_entry + STATUSES+=("$P_STATUS") + NOTES+=("$P_NOTE") + done +} + +# default_index: the entry Enter takes. The last-used entry when it is still +# available, otherwise the first available entry, otherwise none (0). A stale or +# now-unavailable last-used entry must never become an Enter that cannot launch. +DEFAULT_INDEX=0 +DEFAULT_MARK="← default" + +resolve_default() { + local last="" i + DEFAULT_INDEX=0 + DEFAULT_MARK="← default" + if [ -r "$LAST_USED" ]; then + IFS=$' \t\n' read -r last < "$LAST_USED" || last="" + fi + if [ -n "$last" ]; then + for i in "${!ENTRIES[@]}"; do + entry_split "${ENTRIES[$i]}" + if [ "$E_ID" = "$last" ] && [ "${STATUSES[$i]}" = ok ]; then + DEFAULT_INDEX=$((i + 1)) + DEFAULT_MARK="← last" + return 0 + fi + done + fi + for i in "${!ENTRIES[@]}"; do + if [ "${STATUSES[$i]}" = ok ]; then + DEFAULT_INDEX=$((i + 1)) + return 0 + fi + done +} + +render_menu() { + local i n note row padded_label padded_detail default_label="" + printf '\n Firstmate\n' + printf ' %s\n\n' "$FM_LAUNCH_AUTONOMY_NOTICE" + for i in "${!ENTRIES[@]}"; do + n=$((i + 1)) + entry_split "${ENTRIES[$i]}" + entry_detail + [ "$n" -eq "$DEFAULT_INDEX" ] && default_label=$E_LABEL + note=${NOTES[$i]} + [ "$n" -eq "$DEFAULT_INDEX" ] && note=${note:+"$note "}$DEFAULT_MARK + pad padded_label "$E_LABEL" 18 + pad padded_detail "$E_DETAIL" 22 + row=" $n $padded_label $padded_detail $note" + # Trailing spaces are noise in a captured menu; strip them. + row=${row%"${row##*[![:space:]]}"} + if [ "${STATUSES[$i]}" = ok ]; then + printf '%s\n' "$row" + else + printf '%s%s%s\n' "$DIM" "$row" "$RESET" + fi + done + if [ "$DEFAULT_INDEX" -gt 0 ]; then + printf '\n ⏎ %s 1-%d select q quit\n' "$default_label" "${#ENTRIES[@]}" + else + printf '\n nothing is available yet 1-%d select q quit\n' "${#ENTRIES[@]}" + fi +} + +# --- selection -------------------------------------------------------------- +# +# One keypress selects AND launches; Enter takes the default; q quits clean. +# +# An invalid key redraws the prompt with an inline "?" and waits. That is a +# deliberate divergence from firstmate's refuse-don't-reprompt law, and only for +# a TTY: that law is right for a SCRIPTED selection, where a wrong value must +# fail loudly rather than be guessed, and wrong for a human who mistyped a key at +# their own front door. Determinism is preserved either way - the input alphabet +# is closed and the mapping is total; only the response to invalid input differs. +# Non-TTY stdin keeps the original behavior exactly: one line, no reprompt. + +SELECTED=0 + +prompt_line() { # + printf '\r\033[K› %s' "$1" +} + +select_entry_tty() { + local key n status + while :; do + prompt_line "" + IFS= read -rsn1 key || { printf '\n'; exit 1; } + case "$key" in + q|Q) printf '\r\033[K'; exit 0 ;; + '') n=$DEFAULT_INDEX ;; + [0-9]) n=$key ;; + *) prompt_line "?"; continue ;; + esac + if [ "$n" -lt 1 ] || [ "$n" -gt "${#ENTRIES[@]}" ]; then + prompt_line "?" + continue + fi + status=${STATUSES[$((n - 1))]} + if [ "$status" != ok ]; then + prompt_line "$n ${NOTES[$((n - 1))]}" + continue + fi + printf '\r\033[K› %s\n' "$n" + SELECTED=$n + return 0 + done +} + +select_entry_pipe() { + local line n + IFS=$' \t\n' read -r line || line="" + case "$line" in + q|Q) exit 0 ;; + # EOF and a blank line REFUSE here rather than taking the default. Enter + # taking the default is a convenience for a human at a keyboard; for a + # scripted caller the same silence would mean `fm-launch.sh < /dev/null` + # starts a real unattended session nobody chose. A scripted selection must + # be explicit - that is the refuse-don't-reprompt law this path preserves. + '') refuse "No selection was made." "Choose 1-${#ENTRIES[@]}, or q to quit." ;; + # A 0-prefixed number must refuse HERE: it would pass the decimal range + # check below (`test` parses 08 as eight) and then reach $((n - 1)), where + # bash parses the leading zero as octal and dies under set -eu instead of + # refusing with guidance like every other invalid input. + 0[0-9]) refuse "'$line' is not one of the menu choices." "Choose 1-${#ENTRIES[@]}, or q to quit." ;; + [0-9]|[0-9][0-9]) n=$line ;; + *) refuse "'$line' is not one of the menu choices." "Choose 1-${#ENTRIES[@]}, or q to quit." ;; + esac + if [ "$n" -lt 1 ] || [ "$n" -gt "${#ENTRIES[@]}" ]; then + refuse "There is no menu entry $n." "Choose 1-${#ENTRIES[@]}, or q to quit." + fi + if [ "${STATUSES[$((n - 1))]}" != ok ]; then + entry_split "${ENTRIES[$((n - 1))]}" + refuse "$E_LABEL is not available: ${NOTES[$((n - 1))]}." + fi + printf '› %s\n' "$n" + SELECTED=$n +} + +# --- launch ----------------------------------------------------------------- + +# remember_choice : atomic temp + mv, so a launcher killed mid-write leaves +# the previous default intact rather than a truncated file. +remember_choice() { # + local id=$1 tmp + mkdir -p "$STATE" 2>/dev/null || return 0 + tmp="$LAST_USED.tmp.$$" + printf '%s\n' "$id" > "$tmp" 2>/dev/null || return 0 + mv -f "$tmp" "$LAST_USED" 2>/dev/null || rm -f "$tmp" 2>/dev/null || true +} + +# The launch lock closes the reattach guard's startup window. Between this +# launcher's tab creation and the harness TUI registering as a herdr agent, the +# new primary's pane still reads no-agent, so an unlocked second launch would +# pass live_primary_pane and then have fm_backend_herdr_create_task replace the +# just-started tab as a husk, silently killing it. The lock - not any pane +# state - is what distinguishes "no agent YET, a launch is in flight" from "no +# agent because that session died": a dead session's launcher no longer holds +# it, so crash-husk replacement on the next launch stays intact. It is held +# from before the reattach check through the launch-line send and released on +# every exit path by the EXIT trap; attach() releases it explicitly because +# exec never reaches that trap. A killed launcher cannot wedge the home: +# fm_lock_try_acquire reclaims a lock whose recorded pid is dead. +LAUNCH_LOCK="" +LAUNCH_LOCK_HELD=0 + +launch_lock_release() { + if [ "$LAUNCH_LOCK_HELD" = 1 ]; then + LAUNCH_LOCK_HELD=0 + fm_lock_release "$LAUNCH_LOCK" || true + fi +} +trap launch_lock_release EXIT + +# herdr_gate: Herdr is MANDATORY and there is no silent fallback to a bare +# shell. Absence is this launcher's own condition and gets the front door's +# wording; every other verdict belongs to fm_backend_herdr_version_check, whose +# message carries the protocol numbers and is relayed rather than restated. +herdr_gate() { + local out + command -v herdr >/dev/null 2>&1 \ + || refuse "Herdr is required and was not found." \ + "Install it (https://herdr.dev), then relaunch." + if ! out=$(fm_backend_herdr_version_check 2>&1); then + refuse "Herdr is required and this one cannot be used." "${out#error: }" + fi +} + +# live_primary_pane: the pane of a primary already running in this home, or +# empty. Read-only, and skipped entirely when no herdr server is up - if nothing +# is running, nothing can be reattached to. Reattach is checked BEFORE any create +# so a second launch can never leave two primaries contending for this home's +# session lock. Returns non-zero when herdr stops answering mid-check; the +# caller must then FAIL CLOSED - proceeding would create a second primary in +# exactly the race this guard exists to prevent. +live_primary_pane() { # + local session=$1 running wsid tab pane + running=$(fm_backend_herdr_cli "$session" status --json 2>/dev/null | jq -r '.server.running // false' 2>/dev/null) + [ "$running" = true ] || return 0 + wsid=$(fm_backend_herdr_workspace_find "$session") + [ -n "$wsid" ] || return 0 + tab=$(fm_backend_herdr_cli "$session" tab list --workspace "$wsid" 2>/dev/null \ + | jq -r --arg want "$PRIMARY_TAB_LABEL" '.result.tabs[]? | select(.label == $want) | .tab_id' 2>/dev/null | head -1) + [ -n "$tab" ] || return 0 + pane=$(fm_backend_herdr_pane_for_tab "$session" "$wsid" "$tab") || return 1 + [ -n "$pane" ] || return 0 + [ "$(fm_backend_herdr_pane_agent_state "$session" "$pane")" = live ] || return 0 + printf '%s' "$pane" +} + +# attach : `herdr session attach ` is the verified attach form. +# Empirical, herdr 0.7.1: `herdr session --help` lists an `attach` subcommand +# ("Attach to a session"), and `herdr session attach --help` reports exactly +# `Usage: herdr session attach ` with a single required NAME argument. +# The exec replaces this process, so the EXIT trap can never run past it; the +# launch lock is released explicitly first. +attach() { # + local session=$1 + if [ -n "${FM_LAUNCH_NO_ATTACH:-}" ]; then + vlog "attach suppressed by FM_LAUNCH_NO_ATTACH" + return 0 + fi + launch_lock_release + exec herdr session attach "$session" +} + +# wait_for_prompt : a bounded deadline poll for the pane's shell prompt, +# never a bare sleep. It gates on the SHELL, not on the agent, so the attach +# lands well before the model's first token. Exhaustion refuses with the attempt +# count and creates nothing further; it never proceeds hopefully. +wait_for_prompt() { # + local target=$1 attempts=${FM_LAUNCH_READY_ATTEMPTS:-10} nap=${FM_LAUNCH_READY_SLEEP:-0.3} i out + for i in $(seq 1 "$attempts"); do + out=$(fm_backend_herdr_capture "$target" 5 2>/dev/null) || out="" + if [ -n "$(printf '%s' "$out" | tr -d '[:space:]')" ]; then + vlog "shell prompt ready after attempt $i" + return 0 + fi + sleep "$nap" + done + return 1 +} + +launch_entry() { # + local rec=$1 session container seeded ids tab pane + local template modelflag effortflag launch target + entry_split "$rec" + entry_detail + + printf ' %s · %s\n' "$E_LABEL" "$E_DETAIL" + + fm_backend_source herdr || refuse "The Herdr runtime could not be loaded." "Reinstall firstmate, then relaunch." + # The lock primitives. Sourced here, after selection, because fm-wake-lib.sh + # creates $STATE at source time and nothing may touch the home before a choice. + # shellcheck source=bin/fm-wake-lib.sh + . "$SCRIPT_DIR/fm-wake-lib.sh" + + herdr_gate + session=$(fm_backend_herdr_session) + + LAUNCH_LOCK="$STATE/.launch.lock" + fm_lock_try_acquire "$LAUNCH_LOCK" \ + || refuse "Another launch is already starting a session here." \ + "Wait for it to finish, then reattach or retry." + LAUNCH_LOCK_HELD=1 + + pane=$(live_primary_pane "$session") \ + || refuse "Herdr could not confirm whether a session is already running here." \ + "Retry, or run with --verbose." + if [ -n "$pane" ]; then + if [ "$TTY" -eq 1 ]; then + printf '\n A firstmate session is already running here.\n' + printf ' r reattach q quit ' + local key + IFS= read -rsn1 key || key=q + printf '\n' + case "$key" in + r|R|'') attach "$session"; return 0 ;; + *) exit 0 ;; + esac + fi + refuse "A firstmate session is already running here." \ + "Attach to it, or close it before starting another." + fi + + template=$(launch_template "$E_HARNESS" primary) \ + || refuse "$E_LABEL cannot start a firstmate session." "Choose another entry." + modelflag=$(model_flag_for_harness "$E_HARNESS" "$E_MODEL") + effortflag=$(effort_flag_for_harness "$E_HARNESS" "$E_EFFORT") + launch=${template//__MODELFLAG__/$modelflag} + launch=${launch//__EFFORTFLAG__/$effortflag} + # An unset flag placeholder leaves one trailing space (bin/fm-launch-lib.sh + # header); cosmetic in a shell command, trimmed here. + launch=${launch%"${launch##*[![:space:]]}"} + # The pane's shell is a child of the herdr SERVER, which this launcher may + # itself have started - so it can inherit this process's overrides. Clear them + # and pin FM_HOME explicitly, the same shape bin/fm-spawn.sh uses to launch a + # firstmate primary in a secondmate home. + launch="FM_ROOT_OVERRIDE= FM_STATE_OVERRIDE= FM_DATA_OVERRIDE= FM_PROJECTS_OVERRIDE= FM_CONFIG_OVERRIDE= FM_HOME=$(shell_quote "$FM_HOME") $launch" + vlog "launch: $launch" + + fm_backend_herdr_server_ensure "$session" \ + || refuse "The Herdr server did not start." "Retry, or run with --verbose." + container=$(fm_backend_herdr_container_ensure "$FM_HOME") \ + || refuse "The Herdr workspace for this home could not be prepared." "Retry, or run with --verbose." + seeded=${container#*$'\t'} + container=${container%%$'\t'*} + ids=$(fm_backend_herdr_create_task "$container" "$PRIMARY_TAB_LABEL" "$FM_HOME" "$seeded") \ + || refuse "The firstmate session could not be created." "Retry, or run with --verbose." + tab=${ids%% *} + pane=${ids##* } + target="$session:$pane" + vlog "workspace=$container tab=$tab pane=$pane" + + wait_for_prompt "$target" \ + || refuse "The session did not start within ${FM_LAUNCH_READY_ATTEMPTS:-10} attempts." \ + "Retry, or run with --verbose." + + # One line, one round trip: `clear` wipes both the echoed command itself and + # any shell banner in the fresh pane, and the agent starts in its place - so + # the first thing on screen after the attach repaint is firstmate's greeting. + fm_backend_herdr_send_literal "$target" "clear && $launch" \ + || refuse "The session could not be started in its pane." "Retry, or run with --verbose." + fm_backend_herdr_send_key "$target" Enter \ + || refuse "The session could not be started in its pane." "Retry, or run with --verbose." + + remember_choice "$E_ID" + fm_backend_herdr_cli "$session" tab focus "$tab" >/dev/null 2>&1 || true + attach "$session" +} + +# --- main ------------------------------------------------------------------- + +[ -d "$FM_HOME" ] \ + || refuse_at_door "No firstmate home at $FM_HOME." "Set FM_HOME, or finish setup, then relaunch." + +# Sourcing fm-backend.sh here is a local file read - it touches no herdr +# socket, so the no-network, no-`herdr status` promises above hold. It is +# needed at the door because the launcher only ever starts a Herdr session: +# a home explicitly configured for another backend would get a primary that +# fm-send/fm-watch/fm-spawn cannot see and that the reattach guard cannot +# find, so the incoherence must refuse before the menu, not after a choice. +# Only the EXPLICIT setting (FM_BACKEND, then config/backend) is checked; +# auto-detection reflects the terminal this command runs in, not the home's +# configuration, and the session this launcher creates runs inside Herdr. +# shellcheck source=bin/fm-backend.sh +. "$SCRIPT_DIR/fm-backend.sh" + +BACKEND_CONFIGURED=$(fm_backend_configured_name) +if [ -n "$BACKEND_CONFIGURED" ] && [ "$BACKEND_CONFIGURED" != herdr ]; then + if [ -n "${FM_BACKEND:-}" ]; then + refuse_at_door "FM_BACKEND selects the $BACKEND_CONFIGURED backend, and this launcher only starts Herdr sessions." \ + "Unset FM_BACKEND (or set it to herdr), then relaunch." + fi + refuse_at_door "This home's config/backend selects $BACKEND_CONFIGURED, and this launcher only starts Herdr sessions." \ + "Set config/backend to herdr (or remove it), then relaunch." +fi + +# shellcheck source=bin/fm-launch-lib.sh +. "$SCRIPT_DIR/fm-launch-lib.sh" + +load_presets +probe_all +resolve_default +render_menu + +[ "$PRINT_MENU" -eq 1 ] && exit 0 + +if [ "$TTY" -eq 1 ]; then + select_entry_tty +else + select_entry_pipe +fi + +launch_entry "${ENTRIES[$((SELECTED - 1))]}" diff --git a/bin/fm-loopspec.sh b/bin/fm-loopspec.sh new file mode 100755 index 00000000000..4290438cb4e --- /dev/null +++ b/bin/fm-loopspec.sh @@ -0,0 +1,683 @@ +#!/usr/bin/env bash +# fm-loopspec.sh - the single owner of firstmate's canonical LoopSpec +# representation: validation, deterministic selection, and persistent loop state. +# +# A LoopSpec answers one question and only that question: under what conditions +# is this done again, what may execute, how is progress verified, and when must +# it stop. It is not a ticket, a plan, a skill, an ExecutionUnit, a state machine +# or a decomposition node. A node that produces three children has decomposed; +# it has not looped. +# +# WHAT THIS SCRIPT IS NOT. It is not a loop runner. It never schedules, never +# polls, never spawns, never steers, never merges and never executes a skill. +# The intended path reuses what already exists: +# +# existing detector -> durable wake -> deterministic LoopSpec selection -> +# authorised skills -> verification -> persisted state -> terminal state or +# bounded next iteration +# +# This script owns exactly three of those arrows: selection, persisted state, +# and the terminal-state bookkeeping that bounds the next iteration. The wake +# still comes from the existing watcher. The work still happens in a model turn. +# +# EVERYTHING REFUSES BY DEFAULT. Every subcommand that could lead to work prints +# one stable refusal token on stderr and exits 1 rather than proceeding under +# uncertainty. An unavailable verifier can never become a pass. +# +# Usage: +# fm-loopspec.sh validate [...] validate the registry, or these files +# fm-loopspec.sh list id, version and status per spec +# fm-loopspec.sh show print one spec +# fm-loopspec.sh triggers [--summary] the sixteen-trigger register +# fm-loopspec.sh select --trigger [--scope ] +# choose exactly one eligible spec +# fm-loopspec.sh claim --event-key --spec-version [--headroom ] +# open or resume one iteration +# fm-loopspec.sh finish --event-key --terminal +# --verifier-result pass|fail|unavailable [--evidence ]... +# [--progress made|none] +# drive the iteration to a terminal +# fm-loopspec.sh state print persisted loop state +# fm-loopspec.sh --help print this usage +# +# Exit status: 0 proceed, 1 refused (one refusal token on stderr), 2 usage error. +# +# Refusal tokens are stable contract, because callers and tests key on them: +# refuse_invalid_spec the registry does not validate, so nothing may run +# refuse_unknown_spec no spec with that id +# refuse_no_match no spec matches the event +# refuse_not_enabled matched, but draft/specified/ready_not_active/disabled/retired +# refuse_trigger_unimplemented the trigger has no verified execution path +# refuse_ambiguous_tie candidates remain tied after the deterministic key +# refuse_authority_insufficient the spec's authority class needs the captain first +# refuse_version_changed the spec changed under a running iteration +# refuse_evidence_missing fewer evidence items than the verifier requires +# refuse_verifier_unavailable no verifier verdict, so no success terminal +# refuse_verification_mismatch the verifier rejected a success terminal +# refuse_capacity_unknown capacity is bounded but was not supplied +# refuse_capacity_stop capacity is inside the stop band +# refuse_budget_exceeded an iteration budget is spent +# refuse_no_progress consecutive iterations produced no progress +# refuse_duplicate_event this event key was already handled +# refuse_iteration_open another event key already holds the open iteration +# refuse_no_open_iteration nothing to finish +# refuse_unknown_terminal the terminal is not declared by this spec +# refuse_state_unreadable persistent state exists but cannot be read truthfully +# refuse_state_unwritable persistent state cannot be written truthfully +# +# The schema, the trigger register and each spec are data, owned by loopspecs/. +# This script is their only interpreter; it never restates their content. +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +SPEC_DIR="${FM_LOOPSPEC_DIR:-$FM_ROOT/loopspecs}" +SKILL_DIRS=("$FM_ROOT/.agents/skills" "$FM_ROOT/skills") + +usage() { + sed -n '2,/^set -u$/p' "$SCRIPT_DIR/fm-loopspec.sh" | sed 's/^# \{0,1\}//; $d' +} + +die_usage() { + printf 'error: %s\n' "$1" >&2 + exit 2 +} + +refuse() { + # One stable token, then a human-readable reason on the same line. Callers that + # invoke a refusing helper inside a command substitution must propagate the + # status with `|| exit 1`, because an exit inside a subshell ends only that + # subshell. + printf '%s %s\n' "$1" "$2" >&2 + exit 1 +} + +need_jq() { + command -v jq >/dev/null 2>&1 || die_usage "jq is required" +} + +schema_path() { printf '%s/schema.json' "$SPEC_DIR"; } +triggers_path() { printf '%s/triggers.json' "$SPEC_DIR"; } + +# Every *.json under the registry except the two contract files. +spec_files() { + local f base + for f in "$SPEC_DIR"/*.json; do + [ -f "$f" ] || continue + base=$(basename "$f") + case "$base" in + schema.json|triggers.json) continue ;; + esac + printf '%s\n' "$f" + done +} + +installed_skills_json() { + local d entry + local -a names=() + for d in "${SKILL_DIRS[@]}"; do + [ -d "$d" ] || continue + for entry in "$d"/*/; do + [ -d "$entry" ] || continue + names+=("$(basename "$entry")") + done + done + if [ "${#names[@]}" -eq 0 ]; then + printf '[]' + return 0 + fi + printf '%s\n' "${names[@]}" | jq -R -s 'split("\n") | map(select(length > 0))' +} + +# Structural and invariant validation of one spec against schema.json. +# Prints one line per problem, nothing when sound, and returns non-zero if the +# validator itself could not run - an unrunnable validator is never a pass. +validate_structure() { + local spec_file=$1 file_id=$2 skills_json=$3 out rc + out=$(jq -r --slurpfile schema "$(schema_path)" \ + --slurpfile triggers "$(triggers_path)" \ + --arg file_id "$file_id" \ + --argjson skills "$skills_json" ' + # Identifiers are kebab-case; terminal-state names are snake_case tokens. + # Keeping them distinct stops one vocabulary drifting into the other. + def is_slug($v): (($v|type) == "string") and ($v|test("^[a-z0-9][a-z0-9-]*$")); + def is_token($v): (($v|type) == "string") and ($v|test("^[a-z0-9][a-z0-9_]*$")); + def typeok($v; $t; $enums): + if $t == "string" then (($v|type) == "string") and (($v|length) > 0) + elif $t == "slug" then is_slug($v) + elif $t == "token" then is_token($v) + elif $t == "token[]" then (($v|type) == "array") and (all($v[]?; is_token(.))) + elif $t == "int" then (($v|type) == "number") and (($v|floor) == $v) + elif $t == "nonneg_int" then (($v|type) == "number") and (($v|floor) == $v) and ($v >= 0) + elif $t == "pos_int" then (($v|type) == "number") and (($v|floor) == $v) and ($v >= 1) + elif $t == "bool" then (($v|type) == "boolean") + elif $t == "object" then (($v|type) == "object") + elif $t == "string[]" then (($v|type) == "array") and (all($v[]?; (type == "string") and (length > 0))) + elif $t == "slug[]" then (($v|type) == "array") and (all($v[]?; is_slug(.))) + elif $t == "object[]" then (($v|type) == "array") and (($v|length) > 0) and (all($v[]?; type == "object")) + elif ($t|startswith("enum:")) then ((($enums[$t[5:]] // []) | index($v)) != null) + else false + end; + # Only ever yields objects, so a mistyped branch can never crash the walk. + def objects_at($spec; $p): + if $p == "" then [$spec] + elif ($p|endswith("[]")) then + ((try ($spec | getpath($p[0:-2] | split("."))) catch null)) as $arr + | if ($arr|type) == "array" then [$arr[] | select(type == "object")] else [] end + else + ((try ($spec | getpath($p | split("."))) catch null)) as $v + | if ($v|type) == "object" then [$v] else [] end + end; + def fieldname($obj; $field): "\($obj)\(if $obj == "" then "" else "." end)\($field)"; + + . as $spec + | $schema[0] as $S + | $triggers[0] as $T + | ($S.enums) as $enums + | ((try ($spec.trigger.id) catch null)) as $trigger_id + | ((try ($spec.no_progress.terminal) catch null)) as $np_terminal + | (((try ($spec.escalation.on) catch null)) // []) as $esc_on + | (((try ($spec.authority.permitted_skills) catch null)) // []) as $need_skills + | ((try ($spec.authority.class) catch null)) as $authority_class + | (((try ($spec.terminal_states | map(select(type == "object") | .name)) catch null)) // []) as $declared + | [ + ( $S.objects | to_entries[] as $obj + | objects_at($spec; $obj.key)[] as $val + | (($obj.value.required // {}) | keys) as $req + | (($obj.value.optional // {}) | keys) as $opt + | ($val | keys) as $have + | ( + ( ($have - ($req + $opt))[] | "unknown field: \(fieldname($obj.key; .))" ), + ( ($req - $have)[] | "missing required field: \(fieldname($obj.key; .))" ), + ( (($obj.value.required // {}) + ($obj.value.optional // {})) | to_entries[] as $e + | select(($val | has($e.key)) and ((typeok($val[$e.key]; $e.value; $enums)) | not)) + | "wrong type for \(fieldname($obj.key; $e.key)): expected \($e.value)" ) + ) + ), + + ( select($spec.loopspec_schema_version != $S.loopspec_schema_version) + | "schema version \($spec.loopspec_schema_version) does not match the contract version \($S.loopspec_schema_version)" ), + ( select($spec.id != $file_id) + | "id \"\($spec.id)\" does not match its filename \"\($file_id)\"" ), + ( select(($T.triggers | map(.id) | index($trigger_id)) == null) + | "trigger \"\($trigger_id)\" is not in the trigger register" ), + + ( (($S.required_terminal_states - $declared)[]) | "missing required terminal state: \(.)" ), + ( select(($np_terminal != null) and (($declared | index($np_terminal)) == null)) + | "no_progress.terminal \"\($np_terminal)\" is not a declared terminal state" ), + ( (($esc_on - $declared)[]) | "escalation.on names an undeclared terminal state: \(.)" ), + ( ($declared | group_by(.) | map(select(length > 1) | .[0]))[] | "duplicate terminal state: \(.)" ), + + ( select($spec.status == "enabled") + | ( + ( ($T.triggers[] | select(.id == $trigger_id)) as $t + | ( select((($t.execution_path_implemented) // false) | not) + | "status is enabled but trigger \"\($t.id)\" has no implemented execution path" ), + ( select((($t.verified) // false) | not) + | "status is enabled but trigger \"\($t.id)\" is not verified" ), + ( select((($t.enabled) // false) | not) + | "status is enabled but trigger \"\($t.id)\" is not enabled" ) + ), + ( $need_skills[] as $sk | select(($skills | index($sk)) == null) + | "status is enabled but permitted skill \"\($sk)\" is not installed" ), + ( select($authority_class == "captain-required") + | "status is enabled but authority class captain-required cannot self-enable" ) + ) + ) + ] | .[] + ' "$spec_file") + rc=$? + if [ "$rc" -ne 0 ]; then + printf 'the validator could not run against this file\n' + return 1 + fi + [ -z "$out" ] || printf '%s\n' "$out" + return 0 +} + +# Registry-wide validation. Returns 0 only when every spec is sound. +validate_registry() { + local -a files=() + local f base file_id problems skills_json rc=0 count=0 dupes + + if [ "$#" -gt 0 ]; then + files=("$@") + else + while IFS= read -r f; do + [ -n "$f" ] || continue + files+=("$f") + done < <(spec_files) + fi + + [ -f "$(schema_path)" ] || { printf 'refuse_invalid_spec registry schema is missing: %s\n' "$(schema_path)" >&2; return 1; } + [ -f "$(triggers_path)" ] || { printf 'refuse_invalid_spec trigger register is missing: %s\n' "$(triggers_path)" >&2; return 1; } + jq -e . "$(schema_path)" >/dev/null 2>&1 || { printf 'refuse_invalid_spec registry schema is not readable JSON\n' >&2; return 1; } + jq -e . "$(triggers_path)" >/dev/null 2>&1 || { printf 'refuse_invalid_spec trigger register is not readable JSON\n' >&2; return 1; } + + skills_json=$(installed_skills_json) + + for f in "${files[@]}"; do + [ -f "$f" ] || { printf 'refuse_invalid_spec %s: not a file\n' "$f" >&2; rc=1; continue; } + base=$(basename "$f") + file_id=${base%.json} + if ! jq -e . "$f" >/dev/null 2>&1; then + printf 'refuse_invalid_spec %s: not readable JSON\n' "$base" >&2 + rc=1 + continue + fi + problems=$(validate_structure "$f" "$file_id" "$skills_json") || rc=1 + if [ -n "$problems" ]; then + while IFS= read -r line; do + [ -n "$line" ] || continue + printf 'refuse_invalid_spec %s: %s\n' "$base" "$line" >&2 + done <<<"$problems" + rc=1 + continue + fi + count=$((count + 1)) + done + + # Ambiguity is a registry property, so it is checked across specs, not inside one. + if [ "${#files[@]}" -gt 1 ]; then + dupes=$(jq -sr ' + map(select(type == "object")) + | map({k: "\(.trigger.id)/\(.selection.scope)/\(.selection.priority)", id: .id}) + | group_by(.k) | map(select(length > 1)) + | .[] | "\(.[0].k) claimed by \(map(.id) | join(", "))" + ' "${files[@]}" 2>/dev/null) + if [ -n "$dupes" ]; then + while IFS= read -r line; do + [ -n "$line" ] || continue + printf 'refuse_invalid_spec ambiguous selection key: %s\n' "$line" >&2 + done <<<"$dupes" + rc=1 + fi + fi + + [ "$rc" -eq 0 ] && printf 'LOOPSPEC_VALIDATE ok specs=%s\n' "$count" + return "$rc" +} + +spec_path_for() { + printf '%s/%s.json' "$SPEC_DIR" "$1" +} + +state_path_for() { + local id=$1 template=$2 rel + rel=${template///$id} + case "$rel" in + state/*) printf '%s/%s' "$STATE" "${rel#state/}" ;; + *) printf '%s/%s' "$STATE" "$(basename "$rel")" ;; + esac +} + +read_state() { + # Prints the state object, or an empty object when no state exists yet. + # Refuses when a state file exists but cannot be read truthfully. Callers use + # `state=$(read_state ...) || exit 1` so the refusal is not lost in a subshell. + local path=$1 + if [ ! -e "$path" ]; then + printf '{}' + return 0 + fi + [ -f "$path" ] && [ ! -L "$path" ] && [ -r "$path" ] \ + || refuse refuse_state_unreadable "persistent state is not a readable regular file: $path" + jq -e . "$path" >/dev/null 2>&1 \ + || refuse refuse_state_unreadable "persistent state is not readable JSON: $path" + cat "$path" +} + +write_state() { + # Atomic replace, so a crash mid-write can never leave a half-truth on disk. + local path=$1 payload=$2 dir tmp + dir=$(dirname "$path") + mkdir -p "$dir" 2>/dev/null || refuse refuse_state_unwritable "cannot create state directory: $dir" + [ -w "$dir" ] || refuse refuse_state_unwritable "state directory is not writable: $dir" + tmp=$(mktemp "$dir/.loopspec.XXXXXX" 2>/dev/null) \ + || refuse refuse_state_unwritable "cannot create a temporary state file in: $dir" + printf '%s\n' "$payload" >"$tmp" 2>/dev/null \ + || { rm -f "$tmp"; refuse refuse_state_unwritable "cannot write state: $path"; } + mv -f "$tmp" "$path" 2>/dev/null \ + || { rm -f "$tmp"; refuse refuse_state_unwritable "cannot replace state: $path"; } +} + +# Eligibility to START an iteration, enforced wherever an iteration can begin. +# Selection is the ordinary route in, but claiming must gate independently: +# otherwise naming a spec directly would bypass the status and trigger checks and +# an inert spec could still run. Finishing deliberately does NOT gate on this, so +# an iteration already in flight can always be closed out truthfully even after +# its spec is disabled. +assert_runnable() { + local spec=$1 id status trigger class + id=$(printf '%s' "$spec" | jq -r '.id') + status=$(printf '%s' "$spec" | jq -r '.status') + trigger=$(printf '%s' "$spec" | jq -r '.trigger.id') + class=$(printf '%s' "$spec" | jq -r '.authority.class') + + [ "$status" = "enabled" ] \ + || refuse refuse_not_enabled "spec \"$id\" is $status, so it may not start an iteration" + jq -e --arg t "$trigger" ' + (.triggers[] | select(.id == $t)) + | (.execution_path_implemented and .verified and .enabled) + ' "$(triggers_path)" >/dev/null 2>&1 \ + || refuse refuse_trigger_unimplemented "trigger \"$trigger\" has no verified, enabled execution path" + [ "$class" != "captain-required" ] \ + || refuse refuse_authority_insufficient "spec \"$id\" needs the captain's word before an iteration may start" +} + +# Load a spec by id after proving the whole registry validates. A registry that +# does not validate is not a registry anything may be selected from. +load_spec_or_refuse() { + local id=$1 path + path=$(spec_path_for "$id") + [ -f "$path" ] || refuse refuse_unknown_spec "no spec with id \"$id\" in $SPEC_DIR" + validate_registry >/dev/null || exit 1 + cat "$path" +} + +cmd_validate() { + need_jq + validate_registry "$@" +} + +cmd_list() { + need_jq + local f + while IFS= read -r f; do + [ -n "$f" ] || continue + jq -r '"\(.id)\tversion=\(.spec_version)\tstatus=\(.status)\ttrigger=\(.trigger.id)\tscope=\(.selection.scope)\tpriority=\(.selection.priority)"' "$f" 2>/dev/null \ + || printf '%s\tUNREADABLE\n' "$(basename "$f")" + done < <(spec_files) +} + +cmd_show() { + need_jq + [ "$#" -eq 1 ] || die_usage "show requires exactly one spec id" + local path + path=$(spec_path_for "$1") + [ -f "$path" ] || refuse refuse_unknown_spec "no spec with id \"$1\" in $SPEC_DIR" + cat "$path" +} + +cmd_triggers() { + need_jq + local summary=0 + while [ "$#" -gt 0 ]; do + case "$1" in + --summary) summary=1; shift ;; + *) die_usage "unknown option for triggers: $1" ;; + esac + done + [ -f "$(triggers_path)" ] || refuse refuse_invalid_spec "trigger register is missing" + + if [ "$summary" -eq 1 ]; then + # Every count is computed here and stored nowhere, so no figure can rot. + jq -r ' + .triggers as $t + | "LOOPSPEC_TRIGGERS total=\($t|length)", + " specified=\([$t[]|select(.specified)]|length)", + " detector_implemented=\([$t[]|select(.detector_implemented)]|length)", + " deterministically_selectable=\([$t[]|select(.deterministically_selectable)]|length)", + " execution_path_implemented=\([$t[]|select(.execution_path_implemented)]|length)", + " verified=\([$t[]|select(.verified)]|length)", + " enabled=\([$t[]|select(.enabled)]|length)", + " ruled_never_deterministic=\([$t[]|select(.ruled_never_deterministic)]|length)", + " outside_the_sixteen=\(.outside_the_sixteen|length) (not counted in total)" + ' "$(triggers_path)" + return 0 + fi + + jq -r ' + .triggers[] + | "\(.id)\tdetector=\(.detector_status)\tdeterministic=\(.deterministically_selectable)\texec_path=\(.execution_path_implemented)\tverified=\(.verified)\tenabled=\(.enabled)" + ' "$(triggers_path)" +} + +cmd_select() { + need_jq + local trigger="" scope="" + while [ "$#" -gt 0 ]; do + case "$1" in + --trigger) [ "$#" -gt 1 ] || die_usage "--trigger requires a value"; trigger=$2; shift 2 ;; + --scope) [ "$#" -gt 1 ] || die_usage "--scope requires a value"; scope=$2; shift 2 ;; + *) die_usage "unknown option for select: $1" ;; + esac + done + [ -n "$trigger" ] || die_usage "select requires --trigger " + + validate_registry >/dev/null || exit 1 + + local -a files=() + local f + while IFS= read -r f; do + [ -n "$f" ] || continue + files+=("$f") + done < <(spec_files) + [ "${#files[@]}" -gt 0 ] || refuse refuse_no_match "the registry contains no specs" + + # Refuse an unregistered trigger before matching, so a typo can never read as + # "there is no work to do". + jq -e --arg t "$trigger" '(.triggers | map(.id) | index($t)) != null' "$(triggers_path)" >/dev/null 2>&1 \ + || refuse refuse_no_match "trigger \"$trigger\" is not in the register" + + local matched enabled winners lowest count + matched=$(jq -sr --arg t "$trigger" --arg s "$scope" ' + map(select(.trigger.id == $t and ($s == "" or .selection.scope == $s))) + | .[] | "\(.id)\t\(.spec_version)\t\(.status)\t\(.selection.priority)" + ' "${files[@]}") + [ -n "$matched" ] || refuse refuse_no_match "no spec declares trigger \"$trigger\"${scope:+ in scope \"$scope\"}" + + enabled=$(printf '%s\n' "$matched" | awk -F'\t' '$3 == "enabled"') + if [ -z "$enabled" ]; then + refuse refuse_not_enabled "matched specs for \"$trigger\", none enabled: $(printf '%s\n' "$matched" | awk -F'\t' '{printf "%s(%s) ", $1, $3}')" + fi + + # Defence in depth: validation already refuses an enabled spec on an + # unimplemented trigger, so reaching here means the register changed under us. + jq -e --arg t "$trigger" ' + (.triggers[] | select(.id == $t)) + | (.execution_path_implemented and .verified and .enabled) + ' "$(triggers_path)" >/dev/null 2>&1 \ + || refuse refuse_trigger_unimplemented "trigger \"$trigger\" has no verified, enabled execution path" + + # The deterministic selection key is (trigger, scope, priority); lowest wins. + # A remaining tie is ambiguity, and ambiguity refuses rather than picking. + lowest=$(printf '%s\n' "$enabled" | awk -F'\t' '{print $4}' | sort -n | head -1) + winners=$(printf '%s\n' "$enabled" | awk -F'\t' -v p="$lowest" '$4 == p') + count=$(printf '%s\n' "$winners" | wc -l | tr -d ' ') + if [ "$count" -ne 1 ]; then + refuse refuse_ambiguous_tie "priority $lowest is claimed by: $(printf '%s\n' "$winners" | awk -F'\t' '{printf "%s ", $1}')" + fi + + printf 'LOOPSPEC_SELECT %s version=%s priority=%s\n' \ + "$(printf '%s' "$winners" | cut -f1)" \ + "$(printf '%s' "$winners" | cut -f2)" \ + "$lowest" +} + +cmd_claim() { + need_jq + [ "$#" -ge 1 ] || die_usage "claim requires a spec id" + local id=$1; shift + local event_key="" want_version="" headroom="" + while [ "$#" -gt 0 ]; do + case "$1" in + --event-key) [ "$#" -gt 1 ] || die_usage "--event-key requires a value"; event_key=$2; shift 2 ;; + --spec-version) [ "$#" -gt 1 ] || die_usage "--spec-version requires a value"; want_version=$2; shift 2 ;; + --headroom) [ "$#" -gt 1 ] || die_usage "--headroom requires a value"; headroom=$2; shift 2 ;; + *) die_usage "unknown option for claim: $1" ;; + esac + done + [ -n "$event_key" ] || die_usage "claim requires --event-key " + [ -n "$want_version" ] || die_usage "claim requires --spec-version " + + local spec disk_version path state state_version + spec=$(load_spec_or_refuse "$id") || exit 1 + assert_runnable "$spec" + disk_version=$(printf '%s' "$spec" | jq -r '.spec_version') + [ "$disk_version" = "$want_version" ] \ + || refuse refuse_version_changed "spec \"$id\" is version $disk_version on disk, the iteration requested version $want_version" + + path=$(state_path_for "$id" "$(printf '%s' "$spec" | jq -r '.state.path_template')") + state=$(read_state "$path") || exit 1 + + # State that names a different version than the spec on disk means the spec + # moved under a live iteration. Refuse; never silently re-base it. + state_version=$(printf '%s' "$state" | jq -r '.spec_version // empty') + if [ -n "$state_version" ] && [ "$state_version" != "$disk_version" ]; then + refuse refuse_version_changed "persistent state holds version $state_version but the spec on disk is version $disk_version" + fi + + local open open_key iteration no_progress last_key + open=$(printf '%s' "$state" | jq -r '.open // false') + open_key=$(printf '%s' "$state" | jq -r '.event_key // empty') + iteration=$(printf '%s' "$state" | jq -r '.iteration // 0') + no_progress=$(printf '%s' "$state" | jq -r '.consecutive_no_progress // 0') + last_key=$(printf '%s' "$state" | jq -r '.last_event_key // empty') + + # Crash recovery: the same key still holding an open iteration resumes it, at + # the same version and the same iteration number. Nothing is re-counted. + if [ "$open" = "true" ] && [ "$open_key" = "$event_key" ]; then + printf 'LOOPSPEC_CLAIM %s version=%s iteration=%s resumed=true\n' "$id" "$disk_version" "$iteration" + return 0 + fi + [ "$open" != "true" ] \ + || refuse refuse_iteration_open "iteration $iteration is open for event key \"$open_key\"" + + # A repeated wake for an already-finished key does no work and mutates nothing. + [ "$last_key" != "$event_key" ] \ + || refuse refuse_duplicate_event "event key \"$event_key\" was already handled at iteration $iteration" + + # No-progress terminates the loop. The terminal is reached and recorded, not + # merely reported, so the loop cannot be restarted into the same dead end. + local np_max np_terminal + np_max=$(printf '%s' "$spec" | jq -r '.no_progress.max_iterations_without_progress') + np_terminal=$(printf '%s' "$spec" | jq -r '.no_progress.terminal') + if [ "$no_progress" -ge "$np_max" ]; then + write_state "$path" "$(printf '%s' "$state" | jq --arg t "$np_terminal" \ + '.open = false | .event_key = null | .last_terminal = $t | .terminal_kind = "failure"')" + refuse refuse_no_progress "$no_progress consecutive iterations without progress reached terminal \"$np_terminal\"" + fi + + local max_iterations + max_iterations=$(printf '%s' "$spec" | jq -r '.budgets.max_iterations') + [ "$((iteration + 1))" -le "$max_iterations" ] \ + || refuse refuse_budget_exceeded "iteration budget $max_iterations is spent" + + # Capacity bounds STARTING an iteration only. Verifying work already in flight + # is never blocked here, and unknown capacity is never treated as headroom. + local stop_band + stop_band=$(printf '%s' "$spec" | jq -r '.budgets.capacity_stop_band_percent') + if [ "$stop_band" -gt 0 ]; then + [ -n "$headroom" ] \ + || refuse refuse_capacity_unknown "spec bounds capacity at ${stop_band}% but --headroom was not supplied" + case "$headroom" in + ''|*[!0-9]*) die_usage "--headroom must be a whole percentage" ;; + esac + [ "$headroom" -ge "$stop_band" ] \ + || refuse refuse_capacity_stop "headroom ${headroom}% is inside the ${stop_band}% stop band" + fi + + iteration=$((iteration + 1)) + write_state "$path" "$(printf '%s' "$state" | jq \ + --arg id "$id" --arg key "$event_key" --arg status "$(printf '%s' "$spec" | jq -r .status)" \ + --argjson version "$disk_version" --argjson iter "$iteration" \ + '.spec_id = $id | .spec_version = $version | .status = $status + | .iteration = $iter | .event_key = $key | .open = true + | .consecutive_no_progress = (.consecutive_no_progress // 0)')" + printf 'LOOPSPEC_CLAIM %s version=%s iteration=%s resumed=false\n' "$id" "$disk_version" "$iteration" +} + +cmd_finish() { + need_jq + [ "$#" -ge 1 ] || die_usage "finish requires a spec id" + local id=$1; shift + local event_key="" terminal="" verifier="" progress="made" + local -a evidence=() + while [ "$#" -gt 0 ]; do + case "$1" in + --event-key) [ "$#" -gt 1 ] || die_usage "--event-key requires a value"; event_key=$2; shift 2 ;; + --terminal) [ "$#" -gt 1 ] || die_usage "--terminal requires a value"; terminal=$2; shift 2 ;; + --verifier-result) [ "$#" -gt 1 ] || die_usage "--verifier-result requires a value"; verifier=$2; shift 2 ;; + --evidence) [ "$#" -gt 1 ] || die_usage "--evidence requires a value"; evidence+=("$2"); shift 2 ;; + --progress) [ "$#" -gt 1 ] || die_usage "--progress requires a value"; progress=$2; shift 2 ;; + *) die_usage "unknown option for finish: $1" ;; + esac + done + [ -n "$event_key" ] || die_usage "finish requires --event-key " + [ -n "$terminal" ] || die_usage "finish requires --terminal " + [ -n "$verifier" ] || die_usage "finish requires --verifier-result pass|fail|unavailable" + case "$verifier" in pass|fail|unavailable) ;; *) die_usage "--verifier-result must be pass, fail or unavailable" ;; esac + case "$progress" in made|none) ;; *) die_usage "--progress must be made or none" ;; esac + + local spec path state kind open_key + spec=$(load_spec_or_refuse "$id") || exit 1 + path=$(state_path_for "$id" "$(printf '%s' "$spec" | jq -r '.state.path_template')") + state=$(read_state "$path") || exit 1 + + [ "$(printf '%s' "$state" | jq -r '.open // false')" = "true" ] \ + || refuse refuse_no_open_iteration "no open iteration for spec \"$id\"" + open_key=$(printf '%s' "$state" | jq -r '.event_key // empty') + [ "$open_key" = "$event_key" ] \ + || refuse refuse_no_open_iteration "the open iteration belongs to event key \"$open_key\"" + + kind=$(printf '%s' "$spec" | jq -r --arg t "$terminal" '.terminal_states[] | select(.name == $t) | .kind') + [ -n "$kind" ] || refuse refuse_unknown_terminal "\"$terminal\" is not a terminal state declared by spec \"$id\"" + + # An unavailable verifier can never become a pass. Neither can a rejecting one. + if [ "$kind" = "success" ]; then + [ "$verifier" != "unavailable" ] \ + || refuse refuse_verifier_unavailable "terminal \"$terminal\" is a success state and the verifier was unavailable" + [ "$verifier" = "pass" ] \ + || refuse refuse_verification_mismatch "terminal \"$terminal\" is a success state and the verifier returned \"$verifier\"" + local required have + required=$(printf '%s' "$spec" | jq -r '.verification.required_evidence | length') + have=${#evidence[@]} + [ "$have" -ge "$required" ] \ + || refuse refuse_evidence_missing "terminal \"$terminal\" requires $required evidence items, $have supplied" + fi + + local next_no_progress + if [ "$progress" = "made" ]; then + next_no_progress=0 + else + next_no_progress=$(( $(printf '%s' "$state" | jq -r '.consecutive_no_progress // 0') + 1 )) + fi + + write_state "$path" "$(printf '%s' "$state" | jq \ + --arg t "$terminal" --arg k "$kind" --arg key "$event_key" --arg v "$verifier" \ + --argjson np "$next_no_progress" --argjson ev "${#evidence[@]}" \ + '.open = false | .last_terminal = $t | .terminal_kind = $k + | .last_event_key = $key | .event_key = null + | .verifier_result = $v | .evidence_count = $ev + | .consecutive_no_progress = $np')" + printf 'LOOPSPEC_FINISH %s terminal=%s kind=%s verifier=%s\n' "$id" "$terminal" "$kind" "$verifier" +} + +cmd_state() { + need_jq + [ "$#" -eq 1 ] || die_usage "state requires exactly one spec id" + local id=$1 spec path state + spec=$(load_spec_or_refuse "$id") || exit 1 + path=$(state_path_for "$id" "$(printf '%s' "$spec" | jq -r '.state.path_template')") + state=$(read_state "$path") || exit 1 + printf '%s\n' "$state" +} + +[ "$#" -ge 1 ] || { usage >&2; exit 2; } +COMMAND=$1 +shift +case "$COMMAND" in + validate) cmd_validate "$@" ;; + list) cmd_list "$@" ;; + show) cmd_show "$@" ;; + triggers) cmd_triggers "$@" ;; + select) cmd_select "$@" ;; + claim) cmd_claim "$@" ;; + finish) cmd_finish "$@" ;; + state) cmd_state "$@" ;; + -h|--help|help) usage ;; + *) die_usage "unknown command: $COMMAND" ;; +esac diff --git a/bin/fm-merge-local.sh b/bin/fm-merge-local.sh index fdc8011488b..381f463689c 100755 --- a/bin/fm-merge-local.sh +++ b/bin/fm-merge-local.sh @@ -9,6 +9,23 @@ # auto-approves), and only as a clean fast-forward - it refuses a diverged branch # and tells you to have the crewmate rebase. See AGENTS.md prime directives, # project management, and task lifecycle. +# +# Uncommitted work in the project blocks the merge only on a GENUINE COLLISION: +# a path that is both uncommitted in the project checkout and rewritten by this +# exact fast-forward, including an untracked entry and an incoming path that +# cannot both exist because one is a directory where the other is a file. An +# untracked directory git will not descend into, a nested repository above all, +# stands for everything beneath it, because nothing here can see what it holds. An +# untracked or modified path the fast-forward never touches cannot be clobbered +# by it, and an ignored path is declared disposable - git itself overwrites one +# without complaint, and this guard matches git rather than second-guessing it - +# so neither blocks the merge. A refusal names every colliding path and what the +# incoming commits do to it. A state this guard cannot classify confidently - an +# unresolved conflict, a merge, cherry-pick, revert, rebase, bisect, or patch +# application left in progress, an unrecognized status code, a git command that +# fails - refuses instead of proceeding, naming the operation it found and how to +# conclude or abandon it, ahead of the divergence that operation's own commits +# caused, since rebasing is not what settles it. # Usage: fm-merge-local.sh set -eu @@ -46,22 +63,318 @@ git -C "$PROJ" rev-parse --verify --quiet "refs/heads/$BRANCH" >/dev/null || { e DEFAULT=$(default_branch) || { echo "error: cannot determine default branch for $PROJ; expected origin/HEAD, main, or master" >&2; exit 1; } -# The project's main checkout must be on its default branch and clean, so the -# fast-forward lands predictably (firstmate never writes here otherwise). -cur=$(git -C "$PROJ" symbolic-ref --short HEAD 2>/dev/null || echo "") -[ "$cur" = "$DEFAULT" ] || { echo "error: $PROJ is on '$cur', expected default branch '$DEFAULT'; cannot merge safely" >&2; exit 1; } -if [ -n "$(git -C "$PROJ" status --porcelain 2>/dev/null | head -1)" ]; then - echo "error: $PROJ has a dirty working tree; refusing to merge into it" >&2 +# --- unclassifiable-state refusals ------------------------------------------- +# +# An unclassifiable state is not a safe state, so every parse or command failure +# lands here rather than degrading into an empty - and therefore silent - change set. +refuse_unreadable() { # + echo "REFUSED: cannot classify the state of $PROJ, so refusing to merge into it." >&2 + echo " $1" >&2 + echo "Resolve that state in $PROJ (finish or abort the operation in progress, or report an unexpected git status), then retry." >&2 exit 1 +} + +# A merge, cherry-pick, revert, rebase, bisect, or patch application left +# half-finished is such a state, and one no other check can see on its own: once +# its conflicts are staged, git status reports plain modifications +# indistinguishable from ordinary dirt. Resolve each sentinel through +# rev-parse --git-path rather than assuming "$PROJ/.git": in a linked worktree +# the git dir is elsewhere, and a hardcoded path would silently never fire. +IN_PROGRESS_PATH= +in_progress_path() { # + local rel + rel=$(git -C "$PROJ" rev-parse --git-path "$1") \ + || refuse_unreadable "git rev-parse --git-path $1 failed in $PROJ" + case "$rel" in + /*) IN_PROGRESS_PATH=$rel ;; + *) IN_PROGRESS_PATH="$PROJ/$rel" ;; + esac +} + +refuse_if_in_progress() { # + in_progress_path "$1" + [ -e "$IN_PROGRESS_PATH" ] || return 0 + refuse_unreadable "$2" +} + +# `git am` and the apply-backend rebase share the rebase-apply directory but need +# different advice - `git rebase --abort` refuses outright while an am is open. +# git's own status tells them apart by the `applying` marker inside it, so this +# does too, and checks that marker before the directory it lives in. +refuse_if_operation_in_progress() { + refuse_if_in_progress MERGE_HEAD \ + "a merge is in progress here; conclude it with 'git commit' or abandon it with 'git merge --abort'" + refuse_if_in_progress CHERRY_PICK_HEAD \ + "a cherry-pick is in progress here; conclude it with 'git cherry-pick --continue' or abandon it with 'git cherry-pick --abort'" + refuse_if_in_progress REVERT_HEAD \ + "a revert is in progress here; conclude it with 'git revert --continue' or abandon it with 'git revert --abort'" + # A multi-commit cherry-pick or revert keeps its remaining picks in the + # sequencer after the conflicted one is resolved and committed, which drops + # CHERRY_PICK_HEAD while the sequence is still open. That marker is checked + # last of the three so a sequence still carrying its live conflict keeps the + # more specific message above. The interactive rebase does not appear here: it + # keeps its todo list in rebase-merge, which the next sentinel covers. + refuse_if_in_progress sequencer/todo \ + "a cherry-pick or revert sequence is in progress here; conclude it with 'git cherry-pick --continue' or 'git revert --continue', or abandon it with 'git cherry-pick --abort' or 'git revert --abort'" + refuse_if_in_progress rebase-merge \ + "a rebase is in progress here; conclude it with 'git rebase --continue' or abandon it with 'git rebase --abort'" + refuse_if_in_progress rebase-apply/applying \ + "a patch application is in progress here; conclude it with 'git am --continue' or abandon it with 'git am --abort'" + refuse_if_in_progress rebase-apply \ + "a rebase is in progress here; conclude it with 'git rebase --continue' or abandon it with 'git rebase --abort'" + refuse_if_in_progress BISECT_LOG \ + "a bisect is in progress here; end it with 'git bisect reset'" +} + +# --- state checks, ordered specific before generic --------------------------- +# +# The ordering principle every check below obeys, and every check added later +# must obey: a refusal that can NAME the state it found runs before any generic +# check that would describe the same checkout in vaguer terms. A half-finished +# operation is the reason this matters - it parks the checkout on no branch at +# all, and it commits onto the default branch as it works through its remaining +# picks. Either symptom reaches a generic check first if the order is reversed, +# and the operator is told they are on the wrong branch or that the task branch +# has diverged and should be rebased, when what they actually have to do is +# conclude or abandon the operation, which restores the default branch and makes +# the fast-forward valid again. + +cur=$(git -C "$PROJ" symbolic-ref --short HEAD 2>/dev/null || echo "") +if [ -z "$cur" ]; then + # A rebase and a bisect both park the checkout on a detached HEAD, where the + # operation's own name is all there is to report: the listing below classifies + # what such a rebase left behind as ordinary dirt or as a bare conflict, and the + # branch check has nothing but an empty branch name to print. + refuse_if_operation_in_progress fi -# Clean fast-forward only: DEFAULT must be an ancestor of BRANCH. +# --- working-tree listing ---------------------------------------------------- +# +# A blanket "any uncommitted change refuses" check is stricter than git itself +# and self-deadlocking: it blocks the very commit that would settle the +# uncommitted entries (adding ignore rules, dropping a machine-local pointer from +# the index), because the cure sits behind the symptom. So refuse only where the +# fast-forward could actually destroy local work. +# +# Both sides are read NUL-delimited. -z is not an optimization here: it is the +# only status and diff format that emits paths verbatim instead of quoting names +# with spaces or non-ASCII bytes. --untracked-files=all is required so untracked +# entries are individual files rather than a collapsed parent directory, which +# would hide a collision at a nested path. It cannot collapse every such parent: +# a directory git will not descend into - a nested repository is the shape that +# does it - still arrives as a lone "/" entry, so the collision lookup below +# normalizes that trailing slash away and treats the entry as covering everything +# beneath it. Ignored files are absent from this listing (no --ignored), which is +# why they never collide. + +GUARD_TMP= +guard_tmp_cleanup() { + [ -n "$GUARD_TMP" ] || return 0 + rm -rf "$GUARD_TMP" + GUARD_TMP= +} +trap guard_tmp_cleanup EXIT +GUARD_TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-merge-local.XXXXXX") || { + echo "error: cannot create a temporary directory to inspect $PROJ" >&2 + exit 1 +} + +DIRTY_PATHS=() # tracked paths carrying staged or unstaged changes +UNTRACKED_PATHS=() # untracked, non-ignored files +INC_PATHS=() # paths the fast-forward rewrites +INC_ACTIONS=() # what it does to INC_PATHS[i]: changes, removes, or adds + +git -C "$PROJ" status --porcelain=v1 -z --untracked-files=all > "$GUARD_TMP/status" \ + || refuse_unreadable "git status failed in $PROJ" + +# Records are "XY \0". Rename and copy entries carry a second NUL field +# holding the other endpoint, which must be consumed from the same stream or +# every later record is misread as a status code. +while IFS= read -r -d '' entry; do + code=${entry:0:2} + path=${entry:3} + case "$code" in + '??') + UNTRACKED_PATHS+=("$path") + continue + ;; + 'DD'|'AA'|U?|?U) + refuse_unreadable "unresolved merge conflict at '$path'" + ;; + esac + case "${code:0:1}" in + ' '|M|T|A|D|R|C) : ;; + *) refuse_unreadable "unrecognized git status code '$code' at '$path'" ;; + esac + # 'A' on the worktree side is an intent-to-add entry (`git add -N`), which is an + # ordinary dirty tracked path; the collision intersection below decides it, the + # same as a modification. Every code outside these two sets still refuses, + # because a state this cannot classify is not a safe state. + case "${code:1:1}" in + ' '|M|T|A|D|R|C) : ;; + *) refuse_unreadable "unrecognized git status code '$code' at '$path'" ;; + esac + DIRTY_PATHS+=("$path") + case "$code" in + *R*|*C*) + IFS= read -r -d '' other || refuse_unreadable "truncated '$code' rename entry at '$path'" + DIRTY_PATHS+=("$other") + ;; + esac +done < "$GUARD_TMP/status" + +# An operation that kept HEAD attached reaches the sentinels here rather than at +# the detached-HEAD check above, deliberately after the listing: one still +# carrying live conflicts is named by the conflicted path it left behind, and +# these catch it once its resolutions are staged and it has nothing left to show. +refuse_if_operation_in_progress + +# --- generic checks, reached only once no specific state was named ----------- + +# The project's main checkout must be on its default branch, so the fast-forward +# lands predictably (firstmate never writes here otherwise). +[ "$cur" = "$DEFAULT" ] || { echo "error: $PROJ is on '$cur', expected default branch '$DEFAULT'; cannot merge safely" >&2; exit 1; } + +# Clean fast-forward only: DEFAULT must be an ancestor of BRANCH. Settled before +# the incoming change set below, so that set is exactly the paths this +# fast-forward rewrites rather than a diverged two-way diff. if ! git -C "$PROJ" merge-base --is-ancestor "$DEFAULT" "$BRANCH"; then echo "REFUSED: $BRANCH is not a fast-forward of $DEFAULT (it has diverged)." >&2 echo "Have the crewmate rebase $BRANCH onto $DEFAULT, then retry." >&2 exit 1 fi +# --- incoming change set and collision guard --------------------------------- + +# The trailing `--` is load-bearing: without it git applies its revision/filename +# ambiguity check to both arguments and dies outright in any project holding a +# root entry named like its default branch, which would refuse a landing the +# operator cannot settle. No other git call here takes a pathspec, so none of the +# others can read a revision as a filename. +git -C "$PROJ" diff -z --name-status "$DEFAULT" "$BRANCH" -- > "$GUARD_TMP/incoming" \ + || refuse_unreadable "git diff failed between $DEFAULT and $BRANCH in $PROJ" + +# Records are "\0\0", and for rename/copy "\0\0\0" +# (source first, the reverse of the status field order above). +while IFS= read -r -d '' change; do + IFS= read -r -d '' path || refuse_unreadable "truncated '$change' entry in the incoming diff" + case "$change" in + A) + INC_PATHS+=("$path"); INC_ACTIONS+=(adds) + ;; + M|M[0-9]*|T|T[0-9]*) + INC_PATHS+=("$path"); INC_ACTIONS+=(changes) + ;; + D) + INC_PATHS+=("$path"); INC_ACTIONS+=(removes) + ;; + R|R[0-9]*) + IFS= read -r -d '' other || refuse_unreadable "truncated '$change' entry at '$path' in the incoming diff" + INC_PATHS+=("$path"); INC_ACTIONS+=(removes) + INC_PATHS+=("$other"); INC_ACTIONS+=(adds) + ;; + C|C[0-9]*) + IFS= read -r -d '' other || refuse_unreadable "truncated '$change' entry at '$path' in the incoming diff" + INC_PATHS+=("$other"); INC_ACTIONS+=(adds) + ;; + *) + refuse_unreadable "unrecognized git diff status '$change' at '$path'" + ;; + esac +done < "$GUARD_TMP/incoming" + +# Set INC_ACTION to what the fast-forward does to , or fail when it leaves +# it alone. Both lookups below assign a global rather than echoing: a command +# substitution would fork a subshell per dirty path in a large checkout, and +# would also swallow the `exit` a refusal depends on. +INC_ACTION= +incoming_action() { # + local want=$1 i=0 + while [ "$i" -lt "${#INC_PATHS[@]}" ]; do + if [ "${INC_PATHS[$i]}" = "$want" ]; then + INC_ACTION=${INC_ACTIONS[$i]} + return 0 + fi + i=$((i + 1)) + done + return 1 +} + +# Set UNTRACKED_COLLISION to how the fast-forward would have to disturb the +# untracked entry at to land, or fail when it can leave that entry alone. +# git refuses all three shapes, so name them here rather than letting git emit +# its own abort: the merge's file at that exact path, a file where the merge +# needs a directory, and a directory where the merge needs a file. +UNTRACKED_COLLISION= +untracked_collision() { # + local want=$1 i=0 add lead nested + lead="untracked file here" + nested="needs that path to be a directory to create" + # A lone "/" entry is the one shape --untracked-files=all still collapses: + # git reports a directory it will not descend into, a nested repository above + # all, without listing what is inside. Comparing that entry verbatim would build + # the prefix pattern '//*', which matches nothing, so every incoming add + # beneath it would pass this guard and hit git's own abort instead. Strip the + # slash and let the entry stand for its whole subtree: nothing here can see what + # that directory holds, and an unreadable state refuses. + case "$want" in + */) + want=${want%/} + lead="untracked directory here that git will not descend into" + nested="creates a file beneath it at" + ;; + esac + while [ "$i" -lt "${#INC_PATHS[@]}" ]; do + if [ "${INC_ACTIONS[$i]}" = adds ]; then + add=${INC_PATHS[$i]} + if [ "$add" = "$want" ]; then + UNTRACKED_COLLISION="$lead, and this merge creates a file at that path" + return 0 + fi + case "$add" in + "$want"/*) + UNTRACKED_COLLISION="$lead, and this merge $nested '$add'" + return 0 + ;; + esac + case "$want" in + "$add"/*) + UNTRACKED_COLLISION="$lead, and this merge creates a file at '$add', replacing the directory holding it" + return 0 + ;; + esac + fi + i=$((i + 1)) + done + return 1 +} + +COLLISIONS=() +if [ "${#DIRTY_PATHS[@]}" -gt 0 ]; then + for path in "${DIRTY_PATHS[@]}"; do + if incoming_action "$path"; then + COLLISIONS+=("$path - uncommitted changes here, and this merge $INC_ACTION it") + fi + done +fi +if [ "${#UNTRACKED_PATHS[@]}" -gt 0 ]; then + for path in "${UNTRACKED_PATHS[@]}"; do + if untracked_collision "$path"; then + COLLISIONS+=("$path - $UNTRACKED_COLLISION") + fi + done +fi + +if [ "${#COLLISIONS[@]}" -gt 0 ]; then + echo "REFUSED: uncommitted work in $PROJ collides with the $BRANCH fast-forward:" >&2 + for collision in "${COLLISIONS[@]}"; do + echo " $collision" >&2 + done + echo "Commit, stash, or remove exactly those paths, then retry; other uncommitted files do not block this merge." >&2 + exit 1 +fi +guard_tmp_cleanup + before=$(git -C "$PROJ" rev-parse --short "$DEFAULT") git -C "$PROJ" merge --ff-only "$BRANCH" >/dev/null after=$(git -C "$PROJ" rev-parse --short "$DEFAULT") diff --git a/bin/fm-model-registry-lib.sh b/bin/fm-model-registry-lib.sh new file mode 100644 index 00000000000..d30ded6da97 --- /dev/null +++ b/bin/fm-model-registry-lib.sh @@ -0,0 +1,617 @@ +#!/usr/bin/env bash +# fm-model-registry-lib.sh - single owner of config/models.json parsing, schema +# validation, and the zero-budget routing decision. +# +# WHY THIS EXISTS. The fleet's zero-budget rule ("the budget for every API-key +# provider is ZERO ... this is a safety rule, not a preference") lived only as +# prose inside a JSON comment blob in config/crew-dispatch.json. Nothing read it. +# One API key can reach both free and metered models on the same provider, so a +# single mistyped or well-meant model name was a charge, and the fleet had already +# demonstrated once that it will route from a plausible name without checking. +# This library is the enforced copy of that rule; the prose is descriptive. +# +# It is sourced by bin/fm-spawn.sh (the last gate before a dispatch spends money, +# and the only one that sees an explicit --model that bypassed the dispatch config) +# and by bin/fm-bootstrap.sh (which catches a bad model at config-edit time, +# before any dispatch at all). +# +# docs/configuration.md "Model registry (config/models.json)" owns the schema. +# .agents/skills/model-onboarding/SKILL.md owns the admission policy this +# enforces. This header owns the mechanics. +# +# ENFORCEMENT SCOPE - the deliberate asymmetry, ruled 2026-07-28: +# +# Registry ABSENT -> enforcement inert, and bootstrap says so out loud whenever +# the dispatch config names a provider-prefixed model. This +# keeps the change purely additive: a home with no registry +# behaves byte-identically to one built before this existed. +# Unenforced is never SILENT, but it is not a refusal. +# Registry PRESENT -> fail closed everywhere. Malformed JSON, an unknown schema +# version, an unclassified provider, a missing jq, or a +# stale-evidence allowlist entry all REFUSE. A broken safety +# file must never read as an absent one. +# +# The unit of authorization is the MODEL, never the provider and never the +# credential: one key reaching both free and paid models means "this provider is +# fine" is not a safe answer to any question. + +# Idempotent guard: fm-spawn.sh and fm-bootstrap.sh may both be in one process +# tree, and a re-source must not redefine constants under set -u. +if [ -n "${FM_MODEL_REGISTRY_LIB_SOURCED:-}" ]; then + return 0 +fi +FM_MODEL_REGISTRY_LIB_SOURCED=1 + +# The only schema this build understands. A registry carrying anything else is +# refused rather than best-effort parsed: a future format change must fail loudly +# instead of being silently misread by an old binary. +FM_MODEL_SCHEMA_VERSION='fm-model-registry.v1' + +# Evidence source kinds, in the descending authority order the policy defines. +# The freshness rule turns on this ordering: harness-fetched-cache is a cache the +# provider rewrites underneath you, so it may never be the SOLE evidence for a +# cost class or an entitlement. Anything weaker than it (third-party, inference) +# is likewise insufficient alone. +# Listed in descending authority; that ordering is the policy's, and the two +# weakest kinds are anecdote rather than evidence. +FM_MODEL_SOURCE_KINDS='probe provider-entitlement provider-doc harness-static-catalogue harness-fetched-cache third-party inference' + +# Price evidence is a NARROWER set than entitlement evidence, and conflating them +# is a real hole rather than a nicety: a successful probe proves the account gets +# an answer, and says nothing whatsoever about what that answer costs. Only a +# source that actually carries a price can establish one. This is what excludes an +# allowlist entry whose price is known solely from a cache the provider rewrites - +# the case where a repricing would land silently and the allowlist would not notice. +FM_MODEL_SOURCE_PRICE_AUTHORITATIVE='provider-doc harness-static-catalogue' + +# --------------------------------------------------------------------------- +# Location and presence +# --------------------------------------------------------------------------- + +# Echo the registry path for the active home. CONFIG is exported by the callers; +# fall back to the standard layout so the library is usable standalone in tests. +fm_model_registry_path() { + local cfg + cfg="${CONFIG:-${FM_CONFIG_OVERRIDE:-${FM_HOME:-.}/config}}" + printf '%s\n' "$cfg/models.json" +} + +fm_model_registry_present() { + [ -f "$(fm_model_registry_path)" ] +} + +# --------------------------------------------------------------------------- +# Identity +# --------------------------------------------------------------------------- + +# Echo the provider half of a "provider/model" name, or nothing for a bare name. +# A bare name (opus, haiku) is a harness-native model selector: no provider +# credential is involved, so the zero-budget rule has nothing to say about it. +fm_model_provider_of() { + case "${1:-}" in + */*) printf '%s\n' "${1%%/*}" ;; + *) printf '%s\n' '' ;; + esac +} + +# --------------------------------------------------------------------------- +# Schema validation +# --------------------------------------------------------------------------- + +# fm_model_registry_validate +# Echo one human-readable reason and return 1 when the registry is invalid; +# print nothing and return 0 when it is well formed. +# +# The promotion-authority checks encode the captain-delegated ruling of +# 2026-07-28 as a CEILING in each direction: a registry may be equally or more +# conservative than the ruling, never more permissive. The Tier 1 / Tier 0 rows +# are the hard ceiling - no accumulation of evidence may enter those tiers, because +# Tier 1 is triggered by risk rather than capability rank, and a model with a +# spotless record at Tier 2 has demonstrated nothing whatsoever about credential +# handling or destructive-operation judgment. +fm_model_registry_validate() { + local file=$1 err + [ -f "$file" ] || { echo "registry file not found: $file"; return 1; } + if ! command -v jq >/dev/null 2>&1; then + echo "jq is required to validate the model registry" + return 1 + fi + if ! jq -e . "$file" >/dev/null 2>&1; then + echo "malformed JSON" + return 1 + fi + err=$(jq -r \ + --arg schema "$FM_MODEL_SCHEMA_VERSION" \ + --arg kinds "$FM_MODEL_SOURCE_KINDS" \ + --arg pricestrong "$FM_MODEL_SOURCE_PRICE_AUTHORITATIVE" \ + ' + def kinds: ($kinds | split(" ")); + def pricestrong: ($pricestrong | split(" ")); + # The freshness rule, applied to PRICE: at least one source that actually + # carries a price. A probe is deliberately NOT enough - it proves the account + # gets an answer, not what that answer costs. A cache-only price is exactly how + # allowlist names once came to exist that no shipped catalogue could confirm. + def fresh($s): (($s // []) | any(. as $k | pricestrong | index($k))); + def bad_kinds($s): (($s // []) | map(select(. as $k | kinds | index($k) | not))); + def accessclasses: ["A","B","C","D"]; + def postures: ["subscription-flat","api-key","self-hosted"]; + def provstatuses: ["active","blocked","dropped"]; + def costclasses: ["subscription-flat","verified-free","metered","unknown"]; + def modelstatuses: ["rejected","blocked","experimental","approved-fallback","approved-specialist","approved-primary"]; + def obslevels: ["O1","O2","O3","O4"]; + def entries($o): (($o // {}) | to_entries); + + if type != "object" then "top-level value must be an object" + elif (.schema? // null) == null then "missing schema; expected \"" + $schema + "\"" + elif .schema != $schema then "unsupported schema \"" + (.schema | tostring) + "\"; this build understands \"" + $schema + "\"" + + # --- providers --------------------------------------------------------- + elif has("providers") and (.providers | type) != "object" then "providers must be an object" + elif (entries(.providers) | map(select((.value | type) != "object")) | length) > 0 then + "provider entries must be objects: " + (entries(.providers) | map(select((.value | type) != "object")) | map(.key) | join(", ")) + elif (entries(.providers) | map(select((.value.access_class? // null) as $a | ($a == null) or (accessclasses | index($a) | not))) | length) > 0 then + "provider needs access_class one of A, B, C, D: " + (entries(.providers) | map(select((.value.access_class? // null) as $a | ($a == null) or (accessclasses | index($a) | not))) | map(.key) | join(", ")) + elif (entries(.providers) | map(select((.value.cost_posture? // null) as $p | ($p == null) or (postures | index($p) | not))) | length) > 0 then + "provider needs cost_posture one of subscription-flat, api-key, self-hosted: " + (entries(.providers) | map(select((.value.cost_posture? // null) as $p | ($p == null) or (postures | index($p) | not))) | map(.key) | join(", ")) + elif (entries(.providers) | map(select(has("status") and ((.value.status) as $s | provstatuses | index($s) | not))) | length) > 0 then + "provider status must be one of active, blocked, dropped: " + (entries(.providers) | map(select(has("status") and ((.value.status) as $s | provstatuses | index($s) | not))) | map(.key) | join(", ")) + elif (entries(.providers) | map(select((.value.status? // "active") != "active" and (((.value.status_reason? // "") | length) == 0))) | length) > 0 then + "a blocked or dropped provider needs status_reason: " + (entries(.providers) | map(select((.value.status? // "active") != "active" and (((.value.status_reason? // "") | length) == 0))) | map(.key) | join(", ")) + + # --- models ------------------------------------------------------------ + elif has("models") and (.models | type) != "object" then "models must be an object" + elif (entries(.models) | map(select((.value | type) != "object")) | length) > 0 then + "model entries must be objects: " + (entries(.models) | map(select((.value | type) != "object")) | map(.key) | join(", ")) + elif (entries(.models) | map(select((.value.cost_class? // null) as $c | ($c == null) or (costclasses | index($c) | not))) | length) > 0 then + "model needs cost_class one of subscription-flat, verified-free, metered, unknown: " + (entries(.models) | map(select((.value.cost_class? // null) as $c | ($c == null) or (costclasses | index($c) | not))) | map(.key) | join(", ")) + elif (entries(.models) | map(select((.value.status? // null) as $s | ($s == null) or (modelstatuses | index($s) | not))) | length) > 0 then + "model needs a known status: " + (entries(.models) | map(select((.value.status? // null) as $s | ($s == null) or (modelstatuses | index($s) | not))) | map(.key) | join(", ")) + elif (entries(.models) | map(select((.value.status? // "") as $s | (($s == "rejected") or ($s == "blocked")) and (((.value.status_reason? // "") | length) == 0))) | length) > 0 then + "a rejected or blocked model needs status_reason so the refusal is not rediscovered: " + (entries(.models) | map(select((.value.status? // "") as $s | (($s == "rejected") or ($s == "blocked")) and (((.value.status_reason? // "") | length) == 0))) | map(.key) | join(", ")) + elif (entries(.models) | map(select(has("observation_level") and ((.value.observation_level) as $o | obslevels | index($o) | not))) | length) > 0 then + "observation_level must be one of O1, O2, O3, O4: " + (entries(.models) | map(select(has("observation_level") and ((.value.observation_level) as $o | obslevels | index($o) | not))) | map(.key) | join(", ")) + elif (entries(.models) | map(select((.value.evidence?.price?.sources? // null) != null and ((bad_kinds(.value.evidence.price.sources) | length) > 0))) | length) > 0 then + "unknown evidence source kind in: " + (entries(.models) | map(select((.value.evidence?.price?.sources? // null) != null and ((bad_kinds(.value.evidence.price.sources) | length) > 0))) | map(.key) | join(", ")) + + # --- zero_budget allowlist -------------------------------------------- + elif has("zero_budget") and (.zero_budget | type) != "object" then "zero_budget must be an object" + elif (.zero_budget?.allowlist? // null) != null and (.zero_budget.allowlist | type) != "object" then "zero_budget.allowlist must be an object" + elif (entries(.zero_budget?.allowlist) | map(select((.value | type) != "object")) | length) > 0 then + "allowlist entries must be objects: " + (entries(.zero_budget?.allowlist) | map(select((.value | type) != "object")) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select((.value.sources? | type) != "array")) | length) > 0 then + "allowlist entry needs a sources array: " + (entries(.zero_budget?.allowlist) | map(select((.value.sources? | type) != "array")) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select((bad_kinds(.value.sources) | length) > 0)) | length) > 0 then + "unknown evidence source kind in allowlist: " + (entries(.zero_budget?.allowlist) | map(select((bad_kinds(.value.sources) | length) > 0)) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select(fresh(.value.sources) | not)) | length) > 0 then + "allowlist entry has no price evidence beyond a provider-rewritten cache (a probe proves entitlement, not price): " + (entries(.zero_budget?.allowlist) | map(select(fresh(.value.sources) | not)) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select(((.value.verified_at? // "") | length) == 0)) | length) > 0 then + "allowlist entry needs verified_at: " + (entries(.zero_budget?.allowlist) | map(select(((.value.verified_at? // "") | length) == 0)) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select((.value.price_at_verification? | type) != "object")) | length) > 0 then + "allowlist entry needs price_at_verification: " + (entries(.zero_budget?.allowlist) | map(select((.value.price_at_verification? | type) != "object")) | map(.key) | join(", ")) + elif (entries(.zero_budget?.allowlist) | map(select([.value.price_at_verification[]?] | any(. != 0))) | length) > 0 then + "allowlist entry is not priced at zero: " + (entries(.zero_budget?.allowlist) | map(select([.value.price_at_verification[]?] | any(. != 0))) | map(.key) | join(", ")) + + # --- promotion authority ceiling -------------------------------------- + elif has("promotion") and (.promotion | type) != "object" then "promotion must be an object" + elif (.promotion?.authority? // null) != null and (.promotion.authority | type) != "object" then "promotion.authority must be an object" + elif ((.promotion?.authority?.t1_to_t0? // "never-by-evidence") != "never-by-evidence") then + "promotion.authority.t1_to_t0 must be never-by-evidence: Tier 0 is never entered by accumulated evidence" + elif ((.promotion?.authority?.t2_to_t1? // "never-by-evidence") != "never-by-evidence") then + "promotion.authority.t2_to_t1 must be never-by-evidence: Tier 1 is triggered by risk, not capability rank" + elif ((.promotion?.authority?.t3_to_t2? // "captain-confirm") as $a | ["captain-confirm","never-by-evidence"] | index($a) | not) then + "promotion.authority.t3_to_t2 must be captain-confirm or never-by-evidence" + elif ((.promotion?.authority?.t4_to_t3? // "automatic-notify-immediate") as $a | ["automatic-notify-immediate","captain-confirm","never-by-evidence"] | index($a) | not) then + "promotion.authority.t4_to_t3 must be automatic-notify-immediate, captain-confirm, or never-by-evidence" + elif (.promotion?.enabled? // false) == true and (((.promotion?.requires_instrument? // "") | length) == 0) then + "promotion.enabled needs requires_instrument naming the evidence instrument" + else empty + end + ' "$file" 2>/dev/null || true) + if [ -n "$err" ]; then + printf '%s\n' "$err" + return 1 + fi + return 0 +} + +# --------------------------------------------------------------------------- +# The zero-budget decision +# --------------------------------------------------------------------------- + +# fm_model_zero_budget_decision +# Return 0 to allow the dispatch; return 1 and echo one actionable reason to +# refuse it. The reason names the exact model, why it was refused, and where the +# allowlist lives - a refusal a supervisor cannot act on is worse than none. +fm_model_zero_budget_decision() { + local model=${1:-} reg provider verdict + # No model selected: the harness picks its own default on a subscription + # credential. Nothing here to authorize. + [ -n "$model" ] && [ "$model" != default ] || return 0 + + provider=$(fm_model_provider_of "$model") + # A bare model name is a harness-native selector (claude's "opus"), not a + # provider-credentialed route. + [ -n "$provider" ] || return 0 + + reg=$(fm_model_registry_path) + # Ruled 2026-07-28: an absent registry leaves enforcement inert so the change + # stays purely additive for homes that never opted in. Bootstrap is what makes + # that state loud rather than silent. + [ -f "$reg" ] || return 0 + + if ! command -v jq >/dev/null 2>&1; then + echo "jq is required to check $model against the zero-budget allowlist in $reg, and is not installed; refusing rather than dispatching an unchecked model" + return 1 + fi + + verdict=$(jq -r \ + --arg model "$model" \ + --arg provider "$provider" \ + --arg schema "$FM_MODEL_SCHEMA_VERSION" \ + --arg pricestrong "$FM_MODEL_SOURCE_PRICE_AUTHORITATIVE" \ + ' + def pricestrong: ($pricestrong | split(" ")); + def fresh($s): (($s // []) | any(. as $k | pricestrong | index($k))); + if type != "object" or (.schema? // "") != $schema then + "refuse\tregistry schema is missing or unsupported" + else + (.providers[$provider]? // null) as $p + | if $p == null then + "refuse\tprovider \"" + $provider + "\" is not classified in the registry, so its cost posture is unknown" + elif ($p.status? // "active") != "active" then + "refuse\tprovider \"" + $provider + "\" is " + ($p.status | tostring) + ": " + ($p.status_reason? // "no reason recorded") + elif ($p.cost_posture? // "") == "subscription-flat" then "allow" + elif ($p.cost_posture? // "") == "self-hosted" then "allow" + else + (.zero_budget?.allowlist[$model]? // null) as $a + | if $a == null then + "refuse\t" + $provider + " is an API-key provider and \"" + $model + "\" is not on the verified-free allowlist" + elif (fresh($a.sources) | not) then + "refuse\tallowlist entry for \"" + $model + "\" has a price known only from a provider-rewritten cache, which cannot establish one" + elif ([$a.price_at_verification[]?] | any(. != 0)) then + "refuse\tallowlist entry for \"" + $model + "\" records a non-zero price" + else "allow" + end + end + end + ' "$reg" 2>/dev/null || true) + + # A malformed registry yields no verdict. A broken safety file must not read as + # an absent one, so an unparseable answer refuses. + if [ -z "$verdict" ]; then + echo "the model registry $reg could not be read; refusing $model rather than dispatching an unchecked model" + return 1 + fi + case "$verdict" in + allow) return 0 ;; + refuse*) + printf 'zero-budget rule refuses %s: %s (allowlist: %s -> zero_budget.allowlist)\n' \ + "$model" "${verdict#refuse$'\t'}" "$reg" + return 1 + ;; + *) + printf 'zero-budget check for %s produced no usable verdict from %s; refusing\n' "$model" "$reg" + return 1 + ;; + esac +} + +# --------------------------------------------------------------------------- +# Routability +# --------------------------------------------------------------------------- + +# fm_model_routable_decision +# Refuse a model the registry records as rejected or blocked. Return 0 otherwise. +# +# This is a SEPARATE axis from the zero-budget decision and must not be folded +# into it. A model can be perfectly free of cost risk and still be unroutable: the +# model that broke an entire tier was on a flat subscription, so the cost rule had +# nothing to say about it, and only its recorded entitlement refusal makes it +# refusable. Bootstrap catches such a model when a dispatch RULE names it, but the +# explicit --model path bypasses the dispatch config entirely - which is the whole +# reason enforcement lives at spawn too. +# +# Availability is deliberately NOT consulted here. A rate-limited or cooling-down +# model is unavailable, not rejected, and that lives in state/model-health.json on +# its own axis; conflating them would let a transient outage read as a permanent +# routing refusal. +fm_model_routable_decision() { + local model=${1:-} reg verdict + [ -n "$model" ] && [ "$model" != default ] || return 0 + [ -n "$(fm_model_provider_of "$model")" ] || return 0 + reg=$(fm_model_registry_path) + [ -f "$reg" ] || return 0 + command -v jq >/dev/null 2>&1 || return 0 + + verdict=$(jq -r --arg m "$model" --arg schema "$FM_MODEL_SCHEMA_VERSION" ' + if type != "object" or (.schema? // "") != $schema then empty + else + (.models[$m]? // null) as $e + | if $e == null then "allow" + elif (($e.status? // "") == "rejected") or (($e.status? // "") == "blocked") then + "refuse\t" + ($e.status | tostring) + ": " + ($e.status_reason? // "no reason recorded") + else "allow" + end + end + ' "$reg" 2>/dev/null || true) + + case "$verdict" in + allow|'') return 0 ;; + refuse*) + printf 'the model registry records %s as %s (registry: %s)\n' \ + "$model" "${verdict#refuse$'\t'}" "$reg" + return 1 + ;; + esac + return 0 +} + +# --------------------------------------------------------------------------- +# Concurrency / shared quota pool guard +# --------------------------------------------------------------------------- + +# fm_model_concurrency_decision [state-dir] +# Refuse when launching one more worker on this model would exceed the cap the +# registry records for it, or for the quota pool it shares with siblings. +# +# Counts tasks by their recorded state/.meta, which is exactly the set +# teardown has not yet cleaned. That over-counts a dead-but-not-torn-down task, +# which errs toward refusing - the correct direction for a guard whose whole job +# is keeping several workers from collectively breaching one free-tier pool. +fm_model_concurrency_decision() { + local model=${1:-} statedir=${2:-} reg cap pool live meta m pool_of + [ -n "$model" ] && [ "$model" != default ] || return 0 + reg=$(fm_model_registry_path) + [ -f "$reg" ] || return 0 + command -v jq >/dev/null 2>&1 || return 0 + statedir=${statedir:-${STATE:-${FM_STATE_OVERRIDE:-${FM_HOME:-.}/state}}} + + cap=$(jq -r --arg m "$model" '.models[$m]?.limits?.concurrency? // empty' "$reg" 2>/dev/null || true) + case "$cap" in + ''|null) return 0 ;; + *[!0-9]*) return 0 ;; + esac + [ "$cap" -gt 0 ] 2>/dev/null || return 0 + + pool=$(jq -r --arg m "$model" '.models[$m]?.limits?.shared_quota_pool? // empty' "$reg" 2>/dev/null || true) + + live=0 + for meta in "$statedir"/*.meta; do + [ -f "$meta" ] || continue + m=$(sed -n 's/^model=//p' "$meta" | tail -1) + [ -n "$m" ] && [ "$m" != default ] || continue + if [ "$m" = "$model" ]; then + live=$((live + 1)) + continue + fi + # A sibling on the same pool consumes the same free-tier budget, so it counts + # against the same cap. + [ -n "$pool" ] || continue + pool_of=$(jq -r --arg m "$m" '.models[$m]?.limits?.shared_quota_pool? // empty' "$reg" 2>/dev/null || true) + [ "$pool_of" = "$pool" ] && live=$((live + 1)) + done + + if [ "$live" -ge "$cap" ]; then + printf 'concurrency guard refuses %s: %d active worker(s) already on %s, cap is %d (registry: %s -> models."%s".limits.concurrency)\n' \ + "$model" "$live" "${pool:-this model}" "$cap" "$reg" "$model" + return 1 + fi + return 0 +} + +# --------------------------------------------------------------------------- +# Referential integrity: crew-dispatch.json x models.json +# --------------------------------------------------------------------------- + +# fm_model_registry_integrity [now-epoch] +# Echo one line per problem; return 1 when any were found. This is the check that +# catches a bad model at CONFIG-EDIT time rather than at dispatch time: adding a +# rule that names a model with no probe record fails here, before a worker is ever +# launched against it. +fm_model_registry_integrity() { + local dispatch=$1 reg=$2 now=${3:-} out + [ -f "$dispatch" ] || return 0 + [ -f "$reg" ] || return 0 + command -v jq >/dev/null 2>&1 || return 0 + now=${now:-$(date -u +%s)} + out=$(jq -r -n \ + --slurpfile d "$dispatch" \ + --slurpfile r "$reg" \ + --argjson now "$now" \ + ' + def profiles($v): + if ($v | type) == "array" then $v + elif ($v | type) == "object" then [$v] + else [] end; + def named: + ($d[0] // {}) as $cfg + | ([ (($cfg.rules // [])[]? | profiles(.use?)[]?) + , (profiles($cfg.default // null)[]?) ] + | map(.model? // empty) + | map(select(type == "string" and (. != "default") and (test("/")))) + | unique); + def approved: ["approved-primary","approved-fallback","approved-specialist"]; + def secs($iso): + ($iso // "") as $t + | if ($t | length) == 0 then null + else (try ($t | sub("Z$"; "+0000") | strptime("%Y-%m-%dT%H:%M:%S%z") | mktime) + catch (try ($t | sub("Z$"; "") | sub("\\.[0-9]+$"; "") | strptime("%Y-%m-%dT%H:%M") | mktime) catch null)) + end; + ($r[0] // {}) as $reg + | named + | map(. as $m + | ($reg.models[$m]? // null) as $e + | if $e == null then + "MODEL_REGISTRY: " + $m + " is named in config/crew-dispatch.json but absent from config/models.json" + elif (approved | index($e.status? // "") | not) then + "MODEL_REGISTRY: " + $m + " is routed but its registry status is " + ($e.status? // "unset") + + (if ($e.status_reason? // "") != "" then " (" + $e.status_reason + ")" else "" end) + elif (($e.evidence?.probe?.at? // "") | length) == 0 then + "MODEL_REGISTRY: " + $m + " is routed with no live-probe record; every routed model must be probe-verified" + else + (secs($e.evidence.probe.at)) as $at + | (($reg.observation?.levels?[($e.observation_level? // "O4")]?.probe_max_age_days?) // null) as $maxd + | if $at == null then + "MODEL_REGISTRY: " + $m + " has an unparseable probe timestamp \"" + ($e.evidence.probe.at | tostring) + "\"" + elif $maxd != null and (($now - $at) > ($maxd * 86400)) then + "MODEL_REGISTRY: " + $m + " probe evidence is " + ((($now - $at) / 86400) | floor | tostring) + + "d old, past the " + ($maxd | tostring) + "d limit for " + ($e.observation_level? // "O4") + else empty + end + end) + | .[] + ' 2>/dev/null || true) + if [ -n "$out" ]; then + printf '%s\n' "$out" + return 1 + fi + return 0 +} + +# --------------------------------------------------------------------------- +# Price drift +# --------------------------------------------------------------------------- + +# fm_model_price_drift +# Echo one line per drifted model; return 1 when any drifted. Free, local file +# reads only - no network, no probe, no tokens. +# +# This is the only check that can catch a repricing, because a NAME-based +# allowlist is structurally blind to it: the thing that makes a name safe is its +# price, and the price lives in a catalogue the provider rewrites underneath you. +# Comparing two numbers is what makes drift detectable at all. +fm_model_price_drift() { + local reg=$1 out + [ -f "$reg" ] || return 0 + command -v jq >/dev/null 2>&1 || return 0 + # jq cannot open the catalogue files itself, so the comparison is driven from + # the shell: read each declared catalogue once, then diff it against the stored + # prices for that provider's models. + out=$(fm_model_price_drift_scan "$reg") + if [ -n "$out" ]; then + printf '%s\n' "$out" + return 1 + fi + return 0 +} + +# Internal: the actual drift scan. Split out so fm_model_price_drift keeps a +# single simple contract and this part stays testable on its own. +fm_model_price_drift_scan() { + local reg=$1 prov src found + command -v jq >/dev/null 2>&1 || return 0 + while IFS=$'\t' read -r prov src; do + [ -n "$prov" ] && [ -n "$src" ] || continue + # A declared catalogue that is not readable is reported, not silently skipped: + # a drift check that quietly stops running is indistinguishable from one that + # keeps passing. + if [ ! -f "$src" ]; then + printf 'MODEL_PRICE: %s catalogue source is unreadable: %s\n' "$prov" "$src" + continue + fi + found=$(jq -r -n \ + --slurpfile reg "$reg" \ + --slurpfile cat "$src" \ + --arg prov "$prov" \ + ' + def costmap: + (($cat[0] // {})) as $c + | if ($c | type) != "object" then {} + elif ($c[$prov]?.models? | type) == "array" then + ($c[$prov].models | map(select(type == "object")) | map({key: (.id | tostring), value: (.cost // {})}) | from_entries) + else + ($c | to_entries | map(select((.value | type) == "object")) + | map({key: ((.value.id // .key) | tostring), value: (.value.cost // {})}) | from_entries) + end; + costmap as $cm + | ($reg[0] // {}) as $r + | [ ($r.models // {}) | to_entries[] + | select((.value.provider? // (.key | split("/")[0])) == $prov) + | .key as $k | .value as $e + | ($e.model_id? // ($k | split("/") | .[1:] | join("/"))) as $id + | ($cm[$id]? // null) as $now + | select($now != null) + | ( # An allowlisted zero-priced model that is no longer free is the + # repricing case, and it is critical: the route must suspend. + if (($r.zero_budget?.allowlist[$k]? // null) != null) and ([$now[]?] | any(. != 0)) then + "MODEL_PRICE: " + $k + " is on the zero-budget allowlist but its catalogue price is no longer zero (" + ($now | tostring) + "); suspend this route" + elif ($e.price_at_verification? != null) + and (($e.price_at_verification.input? // 0) != ($now.input? // 0) + or ($e.price_at_verification.output? // 0) != ($now.output? // 0)) then + "MODEL_PRICE: " + $k + " price drifted from " + ($e.price_at_verification | tostring) + " to " + ($now | tostring) + else empty + end ) + ] + | .[] + ' 2>/dev/null || true) + [ -z "$found" ] || printf '%s\n' "$found" + done </dev/null || true) +EOF + return 0 +} + +# --------------------------------------------------------------------------- +# Probe classification +# --------------------------------------------------------------------------- + +# fm_model_probe_classify +# Echo one token for the four measured response shapes. The classifier matters +# because a local configuration typo and a server-side entitlement refusal are +# different facts with different handlers: one is a config error, the other means +# the account will never be served that model. +fm_model_probe_classify() { + local rc=${1:-1} out=${2:-} + if [ "$rc" = 0 ]; then + printf 'ok\n' + return 0 + fi + case "$out" in + *'is not supported when using'*) printf 'entitlement-refused\n'; return 0 ;; + *'not found for provider'*) printf 'unknown-model\n'; return 0 ;; + *'Using custom model id'*) printf 'unknown-model\n'; return 0 ;; + *'Unknown provider'*) printf 'client-error\n'; return 0 ;; + esac + printf 'unclassified\n' +} + +# --------------------------------------------------------------------------- +# Promotion dormancy +# --------------------------------------------------------------------------- + +# fm_model_promotion_state [registry-file] +# Echo "active" or "dormant\t". Activation is a +# config and data condition, never a code change: write the instrument, flip the +# flag. Naming WHICH condition is unmet is the point - a dormant trigger nobody +# can check is indistinguishable from a rejected one. +fm_model_promotion_state() { + local reg=${1:-} enabled instrument ledger + reg=${reg:-$(fm_model_registry_path)} + [ -f "$reg" ] || { printf 'dormant\tno model registry at %s\n' "$reg"; return 0; } + command -v jq >/dev/null 2>&1 || { printf 'dormant\tjq is not installed\n'; return 0; } + + enabled=$(jq -r '.promotion?.enabled? // false' "$reg" 2>/dev/null || echo false) + instrument=$(jq -r '.promotion?.requires_instrument? // empty' "$reg" 2>/dev/null || true) + + if [ "$enabled" != true ]; then + printf 'dormant\tpromotion.enabled is false in %s\n' "$reg" + return 0 + fi + if [ -z "$instrument" ]; then + printf 'dormant\tpromotion.requires_instrument is unset\n' + return 0 + fi + # The data half: the named evidence instrument must actually exist and be + # producing terminal task lines. Until it does, P1/P2/P3 are not merely unmet - + # they are not COMPUTABLE, and reporting them as unmet would be a false negative. + ledger="${STATE:-${FM_STATE_OVERRIDE:-${FM_HOME:-.}/state}}/wake-outcome-ledger.jsonl" + if [ ! -s "$ledger" ]; then + printf 'dormant\tevidence instrument "%s" has produced no records yet (%s)\n' "$instrument" "$ledger" + return 0 + fi + if ! grep -q '"kind"[[:space:]]*:[[:space:]]*"task-terminal"' "$ledger" 2>/dev/null; then + printf 'dormant\tevidence instrument "%s" exists but records no task-terminal lines yet\n' "$instrument" + return 0 + fi + printf 'active\n' +} diff --git a/bin/fm-model-verify.sh b/bin/fm-model-verify.sh new file mode 100755 index 00000000000..b64cab555e9 --- /dev/null +++ b/bin/fm-model-verify.sh @@ -0,0 +1,361 @@ +#!/usr/bin/env bash +# fm-model-verify.sh - live entitlement probes and local price-drift checks for +# the models config/models.json says this home routes to. +# Usage: fm-model-verify.sh [--all] [--model ] [--drift-only] +# [--dry-run] [--force-probe] [--json] [--timeout ] +# Default: probe only routed models whose recorded evidence is older +# than their observation level's interval, then run the price-drift +# comparison. Prints ONE line per model that needs firstmate's attention +# and nothing at all when everything is current - it runs on the +# session-start path, so silence is the normal outcome. +# --all probe every routed model regardless of its interval. +# --model probe exactly this model, ignoring the interval. +# --drift-only skip probing entirely; only compare stored prices. +# --dry-run print what would be probed, probe nothing. +# --force-probe probe a model the zero-budget cost decision refuses; the +# only override, and it announces itself on stdout. +# --json emit the health record to stdout instead of the summary. +# --timeout hard ceiling in seconds for the whole sweep (default 90). +# Exits 0 when everything is current, 2 when any line was printed. +# +# NO PROBE WITHOUT A COST VERDICT. Every probe path, including an explicit +# --model, consults fm_model_zero_budget_decision before issuing a live request: +# the instrument that catches an entitlement error is itself a billable act on a +# metered provider, which is why cost class is established before entitlement +# and not after. A typed --model is not authorization to spend money - under the +# zero-budget rule spending is never implied, only explicitly flagged, and the +# flag is --force-probe. +# +# WHY BOTH CHECKS, FOREVER. The two things that decay are not properties of the +# model - they are properties of the ACCOUNT and of the provider's price list, and +# both change without warning. A model can be perfectly stable while its +# entitlement is revoked and its price is raised. So observation never reaches +# zero: at the maintenance floor these two checks still run. +# 1. The entitlement probe catches a routed model this account cannot actually +# use. That failure has happened here: a model was configured from a +# plausible name, never probed, and every dispatch to that tier failed at +# launch until an investigation found it. +# 2. The price-drift check is the ONLY thing that can catch a repricing, +# because a name-based allowlist is structurally blind to it. It is a local +# file read and costs nothing. +# +# STDIN IS ALWAYS CLOSED ON A PROBE. `pi -p` can hang indefinitely with stdin +# open. A wedged probe inside the session-start path would present to supervision +# as a stale session - the monitor becoming the fault - so every probe redirects +# from /dev/null and additionally runs under `timeout`. +# +# THIS SCRIPT NEVER WRITES config/models.json. Availability is volatile and lives +# in state/model-health.json; the registry's status is a durable ROUTING decision. +# Conflating them would make every transient outage permanently degrade the +# routing table, which is exactly the failure the demotion policy is built to +# avoid. A rate-limited model is unavailable, not demoted. +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" + +usage() { + sed -n '2,19p' "$0" | sed 's/^# \{0,1\}//' +} + +case "${1:-}" in + -h|--help) usage; exit 0 ;; +esac + +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +CONFIG="${FM_CONFIG_OVERRIDE:-$FM_HOME/config}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +# shellcheck source=bin/fm-model-registry-lib.sh +. "$SCRIPT_DIR/fm-model-registry-lib.sh" + +ALL=0 +ONE= +DRIFT_ONLY=0 +DRY_RUN=0 +FORCE_PROBE=0 +AS_JSON=0 +TOTAL_TIMEOUT=90 +PROBE_TIMEOUT=25 + +want= +for a in "$@"; do + if [ -n "$want" ]; then + case "$want" in + model) ONE=$a ;; + timeout) TOTAL_TIMEOUT=$a ;; + esac + want= + continue + fi + case "$a" in + --all) ALL=1 ;; + --model) want=model ;; + --model=*) ONE=${a#--model=} ;; + --drift-only) DRIFT_ONLY=1 ;; + --dry-run) DRY_RUN=1 ;; + --force-probe) FORCE_PROBE=1 ;; + --json) AS_JSON=1 ;; + --timeout) want=timeout ;; + --timeout=*) TOTAL_TIMEOUT=${a#--timeout=} ;; + *) echo "error: unknown argument '$a'" >&2; exit 1 ;; + esac +done +[ -z "$want" ] || { echo "error: --$want requires a value" >&2; exit 1; } +case "$TOTAL_TIMEOUT" in + ''|*[!0-9]*) echo "error: --timeout must be a whole number of seconds" >&2; exit 1 ;; +esac + +REG="$CONFIG/models.json" +[ -f "$REG" ] || { [ "$AS_JSON" = 1 ] && echo '{}'; exit 0; } +command -v jq >/dev/null 2>&1 || { echo "MODEL_VERIFY: jq is required and not installed"; exit 0; } + +if ! err=$(fm_model_registry_validate "$REG"); then + echo "MODEL_REGISTRY: invalid config/models.json - $err" + exit 0 +fi + +HEALTH="$STATE/model-health.json" +NOW_EPOCH=$(date -u +%s) +NOW_ISO=$(date -u +%Y-%m-%dT%H:%M:%SZ) + +# --------------------------------------------------------------------------- +# Select the models to probe +# --------------------------------------------------------------------------- +# A model is due when its recorded probe evidence is older than the interval its +# observation level names. Interval-gating is what keeps this affordable on the +# session-start path: in the steady state nothing is due and the sweep costs one +# file read. +select_due() { + local out rc + # Bind .key/.value to variables BEFORE any `x | y` pipe: inside a pipe the + # input rebinds, so `select(approved | index(.value.status))` would index the + # approved ARRAY rather than the entry. That mistake fails the whole query, and + # a swallowed failure here would read as "nothing is due" - a sweep that never + # probes while appearing healthy. Hence the explicit rc check below. + out=$(jq -r --argjson now "$NOW_EPOCH" --argjson all "$ALL" ' + def secs($iso): + ($iso // "") as $t + | if ($t | length) == 0 then null + else (try ($t | sub("Z$"; "+0000") | strptime("%Y-%m-%dT%H:%M:%S%z") | mktime) catch null) + end; + def approved: ["approved-primary","approved-fallback","approved-specialist"]; + . as $r + | ($r.models // {}) | to_entries[] + | .key as $k | .value as $e + | ($e.status? // "") as $st + | select(approved | index($st)) + | (($r.observation?.levels?[($e.observation_level? // "O4")]?.probe_max_age_days?) // 1) as $maxd + | (secs($e.evidence?.probe?.at?)) as $at + | select($all == 1 or $at == null or (($now - $at) > ($maxd * 86400))) + | [$k, ($e.harness? // "pi"), ($e.provider? // ($k | split("/")[0])), + ($e.model_id? // ($k | split("/") | .[1:] | join("/")))] + | @tsv + ' "$REG" 2>&1) + rc=$? + if [ "$rc" -ne 0 ]; then + echo "MODEL_VERIFY: could not determine which models are due for a probe: $out" >&2 + return 1 + fi + printf '%s\n' "$out" +} + +if [ -n "$ONE" ]; then + DUE=$(jq -r --arg k "$ONE" ' + (.models[$k]? // null) as $e + | if $e == null then empty + else [$k, ($e.harness? // "pi"), ($e.provider? // ($k | split("/")[0])), + ($e.model_id? // ($k | split("/") | .[1:] | join("/")))] | @tsv end' "$REG" 2>/dev/null || true) + if [ -z "$DUE" ]; then + echo "MODEL_VERIFY: $ONE is not in config/models.json" + exit 1 + fi +else + DUE=$(select_due) || exit 2 +fi + +# --------------------------------------------------------------------------- +# Cost gate: no live request without a zero-budget verdict +# --------------------------------------------------------------------------- +# The probe that catches an entitlement error is itself a billable act on a +# metered provider, which is why cost class is established BEFORE entitlement +# and not after. Both paths land here - the interval-gated sweep and an explicit +# --model - because a typed model name is not authorization to spend money. A +# refused model is skipped entirely: no request is issued and its prior health +# record is left untouched. --force-probe is the only override, and it announces +# itself so an authorized billable probe is never invisible in the output. +NEEDS_ACTION=0 +if [ "$DRIFT_ONLY" != 1 ] && [ -n "$DUE" ]; then + ALLOWED= + while IFS=$'\t' read -r key harness provider model_id; do + [ -n "$key" ] || continue + if reason=$(fm_model_zero_budget_decision "$key"); then + ALLOWED="${ALLOWED}${key}"$'\t'"${harness}"$'\t'"${provider}"$'\t'"${model_id}"$'\n' + elif [ "$FORCE_PROBE" = 1 ]; then + echo "MODEL_VERIFY: --force-probe overrides the cost refusal for $key - this probe is a billable act: $reason" + ALLOWED="${ALLOWED}${key}"$'\t'"${harness}"$'\t'"${provider}"$'\t'"${model_id}"$'\n' + else + echo "MODEL_VERIFY: refusing to probe $key - $reason (--force-probe is the only override)" + NEEDS_ACTION=1 + fi + done <\t\t\t\t". +# stdin closed (/dev/null 2>&1; then + printf '%s\t%s\t\t\t%s\n' "$key" 'unprobeable' "$harness is not installed" + return 0 + fi + t0=$(date +%s) + out=$(timeout "$PROBE_TIMEOUT" "$harness" -p --provider "$provider" --model "$model_id" \ + --no-tools --no-session --thinking off 'Reply with the single word: ok' \ + &1) + rc=$? + t1=$(date +%s) + lat=$((t1 - t0)) + if [ "$rc" = 124 ]; then + printf '%s\t%s\t%s\t%s\t%s\n' "$key" 'timeout' "$rc" "$lat" "probe exceeded ${PROBE_TIMEOUT}s" + return 0 + fi + shape=$(fm_model_probe_classify "$rc" "$out") + printf '%s\t%s\t%s\t%s\t%s\n' "$key" "$shape" "$rc" "$lat" "$(printf '%s' "$out" | head -1)" +} + +RESULTS= +if [ "$DRIFT_ONLY" != 1 ] && [ -n "$DUE" ]; then + if [ "$DRY_RUN" = 1 ]; then + while IFS=$'\t' read -r key harness provider model_id; do + [ -n "$key" ] || continue + echo "MODEL_VERIFY: would probe $key via $harness ($provider/$model_id)" + done < "$TMPD/$n.out" 2>/dev/null & + done </dev/null || true; done + wait 2>/dev/null || true + RESULTS=$(cat "$TMPD"/*.out 2>/dev/null || true) +fi + +# --------------------------------------------------------------------------- +# Merge into the volatile health record +# --------------------------------------------------------------------------- +mkdir -p "$STATE" +[ -f "$HEALTH" ] || printf '{"models":{}}\n' > "$HEALTH" +if ! jq -e . "$HEALTH" >/dev/null 2>&1; then + printf '{"models":{}}\n' > "$HEALTH" +fi + +if [ -n "$RESULTS" ]; then + merged=$(printf '%s\n' "$RESULTS" | jq -R -s --arg at "$NOW_ISO" --slurpfile prior "$HEALTH" ' + ($prior[0] // {"models":{}}) as $p + | [ split("\n")[] | select(length > 0) | split("\t") + | { key: .[0], + value: { shape: .[1], rc: (.[2] | tonumber? // null), + latency_s: (.[3] | tonumber? // null), + detail: (.[4] // ""), at: $at } } ] + | from_entries + | . as $new + | reduce ($new | to_entries[]) as $e + ($p; .models[$e.key] = ( + ($p.models[$e.key]? // {}) as $old + | ($e.value.shape) as $shape + | $e.value + + { state: (if $shape == "ok" then "available" + elif $shape == "unprobeable" then ($old.state? // "unknown") + else "unavailable" end), + consecutive_failures: ( + if $shape == "ok" then 0 + elif $shape == "unprobeable" then ($old.consecutive_failures? // 0) + else (($old.consecutive_failures? // 0) + 1) end) } + )) + | .updated_at = $at + ') + printf '%s\n' "$merged" > "$HEALTH" +fi + +if [ "$AS_JSON" = 1 ]; then + cat "$HEALTH" + exit 0 +fi + +# --------------------------------------------------------------------------- +# Report only what firstmate should act on +# --------------------------------------------------------------------------- +if [ -n "$RESULTS" ]; then + while IFS=$'\t' read -r key shape rc lat detail; do + [ -n "$key" ] || continue + case "$shape" in + ok) ;; + unprobeable) + echo "MODEL_VERIFY: $key could not be probed - $detail" + NEEDS_ACTION=1 + ;; + entitlement-refused) + echo "MODEL_VERIFY: $key is REFUSED by the provider for this account - it must not be routed to: $detail" + NEEDS_ACTION=1 + ;; + unknown-model) + echo "MODEL_VERIFY: $key was not recognised by its provider (identity error, not an outage): $detail" + NEEDS_ACTION=1 + ;; + client-error) + echo "MODEL_VERIFY: $key failed locally before the request left the machine (configuration error): $detail" + NEEDS_ACTION=1 + ;; + timeout) + echo "MODEL_VERIFY: $key probe timed out after ${PROBE_TIMEOUT}s" + NEEDS_ACTION=1 + ;; + *) + echo "MODEL_VERIFY: $key probe returned an unrecognised result (rc=$rc, ${lat}s): $detail" + NEEDS_ACTION=1 + ;; + esac + done < + local dir=$1 timeout_secs=$2 have_timeout=none + shift 2 + if command -v timeout >/dev/null 2>&1; then have_timeout=timeout + elif command -v gtimeout >/dev/null 2>&1; then have_timeout=gtimeout + elif command -v perl >/dev/null 2>&1; then have_timeout=perl + fi + case "$have_timeout" in + timeout) ( cd "$dir" && timeout "$timeout_secs" no-mistakes "$@" ) ;; + gtimeout) ( cd "$dir" && gtimeout "$timeout_secs" no-mistakes "$@" ) ;; + perl) ( cd "$dir" && perl -e 'my $t = shift; my $pid = fork; die "fork failed" unless defined $pid; if (!$pid) { setpgrp(0, 0); exec @ARGV } local $SIG{ALRM} = sub { kill "TERM", -$pid; select undef, undef, undef, 0.2; kill "KILL", -$pid; exit 124 }; alarm $t; waitpid $pid, 0; exit($? >> 8)' "$timeout_secs" no-mistakes "$@" ) ;; + *) return 1 ;; + esac +} + +fm_nm_run_checked() { # + fm_nm_run_bounded "$@" 2>/dev/null +} + +fm_nm_run() { # + fm_nm_run_checked "$@" || true +} + +fm_nm_trim() { + local s=${1:-} + s="${s#"${s%%[![:space:]]*}"}" + s="${s%"${s##*[![:space:]]}"}" + printf '%s' "$s" +} + +fm_nm_strip_quotes() { + local s + s=$(fm_nm_trim "${1:-}") + case "$s" in + \"*\") s=${s#\"}; s=${s%\"} ;; + esac + fm_nm_trim "$s" +} + +# Scalar value of a TOON key in captured `axi status` output $1. +fm_nm_field() { # + printf '%s\n' "$1" | sed -n "s/^[[:space:]]*$2:[[:space:]]*\(.*\)/\1/p" | head -1 +} + +# 0 if run head $2 matches worktree $1's code identity, per the same rule +# everywhere this attribution is needed: +# - missing/empty head: cannot bind; reject +# - equal commits (short or full SHA): match +# - worktree HEAD is an ancestor of run head: match (pipeline fix commits on +# the same history advanced the run tip past local HEAD) +# - run head is a strict ancestor of worktree HEAD, or diverged: no match +# (local work advanced outside the run, or the branch tip was rewritten) +fm_nm_head_matches_worktree() { # + local wt=$1 run_head=$2 local_full run_full + [ -n "$run_head" ] || return 1 + local_full=$(git -C "$wt" rev-parse HEAD 2>/dev/null) || return 1 + run_full=$(git -C "$wt" rev-parse --verify "${run_head}^{commit}" 2>/dev/null) || return 1 + [ "$run_full" = "$local_full" ] && return 0 + git -C "$wt" merge-base --is-ancestor "$local_full" "$run_full" 2>/dev/null +} diff --git a/bin/fm-on.sh b/bin/fm-on.sh new file mode 100755 index 00000000000..5e24f2cef1d --- /dev/null +++ b/bin/fm-on.sh @@ -0,0 +1,112 @@ +#!/usr/bin/env bash +# Execute one tracked Firstmate command in a configured remote secondmate home. +# +# Usage: +# fm-on.sh [args...] +# +# Routes come only from remote records in data/secondmates.md. A record names an +# SSH config alias, remote Firstmate code root, and remote FM_HOME. A host alias +# may be used directly only when exactly one record selects it; an ambiguous +# alias is refused. The command must be a genuine executable in this checkout's +# bin/fm-*.sh namespace. No per-command table exists. +# +# argv is encoded as one NUL-delimited stream and passed through the fixed +# fm-remote-entrypoint.sh. stdin remains the caller's stdin, stdout and stderr +# remain separate, and ssh's exit status is returned unchanged. OpenSSH never +# receives an auto-retry instruction here. Exit 255 therefore means unavailable +# transport or unknown remote completion and must be reconciled by the semantic +# caller, never blindly repeated by this layer. +# +# The SSH alias keeps normal public-key and strict host-key policy in ~/.ssh. +# This command explicitly disables agent forwarding, forwarding setup, and +# configured SendEnv patterns. The remote entrypoint executes the selected +# command under an empty environment with only its fixed runtime values. +# +# ServerAliveInterval/ServerAliveCountMax arm dead-peer detection so a vanished +# peer (a reboot, a dropped link) becomes a bounded ssh failure (exit 255) +# instead of an indefinite hang on a half-open TCP connection. The remote +# sshd answers keepalive probes independently of whatever the remote command +# is doing, so a legitimately long-but-alive remote command is never falsely +# killed. FM_SSH_ALIVE_INTERVAL and FM_SSH_ALIVE_COUNT_MAX override the +# defaults; the worst-case detection window is roughly interval * count. +set -eu + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +REG="$DATA/secondmates.md" +PROTOCOL=1 + +# shellcheck source=bin/fm-secondmate-registry-lib.sh +. "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" + +die() { printf 'error: %s\n' "$1" >&2; exit 1; } +usage() { sed -n '2,23p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } + +encode_base64() { + base64 | tr -d '\n' +} + +[ "$#" -ge 2 ] || usage +ROUTE=$1 +COMMAND=$2 +shift 2 + +case "$ROUTE" in ''|-*|*[!A-Za-z0-9._-]*) die "remote route must be a safe secondmate id or SSH alias: $ROUTE" ;; esac +case "$COMMAND" in + fm-*.sh) ;; + *) die "remote command must be a basename in the fm-*.sh namespace: $COMMAND" ;; +esac +case "$COMMAND" in */*|*..*) die "remote command must not contain a path or traversal: $COMMAND" ;; esac +LOCAL_COMMAND="$FM_ROOT/bin/$COMMAND" +[ -f "$LOCAL_COMMAND" ] && [ ! -L "$LOCAL_COMMAND" ] && [ -x "$LOCAL_COMMAND" ] \ + || die "remote command is not a genuine tracked executable in this Firstmate checkout: $COMMAND" +git -C "$FM_ROOT" ls-files --error-unmatch "bin/$COMMAND" >/dev/null 2>&1 \ + || die "remote command is not tracked by this Firstmate checkout: $COMMAND" +[ -f "$REG" ] && [ ! -L "$REG" ] || die "no safe secondmate registry at $REG" + +MATCHES=0 +HOST= +ROOT= +HOME_PATH= +while IFS= read -r line || [ -n "$line" ]; do + case "$line" in '- '*) ;; *) continue ;; esac + secondmate_registry_parse_line "$line" || die "malformed secondmate registry entry: $line" + [ "$SECONDMATE_REGISTRY_REMOTE" -eq 1 ] || continue + if [ "$SECONDMATE_REGISTRY_ID" = "$ROUTE" ] || [ "$SECONDMATE_REGISTRY_HOST" = "$ROUTE" ]; then + MATCHES=$((MATCHES + 1)) + HOST=$SECONDMATE_REGISTRY_HOST + ROOT=$SECONDMATE_REGISTRY_ROOT + HOME_PATH=$SECONDMATE_REGISTRY_HOME + fi +done < "$REG" +[ "$MATCHES" -gt 0 ] || die "no remote secondmate or SSH alias matches '$ROUTE'" +[ "$MATCHES" -eq 1 ] || die "remote route '$ROUTE' is ambiguous across $MATCHES configured secondmates; use a secondmate id" +case "$HOST" in ''|-*|*[!A-Za-z0-9._-]*) die "configured SSH alias is unsafe: $HOST" ;; esac +case "$ROOT" in /*) ;; *) die "configured remote root is not absolute: $ROOT" ;; esac +case "$HOME_PATH" in /*) ;; *) die "configured remote home is not absolute: $HOME_PATH" ;; esac +case "$ROOT$HOME_PATH" in *$'\n'*|*$'\r'*|*$'\t'*) die "configured remote root or home contains control characters" ;; esac +for configured_path in "$ROOT" "$HOME_PATH"; do + case "/$configured_path/" in */../*|*/./*) die "configured remote root or home contains traversal components" ;; esac + case "$configured_path" in *'//'*) die "configured remote root or home contains an empty path component" ;; esac +done + +ROOT_B64=$(printf '%s' "$ROOT" | encode_base64) +HOME_B64=$(printf '%s' "$HOME_PATH" | encode_base64) +ARGV_B64=$(printf '%s\0' "$COMMAND" "$@" | encode_base64) +SSH_BIN=${FM_SSH_BIN:-ssh} +ALIVE_INTERVAL=${FM_SSH_ALIVE_INTERVAL:-15} +ALIVE_COUNT_MAX=${FM_SSH_ALIVE_COUNT_MAX:-3} +case "$ALIVE_INTERVAL" in ''|*[!0-9]*) die "FM_SSH_ALIVE_INTERVAL must be a positive integer: $ALIVE_INTERVAL" ;; esac +case "$ALIVE_COUNT_MAX" in ''|*[!0-9]*) die "FM_SSH_ALIVE_COUNT_MAX must be a positive integer: $ALIVE_COUNT_MAX" ;; esac +[ "$ALIVE_INTERVAL" -gt 0 ] || die "FM_SSH_ALIVE_INTERVAL must be a positive integer: $ALIVE_INTERVAL" +[ "$ALIVE_COUNT_MAX" -gt 0 ] || die "FM_SSH_ALIVE_COUNT_MAX must be a positive integer: $ALIVE_COUNT_MAX" + +"$SSH_BIN" \ + -o ForwardAgent=no \ + -o ClearAllForwardings=yes \ + -o 'SendEnv=-*' \ + -o "ServerAliveInterval=$ALIVE_INTERVAL" \ + -o "ServerAliveCountMax=$ALIVE_COUNT_MAX" \ + -- "$HOST" fm-remote-entrypoint.sh "$PROTOCOL" "$ROOT_B64" "$HOME_B64" "$ARGV_B64" diff --git a/bin/fm-pending-reply-lib.sh b/bin/fm-pending-reply-lib.sh index a45c8378ea2..3d656f22b07 100755 --- a/bin/fm-pending-reply-lib.sh +++ b/bin/fm-pending-reply-lib.sh @@ -337,6 +337,20 @@ fm_pending_reply_confirm_delivery() { # return 2 } +# Preserve an expectation when a remote transport disconnect makes delivery +# completion unknowable. This never resolves or retries the request; it moves +# the existing prepared record to the owner's explicit unknown-delivery path. +fm_pending_reply_mark_delivery_unknown() { # + local state=$1 corr=$2 rec phase + rec=$(fm_pending_reply_path "$state" "$corr") + [ -f "$rec" ] || return 1 + phase=$(fm_pending_reply_get "$rec" phase) + case "$phase" in awaiting_report|delivery_unknown) ;; + *) return 1 ;; + esac + fm_pending_reply_set "$rec" phase delivery_unknown +} + fm_pending_reply_reconcile_delivery() { # local state=$1 corr=$2 rec delivered marker entry delivery_state value epoch local grace now age phase @@ -941,7 +955,7 @@ fm_pending_reply_tick_one() { # [secondmate- # Never scrapes secondmate conversation; uses only parent status, backend busy # state, and optional secondmate-home wrong-home path checks. fm_pending_reply_tick() { # - local state=$1 dir rec corr task_id phase delivered meta backend target label busy sm_home harness + local state=$1 dir rec corr task_id phase delivered meta backend target label busy sm_home harness remote_host local observation observation_task found i local -a observation_tasks=() observation_values=() dir=$(fm_pending_reply_dir "$state") @@ -1004,10 +1018,15 @@ fm_pending_reply_tick() { # sm_home= harness= if [ -f "$meta" ]; then + remote_host=$(fm_meta_get "$meta" remote_host) backend=$(fm_backend_of_meta "$meta") target=$(fm_backend_target_of_meta "$meta") sm_home=$(fm_meta_get "$meta" home) harness=$(fm_meta_get "$meta" harness) + if [ -n "$remote_host" ]; then + target="remote:$task_id" + sm_home= + fi if [ -n "$target" ]; then label="fm-$task_id" observation= @@ -1020,7 +1039,13 @@ fm_pending_reply_tick() { # break done if [ "$found" = 0 ]; then - observation=$(fm_pending_reply_backend_observation "$backend" "$target" "$label" "$harness") + if [ -n "$remote_host" ]; then + observation=$("$_FM_PENDING_REPLY_LIB_DIR/fm-on.sh" "$task_id" \ + fm-remote-secondmate-control.sh observe "$task_id" < /dev/null 2>/dev/null || printf 'unknown') + case "$observation" in busy|idle|fallback-idle|unknown) ;; *) observation=unknown ;; esac + else + observation=$(fm_pending_reply_backend_observation "$backend" "$target" "$label" "$harness") + fi observation_tasks+=("$task_id") observation_values+=("$observation") fi diff --git a/bin/fm-pr-check.sh b/bin/fm-pr-check.sh index 57858db3762..f9d187a0965 100755 --- a/bin/fm-pr-check.sh +++ b/bin/fm-pr-check.sh @@ -5,6 +5,15 @@ # live only in a private sidecar and are never interpolated into shell source. # A GitHub pull request URL and a GitLab merge request URL are both accepted, # including a merge request on a self-hosted GitLab instance. +# +# The identity record this writes into is the task's meta while the task is +# live, and the durable landing record bin/fm-teardown.sh leaves behind once the +# task is released with its request still unlanded, so releasing a worker before +# its request lands no longer ends the watch. When neither exists - a task +# released before landing records did - the record is rebuilt from a forge read +# of the request itself. An existing landing record must be valid and name the +# requested URL before any forge read or poll mutation. A request the forge +# cannot resolve refuses instead. # Usage: fm-pr-check.sh set -eu @@ -33,11 +42,51 @@ PROJECT_PATH=$FM_PR_PATH NUMBER=$FM_PR_NUMBER # Task-derived paths are constructed only after the canonical ID validation. -META="$STATE/$ID.meta" -if [ ! -f "$META" ] || [ -L "$META" ] || [ "$(fm_pr_file_link_count "$META")" != 1 ]; then +# A live task records into its meta. A task released before its pull request +# landed has only the durable landing record bin/fm-teardown.sh left behind, and +# a task released before landing records existed has neither - the pull request +# is then the only authority for its own identity, so the record is rebuilt from +# a forge read rather than from anything the caller asserted. A request the forge +# cannot resolve still refuses: absence is never treated as evidence. +RECONSTRUCTED=0 +if ! RECORD=$(fm_pr_identity_record_path "$STATE" "$ID"); then + LANDING="$STATE/$ID.landing" + if [ -e "$LANDING" ] || [ -L "$LANDING" ]; then + echo "error: task landing record is invalid" >&2 + exit 1 + fi + if ! fm_pr_forge_view "$URL"; then + echo "error: task metadata is unavailable" >&2 + echo "No record for task $ID, and $URL could not be resolved at its forge." >&2 + exit 1 + fi + RECONSTRUCTED_HEAD=$FM_PR_FORGE_HEAD + if ! fm_pr_landing_record_write "$STATE" "$ID" "$URL" "$RECONSTRUCTED_HEAD" ""; then + echo "error: task landing record could not be written" >&2 + exit 1 + fi + RECORD=$(fm_pr_identity_record_path "$STATE" "$ID") || { + echo "error: task metadata is unavailable" >&2 + exit 1 + } + RECONSTRUCTED=1 + printf 'rebuilt: state/%s.landing from %s\n' "$ID" "$URL" +fi +if [ ! -f "$RECORD" ] || [ -L "$RECORD" ] || [ "$(fm_pr_file_link_count "$RECORD")" != 1 ]; then echo "error: task metadata is unavailable" >&2 exit 1 fi +if [ "$RECORD" = "$STATE/$ID.landing" ]; then + fm_pr_metadata_identity_parse "$RECORD" || { + echo "error: task landing record is invalid" >&2 + exit 1 + } + RECORD_URL=$FM_PR_META_URL + if [ "$RECORD_URL" != "$URL" ]; then + echo "error: task landing record names $RECORD_URL, but $URL was requested" >&2 + exit 1 + fi +fi # A prior exact merged result may have queued its durable wake immediately # before interruption. @@ -69,48 +118,56 @@ fi # bin/fm-teardown.sh reads the head from the forge at teardown rather than from # metadata and falls back to its provider-agnostic content check, and # bin/fm-review-diff.sh resolves the head from the remote when none is recorded. -WT=$(grep '^worktree=' "$META" | tail -1 | cut -d= -f2- || true) +WT=$(grep '^worktree=' "$RECORD" | tail -1 | cut -d= -f2- || true) PR_HEAD= -if [ "$PROVIDER" = github ] && [ -n "$WT" ] && [ -d "$WT" ] && command -v gh >/dev/null 2>&1; then +if [ "$RECONSTRUCTED" = 1 ]; then + # The rebuild above already asked the forge for this exact pull request. + PR_HEAD=$RECONSTRUCTED_HEAD +elif [ "$PROVIDER" = github ] && [ -n "$WT" ] && [ -d "$WT" ] && command -v gh >/dev/null 2>&1; then if REMOTE_HEAD=$(cd "$WT" && gh pr view "$URL" --json headRefOid -q .headRefOid 2>/dev/null) \ && fm_pr_head_valid "$REMOTE_HEAD"; then PR_HEAD=$REMOTE_HEAD fi +elif fm_pr_forge_view "$URL"; then + # A released task has no worktree left to resolve the request from, so the + # head comes straight from the request. This also recovers the head for a live + # task whose worktree is already gone, which previously recorded none. + PR_HEAD=$FM_PR_FORGE_HEAD fi -META_TMP= +RECORD_TMP= pr_check_cleanup() { fm_pr_poll_cleanup - [ -z "$META_TMP" ] || rm -f -- "$META_TMP" + [ -z "$RECORD_TMP" ] || rm -f -- "$RECORD_TMP" } trap pr_check_cleanup EXIT trap 'exit 1' HUP INT TERM fm_pr_poll_prepare "$STATE" "$ID" "$PROVIDER" "$URL" "$HOST" "$PROJECT_PATH" "$NUMBER" "$SCRIPT_DIR/fm-pr-poll.sh" \ || { echo "error: could not prepare PR poll" >&2; exit 1; } -META_DEVICE=$(fm_pr_file_device "$META") || exit 1 +RECORD_DEVICE=$(fm_pr_file_device "$RECORD") || exit 1 STATE_DEVICE=$(fm_pr_file_device "$STATE") || exit 1 -[ "$META_DEVICE" = "$STATE_DEVICE" ] || { echo "error: task metadata is unavailable" >&2; exit 1; } -META_TMP=$(mktemp "$STATE/.fm-pr-meta.XXXXXX") || exit 1 +[ "$RECORD_DEVICE" = "$STATE_DEVICE" ] || { echo "error: task metadata is unavailable" >&2; exit 1; } +RECORD_TMP=$(mktemp "$STATE/.fm-pr-meta.XXXXXX") || exit 1 while IFS= read -r line || [ -n "$line" ]; do case "$line" in pr=*|pr_head=*) ;; - *) printf '%s\n' "$line" >> "$META_TMP" || exit 1 ;; + *) printf '%s\n' "$line" >> "$RECORD_TMP" || exit 1 ;; esac -done < "$META" -printf 'pr=%s\n' "$URL" >> "$META_TMP" || exit 1 -[ -z "$PR_HEAD" ] || printf 'pr_head=%s\n' "$PR_HEAD" >> "$META_TMP" || exit 1 -chmod 0600 "$META_TMP" || exit 1 -fm_pr_private_file_valid "$META_TMP" 600 "$STATE_DEVICE" || exit 1 -fm_pr_metadata_identity_parse "$META_TMP" || exit 1 +done < "$RECORD" +printf 'pr=%s\n' "$URL" >> "$RECORD_TMP" || exit 1 +[ -z "$PR_HEAD" ] || printf 'pr_head=%s\n' "$PR_HEAD" >> "$RECORD_TMP" || exit 1 +chmod 0600 "$RECORD_TMP" || exit 1 +fm_pr_private_file_valid "$RECORD_TMP" 600 "$STATE_DEVICE" || exit 1 +fm_pr_metadata_identity_parse "$RECORD_TMP" || exit 1 [ "$FM_PR_META_PROVIDER" = "$PROVIDER" ] && [ "$FM_PR_META_URL" = "$URL" ] \ && [ "$FM_PR_META_HOST" = "$HOST" ] && [ "$FM_PR_META_PATH" = "$PROJECT_PATH" ] \ && [ "$FM_PR_META_NUMBER" = "$NUMBER" ] || exit 1 -fm_pr_regular_destination_on_device_or_absent "$META" "$STATE_DEVICE" || exit 1 -mv -f -- "$META_TMP" "$META" || exit 1 -META_TMP= -fm_pr_private_file_valid "$META" 600 "$STATE_DEVICE" || exit 1 -fm_pr_metadata_identity_parse "$META" || exit 1 +fm_pr_regular_destination_on_device_or_absent "$RECORD" "$STATE_DEVICE" || exit 1 +mv -f -- "$RECORD_TMP" "$RECORD" || exit 1 +RECORD_TMP= +fm_pr_private_file_valid "$RECORD" 600 "$STATE_DEVICE" || exit 1 +fm_pr_metadata_identity_parse "$RECORD" || exit 1 [ "$FM_PR_META_PROVIDER" = "$PROVIDER" ] && [ "$FM_PR_META_URL" = "$URL" ] \ && [ "$FM_PR_META_HOST" = "$HOST" ] && [ "$FM_PR_META_PATH" = "$PROJECT_PATH" ] \ && [ "$FM_PR_META_NUMBER" = "$NUMBER" ] || exit 1 diff --git a/bin/fm-pr-lib.sh b/bin/fm-pr-lib.sh index b70d8468894..6f9fb2cc9bf 100755 --- a/bin/fm-pr-lib.sh +++ b/bin/fm-pr-lib.sh @@ -16,6 +16,11 @@ # after its durable wake is appended. # The receipt binds the terminal observation to the canonical registration and # lets a restart finish fixed-path removal without executing state-file bytes. +# +# A task released before its pull request lands keeps no meta, so its identity +# record is the durable landing record bin/fm-teardown.sh leaves behind. Every +# consumer resolves "the task's PR identity record" through +# fm_pr_identity_record_path rather than naming the meta directly. FM_PR_PROVIDER= FM_PR_URL= @@ -34,6 +39,8 @@ FM_PR_META_URL= FM_PR_META_HOST= FM_PR_META_PATH= FM_PR_META_NUMBER= +FM_PR_FORGE_STATE= +FM_PR_FORGE_HEAD= FM_PR_REG_ID= FM_PR_REG_PROVIDER= FM_PR_REG_URL= @@ -327,6 +334,139 @@ fm_pr_metadata_identity_parse() { [ -n "$FM_PR_META_URL" ] } +# Landing record layout: the version tag, then the same key=value lines a meta +# uses, restricted to the minimum needed to land a released task's pull request. +# The tag is first so the file is never mistaken for a meta, and every remaining +# line sits where fm_pr_metadata_identity_parse already accepts it: project= +# before pr=, and only pr_head= after it. +fm_pr_landing_record_valid() { + local file=$1 first + [ -f "$file" ] && [ ! -L "$file" ] || return 1 + [ "$(fm_pr_file_link_count "$file")" = 1 ] || return 1 + IFS= read -r first < "$file" || return 1 + [ "$first" = fm-landing-v1 ] || return 1 + fm_pr_metadata_identity_parse "$file" +} + +# Resolve a task's PR identity record. A live task's meta always wins: a landing +# record is written only once the task is released and must never shadow one. +# The two are validated differently on purpose. A meta's contents are the +# caller's existing contract - bin/fm-pr-check.sh resolves one before any pr= is +# recorded in it - so only its presence is decided here. A landing record has no +# other purpose, so a file at that name that is not a well-formed landing record +# is refused rather than resolved. +fm_pr_identity_record_path() { + local state=$1 id=$2 meta landing + fm_pr_task_id_valid "$id" || return 1 + meta="$state/$id.meta" + landing="$state/$id.landing" + if [ -e "$meta" ] || [ -L "$meta" ]; then + printf '%s\n' "$meta" + return 0 + fi + fm_pr_landing_record_valid "$landing" || return 1 + printf '%s\n' "$landing" +} + +# Write a task's landing record from forge-derived values, replacing any earlier +# one atomically. The head is optional because only some forge CLIs report it, +# and every consumer treats it as evidence rather than as merge authority. +fm_pr_landing_record_write() { + local state=$1 id=$2 url=$3 head=${4-} project=${5-} + local dest tmp state_device canonical_url prev_umask + fm_pr_task_id_valid "$id" || return 1 + fm_pr_url_parse "$url" || return 1 + canonical_url=$FM_PR_URL + [ -z "$head" ] || fm_pr_head_valid "$head" || return 1 + case "$project" in *[![:print:]]*) return 1 ;; esac + [ -d "$state" ] && [ ! -L "$state" ] || return 1 + state_device=$(fm_pr_file_device "$state") || return 1 + dest="$state/$id.landing" + fm_pr_regular_destination_on_device_or_absent "$dest" "$state_device" || return 1 + # The private mode is asserted explicitly below; this only keeps the window + # between creation and that assertion closed, and it is restored immediately + # so callers such as bin/fm-teardown.sh keep their own umask afterwards. + prev_umask=$(umask) + umask 077 + tmp=$(mktemp "$state/.fm-pr-landing.XXXXXX") || { umask "$prev_umask"; return 1; } + umask "$prev_umask" + { + printf '%s\n' fm-landing-v1 + [ -z "$project" ] || printf 'project=%s\n' "$project" + printf 'pr=%s\n' "$canonical_url" + [ -z "$head" ] || printf 'pr_head=%s\n' "$head" + } > "$tmp" || { rm -f -- "$tmp"; return 1; } + if ! chmod 0600 "$tmp" \ + || ! fm_pr_private_file_valid "$tmp" 600 "$state_device" \ + || ! fm_pr_landing_record_valid "$tmp" \ + || [ "$FM_PR_META_URL" != "$canonical_url" ] \ + || ! fm_pr_regular_destination_on_device_or_absent "$dest" "$state_device" \ + || ! mv -f -- "$tmp" "$dest"; then + rm -f -- "$tmp" + return 1 + fi + fm_pr_private_file_valid "$dest" 600 "$state_device" || return 1 + fm_pr_landing_record_valid "$dest" || return 1 + [ "$FM_PR_META_URL" = "$canonical_url" ] +} + +# Ask the forge for its own view of one pull or merge request, so landing +# identity is read from the request rather than asserted by the caller. Returns +# non-zero whenever the forge cannot answer; a caller must treat that as "not +# known", never as evidence about whether the request merged. Each provider is +# read through the same CLI and output shape bin/fm-pr-poll.sh already relies on. +# FM_PR_FORGE_STATE is normalised to open, merged, or closed, and +# FM_PR_FORGE_HEAD is set only when the CLI reports a valid head. +# This re-parses the URL, so it overwrites every FM_PR_* parse global. +fm_pr_forge_view() { + local url=$1 out state head number host path + FM_PR_FORGE_STATE= + FM_PR_FORGE_HEAD= + fm_pr_url_parse "$url" || return 1 + url=$FM_PR_URL + number=$FM_PR_NUMBER + host=$FM_PR_HOST + path=$FM_PR_PATH + case "$FM_PR_PROVIDER" in + github) + command -v gh >/dev/null 2>&1 || return 1 + # gh resolves a full pull request URL without a repository around it, so + # this runs from / rather than from a worktree that may already be gone. + out=$(cd / && gh pr view "$url" --json state,headRefOid \ + -q '.state + "\t" + .headRefOid' 2>/dev/null) || return 1 + state=${out%%$'\t'*} + head=${out#*$'\t'} + [ "$state" != "$out" ] || return 1 + case "$state" in + OPEN) state=open ;; + MERGED) state=merged ;; + CLOSED) state=closed ;; + *) return 1 ;; + esac + # Consumed by bin/fm-pr-check.sh, bin/fm-pr-merge.sh, and bin/fm-teardown.sh. + # shellcheck disable=SC2034 + fm_pr_head_valid "$head" && FM_PR_FORGE_HEAD=$head + ;; + gitlab) + command -v glab >/dev/null 2>&1 || return 1 + # glab cannot take a merge request URL, so the instance comes from the + # validated record through -R, exactly as the merge poll addresses it. + out=$(cd / && glab mr view "$number" -R "https://$host/$path" 2>/dev/null) || return 1 + state=$(printf '%s\n' "$out" | sed -n 's/^state:[[:space:]]*//p' | head -1) + case "$state" in + open|opened) state=open ;; + merged) state=merged ;; + closed) state=closed ;; + *) return 1 ;; + esac + ;; + *) return 1 ;; + esac + # Consumed by bin/fm-pr-check.sh, bin/fm-pr-merge.sh, and bin/fm-teardown.sh. + # shellcheck disable=SC2034 + FM_PR_FORGE_STATE=$state +} + # Sidecar layout: provider, url, host, path, number, one per line. A sidecar # written before the provider tag existed has a URL on its first line and one # line fewer, so it fails both the field count and the provider comparison and @@ -580,19 +720,19 @@ fm_pr_poll_publish_prepared() { } fm_pr_poll_artifacts_valid() { - local state=$1 id=$2 template=$3 state_device check data registration meta data_hash template_hash data_identity check_identity + local state=$1 id=$2 template=$3 state_device check data registration record data_hash template_hash data_identity check_identity fm_pr_task_id_valid "$id" || return 1 [ -d "$state" ] && [ ! -L "$state" ] || return 1 state_device=$(fm_pr_file_device "$state") || return 1 check="$state/$id.check.sh" data="$state/$id.pr-poll" registration="$state/$id.pr-poll-registration" - meta="$state/$id.meta" + record=$(fm_pr_identity_record_path "$state" "$id") || return 1 fm_pr_private_file_valid "$check" 600 "$state_device" || return 1 fm_pr_private_file_valid "$data" 600 "$state_device" || return 1 fm_pr_private_file_valid "$registration" 600 "$state_device" || return 1 - [ -f "$meta" ] && [ ! -L "$meta" ] || return 1 - [ "$(fm_pr_file_link_count "$meta")" = 1 ] || return 1 + [ -f "$record" ] && [ ! -L "$record" ] || return 1 + [ "$(fm_pr_file_link_count "$record")" = 1 ] || return 1 cmp -s "$template" "$check" || return 1 fm_pr_poll_data_parse "$data" || return 1 data_hash=$(fm_pr_sha256 "$data") || return 1 @@ -610,7 +750,7 @@ fm_pr_poll_artifacts_valid() { [ "$FM_PR_REG_TEMPLATE_HASH" = "$template_hash" ] || return 1 [ "$FM_PR_REG_DATA_IDENTITY" = "$data_identity" ] || return 1 [ "$FM_PR_REG_CHECK_IDENTITY" = "$check_identity" ] || return 1 - fm_pr_metadata_identity_parse "$meta" || return 1 + fm_pr_metadata_identity_parse "$record" || return 1 [ "$FM_PR_META_PROVIDER" = "$FM_PR_DATA_PROVIDER" ] || return 1 [ "$FM_PR_META_URL" = "$FM_PR_DATA_URL" ] || return 1 [ "$FM_PR_META_HOST" = "$FM_PR_DATA_HOST" ] || return 1 @@ -721,7 +861,7 @@ fm_pr_poll_retirement_parse() { } fm_pr_poll_retirement_receipt_valid() { - local state=$1 id=$2 receipt state_device meta + local state=$1 id=$2 receipt state_device record fm_pr_task_id_valid "$id" || return 1 [ -d "$state" ] && [ ! -L "$state" ] || return 1 state_device=$(fm_pr_file_device "$state") || return 1 @@ -729,8 +869,8 @@ fm_pr_poll_retirement_receipt_valid() { fm_pr_private_file_valid "$receipt" 600 "$state_device" || return 1 fm_pr_poll_retirement_parse "$receipt" || return 1 [ "$FM_PR_RETIRE_ID" = "$id" ] || return 1 - meta="$state/$id.meta" - fm_pr_metadata_identity_parse "$meta" || return 1 + record=$(fm_pr_identity_record_path "$state" "$id") || return 1 + fm_pr_metadata_identity_parse "$record" || return 1 [ "$FM_PR_META_PROVIDER" = "$FM_PR_RETIRE_PROVIDER" ] || return 1 [ "$FM_PR_META_URL" = "$FM_PR_RETIRE_URL" ] || return 1 [ "$FM_PR_META_HOST" = "$FM_PR_RETIRE_HOST" ] || return 1 diff --git a/bin/fm-pr-merge.sh b/bin/fm-pr-merge.sh index 8226798a673..efc3e9c0945 100755 --- a/bin/fm-pr-merge.sh +++ b/bin/fm-pr-merge.sh @@ -1,13 +1,63 @@ #!/usr/bin/env bash -# Merge a task's PR after recording pr= and any available pr_head= through -# bin/fm-pr-check.sh, so teardown can verify landed work after squash merges. -# The full canonical GitHub PR URL is parsed by bin/fm-pr-lib.sh and the derived -# owner/repository and PR number are passed to gh-axi as separate arguments. +# Merge a task's PR after re-verifying the pull request's current head, then +# record pr= and any available pr_head= through bin/fm-pr-check.sh, so teardown +# can verify landed work after squash merges. The full canonical GitHub PR URL +# is parsed by bin/fm-pr-lib.sh and the derived owner/repository and PR number +# are passed to gh-axi as separate arguments. +# +# Verification re-reads the pull request rather than trusting any recorded +# value, because a PR can go red between an earlier check and the merge and +# state/.meta may carry a stale pr_head=. An early read refuses without +# recording the PR or arming its poll, then a final authoritative read runs after +# fm-pr-check.sh and immediately before the verification metadata write and +# merge. Each `gh pr view` call reads the head, mergeability, review decision, +# and check rollup together, so every state-based refusal names the exact head it +# evaluated once GitHub has supplied a readable head. +# The merge is refused when: +# * no check runs exist on that head - an empty rollup is never read as green, +# which is the whole point of this guard: a cross-repo fork PR held at +# action_required dispatches zero workflows and reports zero failures. The +# refusal names why the set is empty, separating a head with no CI +# configured from one whose workflows are held awaiting approval; +# * any check run is not SUCCESS - a queued, in-progress, skipped, neutral, +# cancelled, or failed run all refuse, so the guard fails closed on anything +# that is not an observed pass. Runs that returned an adverse verdict and +# runs that returned no verdict are counted and reported separately, so a +# head nothing examined is never described as a head something rejected; +# * the pull request is not MERGEABLE - CONFLICTING and a not-yet-computed +# UNKNOWN both refuse; +# * a review requests changes. +# +# --allow-unverified is the captain's explicit override. It is never a default +# and never inferred from the environment: it skips verification entirely and +# records merge_verification=override in the task's meta so an unverified merge +# stays visible afterwards. A verified merge records merge_verification=verified +# and merge_verified_head=; absence of both keys means unknown, never +# verified. Both keys are written before pr= so the metadata identity contract in +# bin/fm-pr-lib.sh still parses. The flag is recognised only before the optional +# -- separator; after it, it is forwarded to gh-axi, which rejects it. +# +# The final verification is not atomically bound to the merge. It narrows the +# remaining race window to the verification metadata write, but a head can still +# change before the merge. Closing that race requires a server-side head +# precondition under decision +# pipeline-reports-green-on-absent-ci-decision-merge-atomic-binding. The real +# `gh pr merge` supports `--match-head-commit SHA`, but gh-axi constructs its gh +# arguments from a fixed allowlist of the method, --auto, --delete-branch, --body, +# and --subject and silently drops other flags. Adopting the precondition later +# therefore requires changing the single gh-axi invocation at the end. +# +# A task released before its pull request lands keeps a durable landing record +# instead of a meta, and this path lands it through that record. A task released +# before landing records existed keeps neither, so its record is rebuilt from a +# forge read of the request. An existing landing record must be valid and name +# the requested URL before any forge read or merge. Either way the request must +# still be open at its forge, and a request that resolves to nothing is refused. # # Merge method defaults to --squash when the caller passes none of --squash, # --merge, --rebase, or --method after the optional -- separator. Extra args # must not include --repo or -R because the repository comes only from the URL. -# Usage: fm-pr-merge.sh [-- ] +# Usage: fm-pr-merge.sh [--allow-unverified] [-- ] set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -37,7 +87,15 @@ PR_OWNER=$FM_PR_OWNER PR_REPO=$FM_PR_REPO PR_NUMBER=$FM_PR_NUMBER shift 2 -[ "${1:-}" = "--" ] && shift + +ALLOW_UNVERIFIED=0 +while [ "$#" -gt 0 ]; do + case "$1" in + --allow-unverified) ALLOW_UNVERIFIED=1; shift ;; + --) shift; break ;; + *) break ;; + esac +done caller_has_merge_method() { local arg @@ -64,21 +122,298 @@ reject_repo_overrides() { reject_repo_overrides "$@" || exit 1 # Task-derived paths are constructed only after the canonical ID validation. -META="$STATE/$ID.meta" -if [ ! -f "$META" ] || [ -L "$META" ]; then - echo "error: task metadata is unavailable" >&2 - exit 1 +# Landing identity comes from whichever durable record the task still has, and +# from the pull request itself when it has none. A live task keeps its meta and +# takes the unchanged path below. A released task keeps only the landing record +# bin/fm-teardown.sh left behind, and a task released before landing records +# existed keeps neither: its pull request is then the authority for its own +# identity and the record is rebuilt from a forge read, never from the caller. +LANDING="$STATE/$ID.landing" +REBUILD=0 +if ! RECORD=$(fm_pr_identity_record_path "$STATE" "$ID"); then + if [ -e "$LANDING" ] || [ -L "$LANDING" ]; then + echo "error: task landing record is invalid" >&2 + exit 1 + fi + RECORD=$LANDING + REBUILD=1 +fi + +# One read of the live pull request, so the head reported in a refusal is the +# same head the checks, mergeability, and review decision were read from. +PR_VERIFY_FIELDS=headRefOid,mergeable,reviewDecision,statusCheckRollup +# A CheckRun carries .conclusion (empty while queued or running); a legacy +# StatusContext carries .state instead. Neither is treated as a pass unless it +# says SUCCESS. +# +# The members are counted in three disjoint buckets, not two, because "ran and +# reported a failure" and "never produced a result" are different facts about a +# head and collapsing them loses the one this guard exists to report. A run that +# failed, errored, timed out, or failed to start returned an adverse verdict; a +# run that is queued, in progress, skipped, neutral, cancelled, stale, or held +# at action_required returned no verdict at all. Both refuse, and each says so +# in its own words. PR_VERIFY_FAILING is the whole adverse set, so anything +# absent from it that is not SUCCESS counts as unrun rather than as a failure. +PR_VERIFY_FAILING='["FAILURE","ERROR","TIMED_OUT","STARTUP_FAILURE"]' +# $s below is jq's own binding, not a shell variable; only the interpolated +# PR_VERIFY_FAILING array is expanded by the shell. +# shellcheck disable=SC2016 +PR_VERIFY_QUERY='"head=\(.headRefOid // "")", +"mergeable=\(.mergeable // "")", +"review=\(.reviewDecision // "")", +"checks=\((.statusCheckRollup // []) | length)", +"unsuccessful=\((.statusCheckRollup // []) | map(select(((.conclusion // .state // "") | ascii_upcase) != "SUCCESS")) | length)", +"failing=\((.statusCheckRollup // []) | map(select(((.conclusion // .state // "") | ascii_upcase) as $s | ('"$PR_VERIFY_FAILING"' | index($s)) != null)) | length)", +"unrun=\((.statusCheckRollup // []) | map(select(((.conclusion // .state // "") | ascii_upcase) as $s | $s != "SUCCESS" and ('"$PR_VERIFY_FAILING"' | index($s)) == null)) | length)"' + +VERIFIED_HEAD= + +# An empty rollup has more than one cause, and the two common ones need +# different work from the captain: a repository with no CI configured for this +# head, and a cross-repo fork pull request whose workflows exist but are held at +# action_required until a maintainer approves them. GitHub reports neither as a +# check run, so both arrive here as the same empty list, but the check-suite +# read below separates them. This only enriches an already-decided refusal: it +# runs on the refusal path alone, reports nothing when it cannot read the +# suites, and can never turn a refusal into a merge. +empty_rollup_evidence() { + local head=$1 counts total held extra + command -v gh >/dev/null 2>&1 || return 0 + counts=$(gh api "repos/$PR_OWNER/$PR_REPO/commits/$head/check-suites" \ + -q '"\(.total_count // 0) \([.check_suites[]? | select(((.conclusion // "") | ascii_downcase) == "action_required")] | length)"' \ + 2>/dev/null) || return 0 + # Exactly two whole numbers, or this response was not the one asked for and + # the refusal stands with no added detail rather than an invented one. + read -r total held extra <<< "$counts" || return 0 + [ -z "$extra" ] || return 0 + [ -n "$total" ] && [ -z "${total//[0-9]/}" ] || return 0 + [ -n "$held" ] && [ -z "${held//[0-9]/}" ] || return 0 + if [ "$held" -gt 0 ]; then + printf ' (%s check suite(s) on it are held at action_required, so its workflows are waiting on a maintainer to approve them and will not run on their own)' \ + "$held" + elif [ "$total" -eq 0 ]; then + printf ' (no check suite exists for it either, so no CI is configured to run on this head)' + fi + return 0 +} + +verify_current_head() { + local output line joined + local head='' mergeable='' review='' checks='' unsuccessful='' failing='' unrun='' + local -a reasons=() + + command -v gh >/dev/null 2>&1 || { + echo "error: refusing to merge: the pull request could not be verified because gh is not on PATH" >&2 + return 1 + } + output=$(gh pr view "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" \ + --json "$PR_VERIFY_FIELDS" -q "$PR_VERIFY_QUERY" 2>/dev/null) || { + echo "error: refusing to merge: the pull request state could not be read from GitHub" >&2 + return 1 + } + + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + head=*) head=${line#head=} ;; + mergeable=*) mergeable=${line#mergeable=} ;; + review=*) review=${line#review=} ;; + checks=*) checks=${line#checks=} ;; + unsuccessful=*) unsuccessful=${line#unsuccessful=} ;; + failing=*) failing=${line#failing=} ;; + unrun=*) unrun=${line#unrun=} ;; + esac + done <<< "$output" + + fm_pr_head_valid "$head" || { + echo "error: refusing to merge: the pull request head commit could not be read from GitHub" >&2 + return 1 + } + # Each count is validated on its own. Concatenating them would let one empty + # field hide behind the other's digits and reach the comparisons below as an + # empty string, which compares as neither zero nor positive and would merge. + if [ -z "$checks" ] || [ -n "${checks//[0-9]/}" ] \ + || [ -z "$unsuccessful" ] || [ -n "${unsuccessful//[0-9]/}" ] \ + || [ -z "$failing" ] || [ -n "${failing//[0-9]/}" ] \ + || [ -z "$unrun" ] || [ -n "${unrun//[0-9]/}" ]; then + printf 'error: refusing to merge head %s: the check rollup could not be read from GitHub\n' \ + "$head" >&2 + return 1 + fi + # The two disjoint buckets must account for exactly the members that are not + # successes. A response that breaks that identity was not understood, and an + # unreadable rollup is reported as unreadable rather than resolved either way. + if [ "$((failing + unrun))" -ne "$unsuccessful" ]; then + printf 'error: refusing to merge head %s: the check rollup could not be read from GitHub\n' \ + "$head" >&2 + return 1 + fi + + [ "$mergeable" = MERGEABLE ] \ + || reasons+=("the pull request is not mergeable (mergeable=${mergeable:-unreported})") + [ "$review" != CHANGES_REQUESTED ] || reasons+=("a review requests changes") + # Zero check runs and all-successful check runs both report zero failures, so + # the empty rollup is refused on its own count and never folded into the + # counts below. A non-empty rollup reports its failed and its unrun members + # separately, so "this was examined and found broken" never reaches the + # captain wearing the words of "this was never examined", or the reverse. + if [ "$checks" -eq 0 ]; then + reasons+=("no check runs exist on this head$(empty_rollup_evidence "$head")") + else + [ "$failing" -eq 0 ] \ + || reasons+=("$failing of $checks check runs failed") + [ "$unrun" -eq 0 ] \ + || reasons+=("$unrun of $checks check runs reported no result (queued, in progress, skipped, neutral, cancelled, or held for approval)") + fi + + if [ "${#reasons[@]}" -gt 0 ]; then + joined=$(printf '%s; ' "${reasons[@]}") + printf 'error: refusing to merge head %s: %s\n' "$head" "${joined%; }" >&2 + return 1 + fi + VERIFIED_HEAD=$head +} + +MERGE_META_TMP= +merge_meta_cleanup() { + [ -z "$MERGE_META_TMP" ] || rm -f -- "$MERGE_META_TMP" + MERGE_META_TMP= +} +trap merge_meta_cleanup EXIT +trap 'exit 1' HUP INT TERM + +# Record how this merge was authorised. The two keys are emitted before any +# pr=/pr_head= lines so fm_pr_metadata_identity_parse, which refuses any unknown +# key after pr=, still accepts the file at every instant. +record_merge_verification() { + local status=$1 head=$2 line state_device meta_device + state_device=$(fm_pr_file_device "$STATE") || return 1 + meta_device=$(fm_pr_file_device "$META") || return 1 + [ "$meta_device" = "$state_device" ] || return 1 + MERGE_META_TMP=$(mktemp "$STATE/.fm-pr-merge-meta.XXXXXX") || return 1 + { + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + merge_verification=*|merge_verified_head=*|pr=*|pr_head=*) ;; + *) printf '%s\n' "$line" ;; + esac + done < "$META" + printf 'merge_verification=%s\n' "$status" + [ -z "$head" ] || printf 'merge_verified_head=%s\n' "$head" + while IFS= read -r line || [ -n "$line" ]; do + case "$line" in + pr=*|pr_head=*) printf '%s\n' "$line" ;; + esac + done < "$META" + } > "$MERGE_META_TMP" || return 1 + chmod 0600 "$MERGE_META_TMP" || return 1 + fm_pr_private_file_valid "$MERGE_META_TMP" 600 "$state_device" || return 1 + fm_pr_regular_destination_on_device_or_absent "$META" "$state_device" || return 1 + mv -f -- "$MERGE_META_TMP" "$META" || return 1 + MERGE_META_TMP= +} + +# This fork carries a released task's landing record as well as a live task's +# meta, so the record is resolved in three ordered stages rather than one. Its +# own refusals - a torn-down meta, an invalid landing record, a record naming +# another request - come first and cost no forge read, exactly as before. +# Verification runs next, so nothing is recorded and no poll is armed for a head +# it refuses. Only then does the task's identity get written. +LIVE_TASK=0 +if [ "$RECORD" != "$LANDING" ]; then + if [ ! -f "$RECORD" ] || [ -L "$RECORD" ]; then + echo "error: task metadata is unavailable" >&2 + exit 1 + fi + # Recording and arming are deferred until after verification below, so a head + # the guard refuses leaves the task with no pr= and no armed poll. + LIVE_TASK=1 +else + if [ "$REBUILD" = 0 ]; then + fm_pr_metadata_identity_parse "$RECORD" || { + echo "error: task landing record is invalid" >&2 + exit 1 + } + RECORD_URL=$FM_PR_META_URL + if [ "$RECORD_URL" != "$URL" ]; then + echo "error: task landing record names $RECORD_URL, but $URL was requested" >&2 + exit 1 + fi + fi + # A released task has no worktree, no worker, and nothing left to tear down, + # and the merge below is synchronous, so no merge poll is armed for it. The + # forge decides here instead: the request must still be open, and the landing + # record is then written from what the forge reported rather than from a stale + # local value or anything the caller asserted. + RECORD_PROJECT= + [ "$REBUILD" = 1 ] \ + || RECORD_PROJECT=$(grep '^project=' "$RECORD" | tail -1 | cut -d= -f2- || true) + if ! fm_pr_forge_view "$URL"; then + if [ "$REBUILD" = 1 ]; then + echo "error: task metadata is unavailable" >&2 + echo "No record for task $ID, and $URL could not be resolved at its forge." >&2 + else + echo "error: $URL could not be resolved at its forge" >&2 + fi + exit 1 + fi + if [ "$FM_PR_FORGE_STATE" != open ]; then + if [ "$REBUILD" = 1 ]; then + echo "error: task metadata is unavailable" >&2 + echo "No record for task $ID, and $URL is $FM_PR_FORGE_STATE at its forge rather than an open pull request." >&2 + else + echo "error: $URL is $FM_PR_FORGE_STATE at its forge rather than an open pull request" >&2 + fi + exit 1 + fi + fm_pr_landing_record_write "$STATE" "$ID" "$URL" "$FM_PR_FORGE_HEAD" "$RECORD_PROJECT" || { + echo "error: task landing record could not be written" >&2 + exit 1 + } + [ "$REBUILD" = 0 ] || printf 'rebuilt: state/%s.landing from %s\n' "$ID" "$URL" +fi + +if [ "$ALLOW_UNVERIFIED" -ne 1 ]; then + verify_current_head || exit 1 fi -"$SCRIPT_DIR/fm-pr-check.sh" "$ID" "$URL" -grep -qxF "pr=$URL" "$META" || { +[ "$LIVE_TASK" = 0 ] || "$SCRIPT_DIR/fm-pr-check.sh" "$ID" "$URL" +META=$RECORD +grep -qxF "pr=$URL" "$RECORD" || { echo "error: PR metadata recording failed" >&2 exit 1 } +if [ "$ALLOW_UNVERIFIED" -eq 1 ]; then + MERGE_VERIFICATION=override + VERIFIED_HEAD= +else + VERIFIED_HEAD= + verify_current_head || exit 1 + MERGE_VERIFICATION=verified +fi + +record_merge_verification "$MERGE_VERIFICATION" "$VERIFIED_HEAD" || { + echo "error: merge verification metadata could not be recorded" >&2 + exit 1 +} +grep -qxF "merge_verification=$MERGE_VERIFICATION" "$META" || { + echo "error: merge verification metadata could not be recorded" >&2 + exit 1 +} + merge_args=() if ! caller_has_merge_method "$@"; then merge_args=(--squash) fi gh-axi pr merge "$PR_NUMBER" --repo "$PR_OWNER/$PR_REPO" "${merge_args[@]+"${merge_args[@]}"}" "$@" + +# The landing record exists only so a released task's pull request can still be +# landed here, so it is spent once that merge succeeds. It is retained while a +# merge poll is still armed against it, because the poll's registration binds to +# it and the watcher would otherwise reject an authentic check. +if [ "$RECORD" = "$LANDING" ] && [ -f "$LANDING" ] && [ ! -L "$LANDING" ] \ + && [ ! -e "$STATE/$ID.check.sh" ] && [ ! -L "$STATE/$ID.check.sh" ] \ + && [ ! -e "$STATE/$ID.pr-poll" ] && [ ! -L "$STATE/$ID.pr-poll" ]; then + rm -f -- "$LANDING" || true +fi diff --git a/bin/fm-pr-poll.sh b/bin/fm-pr-poll.sh index ed705ce7073..0339bdf97b3 100755 --- a/bin/fm-pr-poll.sh +++ b/bin/fm-pr-poll.sh @@ -1,9 +1,11 @@ #!/usr/bin/env bash # Static watcher program for a validated PR/MR poll sidecar. -# It emits exactly one merged line for a merged PR or MR and stays silent -# otherwise, including on every error, so a failed lookup can never be read as -# a merge. The provider-tagged identity is data in the sidecar and is never -# interpolated into this source: these bytes are identical for every task. +# It emits exactly one merged line for a merged PR or MR, exactly one +# "dirty " line for an open GitHub pull request the forge reports as +# conflicting, and stays silent otherwise, including on every error, so a +# failed lookup can never be read as either result. The provider-tagged +# identity is data in the sidecar and is never interpolated into this source: +# these bytes are identical for every task. # Each provider is read through its own standard CLI, gh for GitHub and glab # for GitLab, so an upstream checkout needs no extra tooling to follow either. set -u @@ -62,8 +64,47 @@ case "$provider" in .|..|*[!A-Za-z0-9._-]*) exit 0 ;; esac [ "$url" = "https://github.com/$owner/$repo/pull/$number" ] || exit 0 - state=$(gh pr view "$url" --json state -q .state 2>/dev/null) || exit 0 - [ "$state" = MERGED ] && printf '%s\n' merged + # One request carries all three fields, so conflict detection adds no call + # to the merge poll. The stable GraphQL "mergeable" field is read rather + # than mergeStateStatus: the two agree on a conflict, but mergeStateStatus + # rides a preview header and is likelier to be absent on GitHub Enterprise + # Server, and its non-conflict values also report failing or pending checks, + # which are a different signal this poll must stay silent about. + # A base push makes GitHub recompute mergeability on its own, so by the time + # this poll runs the answer is normally settled; the transient UNKNOWN it can + # return meanwhile is silence here and resolves on the next poll rather than + # costing this static program a retry and a timing dependency. + raw=$(gh pr view "$url" --json state,mergeable,headRefOid \ + -q '[.state,.mergeable,.headRefOid]|@tsv' 2>/dev/null) || exit 0 + tab=$(printf '\t') + # Exactly three tab-separated fields, revalidated below against exact + # allowlists rather than trusted, so unexpected output stays silent. + case "$raw" in + *"$tab"*"$tab"*) ;; + *) exit 0 ;; + esac + state=${raw%%"$tab"*} + rest=${raw#*"$tab"} + mergeable=${rest%%"$tab"*} + head=${rest#*"$tab"} + case "$head" in + *"$tab"*) exit 0 ;; + esac + # Merged is decided first and alone, so a merged or closed pull request + # keeps the exact result it had before conflict reporting existed. + if [ "$state" = MERGED ]; then + printf '%s\n' merged + exit 0 + fi + [ "$state" = OPEN ] || exit 0 + [ "$mergeable" = CONFLICTING ] || exit 0 + # The head commit identifies the conflict episode for the watcher's dedupe. + # An unreadable head is reported as such instead of suppressing the wake. + case "$head" in + ''|*[!0-9a-f]*) head=unknown ;; + *) [ "${#head}" -ge 7 ] && [ "${#head}" -le 64 ] || head=unknown ;; + esac + printf 'dirty %s\n' "$head" ;; gitlab) [ "${#host}" -ge 1 ] && [ "${#host}" -le 253 ] || exit 0 diff --git a/bin/fm-procevent-remote-reply.sh b/bin/fm-procevent-remote-reply.sh new file mode 100755 index 00000000000..7518178a400 --- /dev/null +++ b/bin/fm-procevent-remote-reply.sh @@ -0,0 +1,457 @@ +#!/usr/bin/env bash +# Remote-secondmate reply adapter for the generic process-event runner. +# +# Usage: +# fm-procevent-remote-reply.sh arm +# fm-procevent-remote-reply.sh handle +# fm-procevent-remote-reply.sh classify +# fm-procevent-remote-reply.sh terminal +# fm-procevent-remote-reply.sh source-id +# fm-procevent-remote-reply.sh retire +# +# `arm` registers one blocking, non-destructive delta source for the remote +# home's state/parent-replies.status log. The process-event runner owns blocking, +# capture, publication, and one machine-wide source owner. Each captured delta is +# terminal for that exact registration; `handle` validates and idempotently +# ingests it, acknowledges the captured generation, then registers the next +# cursor-anchored source. A continuity break is escalated and not re-armed. +# +# Ingest accepts only bounded, printable status lines with an allowed lifecycle +# verb and corr=<16hex>. Exact lines are appended at most once to the parent's +# state/.status. A data/*.md pointer is fetched through the path-confined +# remote file reader and rewritten to its local private copy before append. +set -u + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" +FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" +STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" +DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" +CURSOR_DIR="$STATE/remote-replies" +REMOTE_LOG='state/parent-replies.status' +WAIT_SECONDS=${FM_REMOTE_REPLY_WAIT_SECONDS:-55} +MAX_LINE_BYTES=${FM_REMOTE_REPLY_MAX_LINE_BYTES:-2048} +MAX_DOC_BYTES=${FM_REMOTE_REPLY_MAX_DOC_BYTES:-262144} + +# shellcheck source=bin/fm-wake-lib.sh +. "$SCRIPT_DIR/fm-wake-lib.sh" +# shellcheck source=bin/fm-secondmate-registry-lib.sh +. "$SCRIPT_DIR/fm-secondmate-registry-lib.sh" +# shellcheck source=bin/fm-pending-reply-lib.sh +. "$SCRIPT_DIR/fm-pending-reply-lib.sh" + +die() { printf 'error: %s\n' "$1" >&2; exit 1; } +usage() { sed -n '2,22p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } + +sha256_file() { + if command -v shasum >/dev/null 2>&1; then + shasum -a 256 "$1" | awk '{print $1}' + elif command -v sha256sum >/dev/null 2>&1; then + sha256sum "$1" | awk '{print $1}' + else + die "no SHA-256 tool is available" + fi +} + +empty_hash() { + local tmp + tmp=$(mktemp "${TMPDIR:-/tmp}/fm-empty-hash.XXXXXX") || return 1 + : > "$tmp" + sha256_file "$tmp" + rm -f -- "$tmp" +} + +validate_id() { + case "$1" in ''|*[!A-Za-z0-9._-]*) die "invalid secondmate id: $1" ;; esac +} + +source_id() { + validate_id "$1" + printf 'remote-reply-%s\n' "$1" +} + +cursor_path() { printf '%s/%s.cursor\n' "$CURSOR_DIR" "$1"; } +ingest_receipt_path() { printf '%s/%s.%s.ingested\n' "$CURSOR_DIR" "$1" "$2"; } + +read_cursor() { # ; sets CURSOR_OFFSET and CURSOR_HASH + local path=$1 offset hash schema + path=$(cursor_path "$path") + CURSOR_OFFSET=0 + CURSOR_HASH=$(empty_hash) || die "cannot establish the empty cursor hash" + [ -e "$path" ] || return 0 + [ -f "$path" ] && [ ! -L "$path" ] || die "reply cursor is unsafe: $path" + schema=$(sed -n 's/^schema=//p' "$path") + offset=$(sed -n 's/^offset=//p' "$path") + hash=$(sed -n 's/^prefix_sha256=//p' "$path") + [ "$schema" = fm-remote-reply-cursor.v1 ] || die "reply cursor has an incompatible schema: $path" + case "$offset" in ''|*[!0-9]*) die "reply cursor has an invalid offset: $path" ;; esac + case "$hash" in *[!A-Fa-f0-9]*|'') die "reply cursor has an invalid hash: $path" ;; esac + [ "${#hash}" -eq 64 ] || die "reply cursor has an invalid hash length: $path" + CURSOR_OFFSET=$offset + CURSOR_HASH=$(printf '%s' "$hash" | tr 'A-F' 'a-f') +} + +write_cursor() { # + local id=$1 offset=$2 hash=$3 path tmp + mkdir -p "$CURSOR_DIR" || return 1 + chmod 700 "$CURSOR_DIR" 2>/dev/null || true + path=$(cursor_path "$id") + [ ! -L "$path" ] || return 1 + tmp=$(umask 077; mktemp "$CURSOR_DIR/.cursor.XXXXXX") || return 1 + { + printf 'schema=fm-remote-reply-cursor.v1\n' + printf 'offset=%s\n' "$offset" + printf 'prefix_sha256=%s\n' "$hash" + } > "$tmp" || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + mv -f -- "$tmp" "$path" +} + +ingest_receipt_matches() { # + local path stored actual count + path=$(ingest_receipt_path "$1" "$2") + [ -e "$path" ] || [ -L "$path" ] || return 1 + [ -f "$path" ] && [ ! -L "$path" ] || die "remote reply ingestion receipt is unsafe: $path" + count=$(grep -c '^result_sha256=' "$path" 2>/dev/null || true) + [ "$count" -eq 1 ] || die "remote reply ingestion receipt is malformed: $path" + stored=$(sed -n 's/^result_sha256=//p' "$path") + case "$stored" in *[!A-Fa-f0-9]*|'') die "remote reply ingestion receipt is malformed: $path" ;; esac + [ "${#stored}" -eq 64 ] || die "remote reply ingestion receipt is malformed: $path" + actual=$(sha256_file "$3") || die "cannot hash remote reply result" + [ "$stored" = "$actual" ] || die "remote reply generation conflicts with its ingestion receipt" +} + +write_ingest_receipt() { # + local id=$1 seq=$2 result=$3 path tmp hash + mkdir -p "$CURSOR_DIR" || return 1 + chmod 700 "$CURSOR_DIR" 2>/dev/null || true + path=$(ingest_receipt_path "$id" "$seq") + if [ -e "$path" ] || [ -L "$path" ]; then + ingest_receipt_matches "$id" "$seq" "$result" + return $? + fi + hash=$(sha256_file "$result") || return 1 + tmp=$(umask 077; mktemp "$CURSOR_DIR/.ingested.XXXXXX") || return 1 + printf 'result_sha256=%s\n' "$hash" > "$tmp" \ + || { rm -f -- "$tmp"; return 1; } + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + if ! mv -f -- "$tmp" "$path"; then + rm -f -- "$tmp" + return 1 + fi +} + +result_field() { # + local count + count=$(grep -c "^$2=" "$1" 2>/dev/null || true) + [ "$count" -eq 1 ] || return 1 + grep "^$2=" "$1" | cut -d= -f2- +} + +classify_result() { + local file=$1 schema status + [ -f "$file" ] && [ ! -L "$file" ] || { printf 'malformed\n'; return 0; } + schema=$(result_field "$file" schema 2>/dev/null || true) + status=$(result_field "$file" status 2>/dev/null || true) + [ "$schema" = fm-remote-delta.v1 ] || { printf 'malformed\n'; return 0; } + case "$status" in + delta) printf 'delta\n' ;; + continuity-broken) printf 'continuity-broken\n' ;; + *) printf 'malformed\n' ;; + esac +} + +remote_route_exists() { + local id=$1 remote + remote=$(secondmate_registry_field "$DATA/secondmates.md" "$id" remote 2>/dev/null || true) + [ "$remote" = 1 ] || die "secondmate $id is not a configured remote route" +} + +cmd_arm_locked() { + local id=${1:-} sid + validate_id "$id" + remote_route_exists "$id" + read_cursor "$id" + sid=$(source_id "$id") + "$SCRIPT_DIR/fm-procevent.sh" register remote-reply "$sid" -- \ + "$SCRIPT_DIR/fm-procevent-remote-reply.sh" source "$id" || return 1 + printf 'armed: %s offset=%s\n' "$sid" "$CURSOR_OFFSET" +} + +cmd_arm() { + local id=${1:-} lock + validate_id "$id" + lock=$(secondmate_reply_lifecycle_lock_path "$STATE" "$id") + ( + fm_lock_acquire_wait "$lock" || die "cannot lock remote reply lifecycle for $id" + trap 'fm_lock_release "$lock"' EXIT + cmd_arm_locked "$id" + ) +} + +cmd_source() { + local id=${1:-} + validate_id "$id" + read_cursor "$id" + exec "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-delta-read.sh \ + "$REMOTE_LOG" "$CURSOR_OFFSET" "$CURSOR_HASH" "$WAIT_SECONDS" < /dev/null +} + +safe_doc_path() { + case "$1" in + data/*.md) ;; + *) return 1 ;; + esac + case "/$1/" in */../*|*/./*) return 1 ;; esac + case "$1" in *'//'*) return 1 ;; esac + return 0 +} + +fetch_document() { # + local id=$1 rel=$2 result_var=$3 base destination parent parent_real tmp local_rel + safe_doc_path "$rel" || return 1 + base="$DATA/remote-secondmates/$id" + destination="$base/$rel" + parent=$(dirname "$destination") + mkdir -p "$parent" || return 1 + [ ! -L "$base" ] && [ ! -L "$parent" ] || return 1 + parent_real=$(CDPATH='' cd -- "$parent" 2>/dev/null && pwd -P) || return 1 + case "$parent_real" in "$base"|"$base"/*) ;; *) return 1 ;; esac + [ ! -L "$destination" ] || return 1 + tmp=$(umask 077; mktemp "$parent/.remote-doc.XXXXXX") || return 1 + if ! "$SCRIPT_DIR/fm-on.sh" "$id" fm-remote-file.sh get "$rel" "$MAX_DOC_BYTES" < /dev/null > "$tmp"; then + rm -f -- "$tmp" + return 1 + fi + chmod 600 "$tmp" || { rm -f -- "$tmp"; return 1; } + mv -f -- "$tmp" "$destination" || { rm -f -- "$tmp"; return 1; } + local_rel="data/remote-secondmates/$id/$rel" + printf -v "$result_var" '%s' "$local_rel" +} + +line_valid() { # + local line=$1 bytes + [ -n "$line" ] || return 1 + bytes=$(printf '%s' "$line" | LC_ALL=C wc -c | tr -d ' ') + [ "$bytes" -le "$MAX_LINE_BYTES" ] || return 1 + [ -z "$(printf '%s' "$line" | LC_ALL=C tr -d '\11\40-\176')" ] || return 1 + printf '%s' "$line" | grep -Eq '^(working|needs-decision|blocked|paused|done|failed|resolved)([[:space:]]+\[[^]]+\])?:' || return 1 + printf '%s' "$line" | grep -Eq 'corr=[A-Fa-f0-9]{16}' +} + +cmd_ingest() { + local id=${1:-} result=${2:-} seq=${3:-} class blank payload schema status path from to from_hash to_hash payload_hash payload_bytes reason + local actual_bytes actual_hash line doc local_doc rewritten appended=0 cursor_already=0 lock status_file tmp + validate_id "$id" + [ -f "$result" ] && [ ! -L "$result" ] || die "result file is unavailable or unsafe: $result" + class=$(classify_result "$result") + [ "$class" != malformed ] || die "remote reply result is malformed" + schema=$(result_field "$result" schema) || die "result schema is ambiguous" + status=$(result_field "$result" status) || die "result status is ambiguous" + path=$(result_field "$result" path) || die "result path is ambiguous" + from=$(result_field "$result" from_offset) || die "result start offset is ambiguous" + to=$(result_field "$result" to_offset) || die "result end offset is ambiguous" + from_hash=$(result_field "$result" from_prefix_sha256) || die "result start hash is ambiguous" + to_hash=$(result_field "$result" to_prefix_sha256) || die "result end hash is ambiguous" + payload_hash=$(result_field "$result" payload_sha256) || die "result payload hash is ambiguous" + payload_bytes=$(result_field "$result" payload_bytes) || die "result payload size is ambiguous" + reason=$(result_field "$result" reason) || die "result reason is ambiguous" + [ "$schema" = fm-remote-delta.v1 ] && [ "$path" = "$REMOTE_LOG" ] || die "result identifies the wrong source" + case "$from$to$payload_bytes" in *[!0-9]*) die "result carries a nonnumeric size or offset" ;; esac + for hash in "$from_hash" "$to_hash" "$payload_hash"; do + case "$hash" in *[!A-Fa-f0-9]*|'') die "result carries an invalid SHA-256 value" ;; esac + [ "${#hash}" -eq 64 ] || die "result carries an invalid SHA-256 length" + done + blank=$(grep -n -m 1 '^$' "$result" | cut -d: -f1) + case "$blank" in ''|*[!0-9]*) die "result has no payload boundary" ;; esac + tmp=$(mktemp -d "${TMPDIR:-/tmp}/fm-remote-reply-ingest.XXXXXX") || die "cannot create ingest staging directory" + trap 'rm -rf -- "$tmp"' EXIT + payload="$tmp/payload" + tail -n "+$((blank + 1))" "$result" > "$payload" + actual_bytes=$(LC_ALL=C wc -c < "$payload" | tr -d ' ') + actual_hash=$(sha256_file "$payload") + [ "$actual_bytes" -eq "$payload_bytes" ] && [ "$actual_hash" = "$payload_hash" ] \ + || die "result payload bytes do not match its committed digest" + status_file="$STATE/$id.status" + mkdir -p "$STATE" || die "cannot create parent state directory" + [ ! -L "$status_file" ] || die "parent status log is a symlink" + lock="$STATE/.remote-reply-ingest-$id.lock" + fm_lock_acquire_wait "$lock" || die "cannot lock remote reply ingest for $id" + read_cursor "$id" + if [ "$CURSOR_OFFSET" -eq "$to" ] && [ "$CURSOR_HASH" = "$to_hash" ]; then + cursor_already=1 + elif [ "$CURSOR_OFFSET" -ne "$from" ] || [ "$CURSOR_HASH" != "$from_hash" ]; then + die "result does not continue the current cursor for $id" + fi + if [ "$class" = continuity-broken ]; then + line="blocked [key=remote-reply-continuity-$id]: remote reply continuity broke for $id ($reason)" + if ! grep -Fqx -- "$line" "$status_file" 2>/dev/null; then + printf '%s\n' "$line" >> "$status_file" || { fm_lock_release "$lock"; die "cannot append continuity escalation"; } + fi + fm_lock_release "$lock" + printf 'continuity-broken: %s (%s)\n' "$id" "$reason" + return 3 + fi + [ "$status" = delta ] && [ "$payload_bytes" -gt 0 ] || { fm_lock_release "$lock"; die "delta result has no payload"; } + while IFS= read -r line || [ -n "$line" ]; do + line_valid "$line" || { fm_lock_release "$lock"; die "delta contains an invalid or uncorrelated status line"; } + rewritten=$line + while IFS= read -r doc; do + [ -n "$doc" ] || continue + fetch_document "$id" "$doc" local_doc || { fm_lock_release "$lock"; die "could not fetch referenced remote document: $doc"; } + rewritten=${rewritten//"$doc"/"$local_doc"} + done < <(printf '%s\n' "$line" | grep -Eo 'data/[A-Za-z0-9._/-]+\.md' | awk '!seen[$0]++') + if ! grep -Fqx -- "$rewritten" "$status_file" 2>/dev/null; then + printf '%s\n' "$rewritten" >> "$status_file" || { fm_lock_release "$lock"; die "cannot append remote reply"; } + appended=$((appended + 1)) + fi + done < "$payload" + while IFS= read -r corr; do + [ -n "$corr" ] || continue + fm_pending_reply_try_resolve "$STATE" "$corr" "$status_file" >/dev/null 2>&1 || true + done < <(grep -Eo 'corr=[A-Fa-f0-9]{16}' "$payload" | cut -d= -f2- | tr 'A-F' 'a-f' | awk '!seen[$0]++') + if [ -n "$seq" ]; then + write_ingest_receipt "$id" "$seq" "$result" \ + || { fm_lock_release "$lock"; die "cannot commit remote reply ingestion receipt"; } + fi + if [ "$cursor_already" -eq 0 ]; then + write_cursor "$id" "$to" "$to_hash" || { fm_lock_release "$lock"; die "cannot commit remote reply cursor"; } + fi + fm_lock_release "$lock" + trap - EXIT + rm -rf -- "$tmp" + printf 'ingested: %s appended=%s offset=%s\n' "$id" "$appended" "$to" +} + +cmd_handle_locked() { + local id=${1:-} seq=${2:-} result=${3:-} sid class rc=0 to + validate_id "$id" + case "$seq" in ''|*[!0-9]*) die "sequence must be a nonnegative integer" ;; esac + sid=$(source_id "$id") + class=$(classify_result "$result") + [ "$class" != malformed ] || die "remote reply result is malformed" + if ingest_receipt_matches "$id" "$seq" "$result"; then + to=$(result_field "$result" to_offset) || die "result end offset is ambiguous" + printf 'ingested: %s appended=0 offset=%s\n' "$id" "$to" + else + cmd_ingest "$id" "$result" "$seq" || rc=$? + fi + if [ "$rc" -ne 0 ] && [ "$rc" -ne 3 ]; then + return "$rc" + fi + if [ "$class" = delta ]; then + cmd_arm_locked "$id" || return 1 + fi + "$SCRIPT_DIR/fm-procevent.sh" handled "$sid" "$seq" || return 1 + return "$rc" +} + +cmd_handle() { + local id=${1:-} lock + validate_id "$id" + lock=$(secondmate_reply_lifecycle_lock_path "$STATE" "$id") + ( + fm_lock_acquire_wait "$lock" || die "cannot lock remote reply lifecycle for $id" + trap 'fm_lock_release "$lock"' EXIT + cmd_handle_locked "$@" + ) +} + +retirement_capture_scan() { + local id=$1 sid inbox path base seq pending=0 + sid=$(source_id "$id") + inbox="$STATE/procevent-inbox" + [ -e "$inbox" ] || return 1 + [ -d "$inbox" ] && [ ! -L "$inbox" ] || die "remote reply inbox is unsafe" + for path in "$inbox/$sid".*.result "$inbox/$sid".*.adapter "$inbox/$sid".*.handled; do + [ -e "$path" ] || [ -L "$path" ] || continue + [ -f "$path" ] && [ ! -L "$path" ] || die "remote reply capture is unsafe: $path" + done + for path in "$inbox/$sid".*.result; do + [ -e "$path" ] || continue + base=${path%.result} + seq=${base##*.} + case "$seq" in ''|*[!0-9]*) die "remote reply capture has an invalid generation: $path" ;; esac + [ -f "$base.adapter" ] && [ ! -L "$base.adapter" ] \ + || die "remote reply capture has no safe adapter record: $path" + [ -e "$base.handled" ] || pending=$((pending + 1)) + done + RETIREMENT_PENDING=$pending + RETIREMENT_INBOX=$inbox + return 0 +} + +cmd_retire_quiesce_locked() { + local id=${1:-} force=${2:-} sid + validate_id "$id" + [ -z "$force" ] || [ "$force" = --force ] || die "invalid retirement option: $force" + sid=$(source_id "$id") + "$SCRIPT_DIR/fm-procevent.sh" retire "$sid" || return 1 + RETIREMENT_PENDING=0 + retirement_capture_scan "$id" || true + if [ "$force" != --force ] && [ "$RETIREMENT_PENDING" -gt 0 ]; then + die "remote reply retirement refused with $RETIREMENT_PENDING unhandled captured result(s)" + fi +} + +cmd_retire_finalize_locked() { + local id=${1:-} force=${2:-} sid path + validate_id "$id" + [ -z "$force" ] || [ "$force" = --force ] || die "invalid retirement option: $force" + sid=$(source_id "$id") + RETIREMENT_PENDING=0 + if retirement_capture_scan "$id"; then + if [ "$force" != --force ] && [ "$RETIREMENT_PENDING" -gt 0 ]; then + die "remote reply retirement refused with $RETIREMENT_PENDING unhandled captured result(s)" + fi + if [ "$force" = --force ]; then + for path in "$RETIREMENT_INBOX/$sid".*.result "$RETIREMENT_INBOX/$sid".*.adapter "$RETIREMENT_INBOX/$sid".*.handled; do + [ -e "$path" ] || continue + rm -f -- "$path" || die "cannot discard remote reply capture: $path" + done + fi + fi + rm -f -- "$(cursor_path "$id")" + rm -f -- "$CURSOR_DIR/$id".*.ingested +} + +cmd_retire() { + local id=${1:-} force=${2:-} lock + validate_id "$id" + lock=$(secondmate_reply_lifecycle_lock_path "$STATE" "$id") + ( + fm_lock_acquire_wait "$lock" || die "cannot lock remote reply lifecycle for $id" + trap 'fm_lock_release "$lock"' EXIT + cmd_retire_quiesce_locked "$id" "$force" || return 1 + cmd_retire_finalize_locked "$id" "$force" + ) +} + +require_parent_lifecycle_lock() { + local id=$1 lock owner pid + lock=$(secondmate_reply_lifecycle_lock_path "$STATE" "$id") + if [ -L "$lock" ]; then + owner=$(fm_lock_link_owner "$lock" 2>/dev/null || true) + [ -n "$owner" ] || die "remote reply lifecycle lock ownership is invalid" + else + owner=$lock + fi + pid=$(cat "$owner/pid" 2>/dev/null || true) + [ "$pid" = "$PPID" ] || die "remote reply lifecycle lock is not held by the caller" +} + +case "${1:-}" in + arm) shift; [ "$#" -eq 1 ] || usage; cmd_arm "$@" ;; + arm-locked) shift; [ "$#" -eq 1 ] || usage; require_parent_lifecycle_lock "$1"; cmd_arm_locked "$@" ;; + source) shift; [ "$#" -eq 1 ] || usage; cmd_source "$@" ;; + handle) shift; [ "$#" -eq 3 ] || usage; cmd_handle "$@" ;; + ingest) shift; [ "$#" -eq 2 ] || usage; cmd_ingest "$@" ;; + classify) shift; [ "$#" -eq 1 ] || usage; classify_result "$1" ;; + terminal) shift; [ "$#" -eq 1 ] || usage; [ -s "$1" ] ;; + source-id) shift; [ "$#" -eq 1 ] || usage; source_id "$1" ;; + retire) shift; [ "$#" -ge 1 ] && [ "$#" -le 2 ] || usage; cmd_retire "$@" ;; + retire-quiesce-locked) shift; [ "$#" -ge 1 ] && [ "$#" -le 2 ] || usage; require_parent_lifecycle_lock "$1"; cmd_retire_quiesce_locked "$@" ;; + retire-finalize-locked) shift; [ "$#" -ge 1 ] && [ "$#" -le 2 ] || usage; require_parent_lifecycle_lock "$1"; cmd_retire_finalize_locked "$@" ;; + ''|-h|--help|help) usage ;; + *) die "unknown command: $1" ;; +esac diff --git a/bin/fm-project-mode.sh b/bin/fm-project-mode.sh index 6a6754c10d7..6a97ce2dfed 100755 --- a/bin/fm-project-mode.sh +++ b/bin/fm-project-mode.sh @@ -1,24 +1,41 @@ #!/usr/bin/env bash -# Resolve a project's delivery mode and yolo flag from the data/projects.md registry. +# Resolve a project's REGISTERED delivery posture from the data/projects.md registry. # Prints two words to stdout: " " where mode is one of # no-mistakes|direct-PR|local-only and yolo is on|off. # +# MECHANICAL CONSUMERS ONLY. This answers "what posture did the captain register +# for this project", never "how does this task ship". A task's delivery mode and +# yolo are resolved by firstmate at intake and passed explicitly to +# bin/fm-brief.sh, bin/fm-spawn.sh, and bin/fm-promote.sh (AGENTS.md section 7). +# The consumers are bin/fm-fleet-sync.sh (skip local-only clones), +# bin/fm-home-seed.sh (refuse local-only seeding, run no-mistakes init), and +# bin/fm-spawn.sh's advisory registry-deviation notice. +# # Registry line format (data/projects.md): # - - (added ) -> no-mistakes off (legacy default) # - [] - (added ) -> off # - [ +yolo] - (added ) -> on # -# mode = how a finished change reaches main: -# no-mistakes full pipeline -> PR -> captain merge (default) -# direct-PR push + PR via gh-axi, no pipeline -> captain merge -# local-only local branch, no remote/PR -> captain approve -> guarded local merge +# Registered modes: +# no-mistakes full pipeline -> PR -> configured merge authority (default) +# direct-PR push + PR via gh-axi, no pipeline +# local-only local branch, no remote/PR, guarded local merge +# no-mistakes-prod-only a conditional policy, not a task mode: firstmate +# classifies each task's surface at intake (the +# project-management skill owns that classification). +# Mechanical output maps it to its most rigorous leg, +# no-mistakes, so sync, seeding, and init treat such a +# project as the remote-backed pipeline project it is. # yolo (orthogonal) = when on, firstmate may make routine approval decisions itself. # AGENTS.md section 7 is the single owner of authority exceptions, including # ask-user contract expansion and stronger captain boundaries. # +# --raw prints the registered annotation unmapped, so a caller that must tell a +# conditional policy apart from a flat mode sees "no-mistakes-prod-only" itself. +# # An unknown/missing project or unknown mode falls back to "no-mistakes off" and warns # to stderr, so a typo never silently drops the gate. -# Usage: fm-project-mode.sh +# Usage: fm-project-mode.sh [--raw] set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -26,7 +43,12 @@ FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" DATA="${FM_DATA_OVERRIDE:-$FM_HOME/data}" REG="$DATA/projects.md" -NAME=${1:?usage: fm-project-mode.sh } +RAW=0 +if [ "${1:-}" = "--raw" ]; then + RAW=1 + shift +fi +NAME=${1:?usage: fm-project-mode.sh [--raw] } if [ ! -f "$REG" ]; then echo "warn: no registry at $REG; defaulting $NAME to no-mistakes off" >&2 @@ -59,8 +81,13 @@ fi mode=${parsed%% *} yolo=${parsed##* } case "$mode" in - no-mistakes|direct-PR|local-only) ;; + no-mistakes|direct-PR|local-only|no-mistakes-prod-only) ;; *) echo "warn: unknown mode \"$mode\" for $NAME; defaulting to no-mistakes off" >&2; mode=no-mistakes; yolo=off ;; esac case "$yolo" in on|off) ;; *) yolo=off ;; esac +# A conditional policy is not a task mode. Mechanical callers get its most +# rigorous leg; --raw callers get the annotation itself (see the header). +if [ "$RAW" -eq 0 ] && [ "$mode" = no-mistakes-prod-only ]; then + mode=no-mistakes +fi echo "$mode $yolo" diff --git a/bin/fm-promote.sh b/bin/fm-promote.sh index 827c17998f2..92c3e53448a 100755 --- a/bin/fm-promote.sh +++ b/bin/fm-promote.sh @@ -5,25 +5,84 @@ # again. After promoting, send the crewmate its ship instructions via fm-send.sh # (inventory scratch state, reset to a clean default-branch base, carry over only # intended fix changes, create branch fm/, implement, then report done -# according to the project's delivery mode). -# Usage: fm-promote.sh +# according to this task's delivery mode). +# A scout records no delivery posture, so promotion is where this task's delivery +# contract is decided: --mode and --yolo are REQUIRED and written into the meta +# alongside the kind= flip. Firstmate resolves both at promotion time, having just +# read the scout's report (AGENTS.md section 7); data/projects.md holds the +# captain's standing posture as context, and this script never looks it up. +# no-mistakes-prod-only is a registry policy rather than a task mode and is refused. +# Usage: fm-promote.sh --mode --yolo set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" FM_ROOT="${FM_ROOT_OVERRIDE:-$(cd "$SCRIPT_DIR/.." && pwd)}" FM_HOME="${FM_HOME:-${FM_ROOT_OVERRIDE:-$FM_ROOT}}" STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" + +MODE= +YOLO= +MODE_SET=0 +YOLO_SET=0 +POS=() +want_value= +for a in "$@"; do + if [ -n "$want_value" ]; then + case "$a" in + --*) echo "error: --$want_value requires a value" >&2; exit 1 ;; + esac + case "$want_value" in + mode) MODE=$a; MODE_SET=1 ;; + yolo) YOLO=$a; YOLO_SET=1 ;; + esac + want_value= + continue + fi + case "$a" in + --mode) want_value=mode ;; + --mode=*) MODE=${a#--mode=}; MODE_SET=1 ;; + --yolo) want_value=yolo ;; + --yolo=*) YOLO=${a#--yolo=}; YOLO_SET=1 ;; + *) POS+=("$a") ;; + esac +done +[ -z "$want_value" ] || { echo "error: --$want_value requires a value" >&2; exit 1; } +[ "${#POS[@]}" -ge 1 ] || { echo "usage: fm-promote.sh --mode --yolo " >&2; exit 1; } +[ "$MODE_SET" -eq 1 ] || { + echo "error: promotion requires --mode ; decide it now from the scout's findings and the project's registered posture in data/projects.md" >&2 + exit 1 +} +[ "$YOLO_SET" -eq 1 ] || { + echo "error: promotion requires --yolo ; it is this task's routine approval authority, not a project lookup" >&2 + exit 1 +} +case "$MODE" in + no-mistakes|direct-PR|local-only) ;; + no-mistakes-prod-only) + echo "error: no-mistakes-prod-only is a registry policy, not a task mode; classify this task's surface and resolve it to no-mistakes or direct-PR" >&2 + exit 1 ;; + *) echo "error: --mode must be one of no-mistakes, direct-PR, local-only (got '$MODE')" >&2; exit 1 ;; +esac +case "$YOLO" in + on|off) ;; + *) echo "error: --yolo must be on or off (got '$YOLO')" >&2; exit 1 ;; +esac + "$FM_ROOT/bin/fm-guard.sh" || true -ID=$1 +ID=${POS[0]} META="$STATE/$ID.meta" [ -f "$META" ] || { echo "error: no meta for task $ID at $META" >&2; exit 1; } grep -qx 'kind=scout' "$META" || { echo "error: task $ID is not a scout task (kind=scout not in meta)" >&2; exit 1; } TMP="$META.tmp" -grep -v '^kind=' "$META" > "$TMP" -echo "kind=ship" >> "$TMP" +grep -v -e '^kind=' -e '^mode=' -e '^yolo=' "$META" > "$TMP" +{ + echo "kind=ship" + echo "mode=$MODE" + echo "yolo=$YOLO" +} >> "$TMP" mv "$TMP" "$META" HOME_Q=$(printf '%q' "$FM_HOME") -echo "promoted $ID to ship (teardown protection restored)" -echo "next: FM_HOME=$HOME_Q bin/fm-send.sh fm-$ID ''" +echo "promoted $ID to ship mode=$MODE yolo=$YOLO (teardown protection restored)" +echo "next: FM_HOME=$HOME_Q bin/fm-send.sh fm-$ID ''" diff --git a/bin/fm-push-transition-lib.sh b/bin/fm-push-transition-lib.sh index f5711f21e7a..1fec2d4f61a 100644 --- a/bin/fm-push-transition-lib.sh +++ b/bin/fm-push-transition-lib.sh @@ -129,7 +129,7 @@ handle_push_transition() { # fm_backend_commit_transition "$backend" "$STATE" "$session" "$record" || exit 1 return fi - reason="stale: $window (herdr: agent $to - waiting on human, escalated immediately, not via wedge timer)" + reason="stale: $window ($(stale_detail_blocked_on_human "$backend" "$to"))" fm_wake_append stale "$window" "$reason" || exit 1 fm_backend_commit_transition "$backend" "$STATE" "$session" "$record" || exit 1 mark_surfaced "$STATE/$task.status" diff --git a/bin/fm-quota-axi-lib.sh b/bin/fm-quota-axi-lib.sh index 441c9ce2c92..98bb30c1171 100644 --- a/bin/fm-quota-axi-lib.sh +++ b/bin/fm-quota-axi-lib.sh @@ -7,6 +7,10 @@ # those fields a dispatch candidate cannot be checked against the authentication # surface it actually uses, which is how one harness's expired CLI token used to # produce a captain-facing sign-out claim for a candidate that never read it. +# 0.1.16 already emits schemaVersion 3 with per-model effectiveAvailability; +# 0.1.17 only adds optional runway under that same schema. quota-array-dispatch +# treats absent runway or pace as disclosed uncertainty, so the floor stays +# 0.1.16 rather than tracking the latest additive field. # # This file is the single owner of that version number. bin/fm-bootstrap.sh # turns a failing check into the operator-facing MISSING diagnostic, which is diff --git a/bin/fm-remote-delta-read.sh b/bin/fm-remote-delta-read.sh new file mode 100755 index 00000000000..28670b22d23 --- /dev/null +++ b/bin/fm-remote-delta-read.sh @@ -0,0 +1,182 @@ +#!/usr/bin/env bash +# Blocking, non-destructive delta read for a remote secondmate append-only log. +# +# Usage: +# fm-remote-delta-read.sh [wait-seconds] +# +# The reader validates continuity by hashing the exact prefix represented by the +# caller's cursor. It then blocks until at least one complete appended line is +# available, returns at most 65536 payload bytes, and never truncates or consumes +# the source. A shortened or changed prefix returns a structured continuity-break +# result instead of silently rebasing the cursor. +set -eu + +FM_HOME=${FM_HOME:?FM_HOME is required} +MAX_BYTES=${FM_REMOTE_DELTA_MAX_BYTES:-65536} +POLL_SECONDS=${FM_REMOTE_DELTA_POLL_SECONDS:-0.2} + +die() { printf 'error: %s\n' "$1" >&2; exit 1; } +usage() { sed -n '2,11p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } + +sha256_file() { + if command -v shasum >/dev/null 2>&1; then + shasum -a 256 "$1" | awk '{print $1}' + elif command -v sha256sum >/dev/null 2>&1; then + sha256sum "$1" | awk '{print $1}' + else + die "no SHA-256 tool is available" + fi +} + +copy_prefix() { # + if [ "$2" -eq 0 ]; then + : > "$3" + else + head -c "$2" "$1" > "$3" + fi +} + +snapshot_log() { # + local file=$1 destination=$2 size_file=$3 parent base actual_parent + parent=$(dirname "$file") + base=$(basename "$file") + ( + CDPATH='' cd -- "$parent" 2>/dev/null || exit 1 + actual_parent=$(pwd -P) || exit 1 + [ "$actual_parent" = "$parent" ] || exit 1 + perl -MFcntl=:DEFAULT -e ' + my ($path, $destination, $size_file, $offset, $max_bytes) = @ARGV; + sysopen(my $source, $path, O_RDONLY | O_NOFOLLOW) or exit 1; + my @stat = stat $source or exit 1; + exit 1 unless -f _; + my $size = $stat[7]; + exit 1 unless $size =~ /\A\d+\z/; + my $limit = $size; + my $bound = $offset + $max_bytes; + $limit = $bound if $limit > $bound; + open(my $output, ">", $destination) or exit 1; + binmode $source; + binmode $output; + my $remaining = $limit; + while ($remaining > 0) { + my $wanted = $remaining > 65536 ? 65536 : $remaining; + my $read = read($source, my $buffer, $wanted); + exit 1 unless defined $read && $read > 0; + print {$output} $buffer or exit 1; + $remaining -= $read; + } + close $output or exit 1; + open(my $size_output, ">", $size_file) or exit 1; + print {$size_output} "$size\n" or exit 1; + close $size_output or exit 1; + ' "$base" "$destination" "$size_file" "$OFFSET" "$MAX_BYTES" + ) +} + +resolve_log() { # + local rel=$1 home_real parent_real parent base path + case "$rel" in ''|/*|*'//'*) die "log must be a nonempty relative path" ;; esac + case "/$rel/" in */../*|*/./*) die "log traversal is not allowed: $rel" ;; esac + case "$rel" in *$'\n'*|*$'\r'*|*$'\t'*) die "log path contains control characters" ;; esac + home_real=$(CDPATH='' cd -- "$FM_HOME" 2>/dev/null && pwd -P) || die "FM_HOME is unavailable" + parent=$(dirname "$rel") + base=$(basename "$rel") + parent_real=$(CDPATH='' cd -- "$FM_HOME/$parent" 2>/dev/null && pwd -P) || die "log parent is unavailable: $rel" + case "$parent_real" in "$home_real"|"$home_real"/*) ;; *) die "log escapes FM_HOME: $rel" ;; esac + path="$parent_real/$base" + if [ -e "$path" ] || [ -L "$path" ]; then + [ -f "$path" ] && [ ! -L "$path" ] || die "log is not a non-symlink regular file: $rel" + fi + printf '%s\n' "$path" +} + +emit_break() { # + printf 'schema=fm-remote-delta.v1\n' + printf 'status=continuity-broken\n' + printf 'path=%s\n' "$REL" + printf 'from_offset=%s\n' "$OFFSET" + printf 'to_offset=%s\n' "$2" + printf 'from_prefix_sha256=%s\n' "$PREFIX" + printf 'to_prefix_sha256=%s\n' "$3" + printf 'payload_sha256=%s\n' "$EMPTY_HASH" + printf 'payload_bytes=0\n' + printf 'reason=%s\n\n' "$1" +} + +[ "$#" -ge 3 ] && [ "$#" -le 4 ] || usage +REL=$1 +OFFSET=$2 +PREFIX=$3 +WAIT=${4:-55} +case "$OFFSET" in ''|*[!0-9]*) die "offset must be a nonnegative integer" ;; esac +case "$PREFIX" in *[!A-Fa-f0-9]*|'') die "prefix-sha256 must be hexadecimal" ;; esac +[ "${#PREFIX}" -eq 64 ] || die "prefix-sha256 must be 64 hexadecimal characters" +PREFIX=$(printf '%s' "$PREFIX" | tr 'A-F' 'a-f') +case "$WAIT" in ''|*[!0-9]*) die "wait-seconds must be a nonnegative integer" ;; esac +[ "$WAIT" -le 300 ] || die "wait-seconds exceeds the 300-second safety bound" +case "$MAX_BYTES" in ''|*[!0-9]*|0) die "FM_REMOTE_DELTA_MAX_BYTES must be a positive integer" ;; esac +[ "$MAX_BYTES" -le 1048576 ] || die "FM_REMOTE_DELTA_MAX_BYTES exceeds the safety bound" + +LOG=$(resolve_log "$REL") +TMP=$(mktemp -d "${TMPDIR:-/tmp}/fm-remote-delta.XXXXXX") || die "cannot create delta staging directory" +trap 'rm -rf -- "$TMP"' EXIT +: > "$TMP/empty" +EMPTY_HASH=$(sha256_file "$TMP/empty") +START=$(date +%s) +while :; do + if [ -e "$LOG" ] || [ -L "$LOG" ]; then + [ -f "$LOG" ] && [ ! -L "$LOG" ] || die "log changed into an unsafe file: $REL" + snapshot_log "$LOG" "$TMP/source" "$TMP/size" \ + || die "log could not be captured safely: $REL" + SIZE=$(tr -d ' ' < "$TMP/size") + if [ "$SIZE" -lt "$OFFSET" ]; then + copy_prefix "$TMP/source" "$SIZE" "$TMP/prefix" + ACTUAL=$(sha256_file "$TMP/prefix") + emit_break truncated "$SIZE" "$ACTUAL" + exit 0 + fi + copy_prefix "$TMP/source" "$OFFSET" "$TMP/prefix" + ACTUAL=$(sha256_file "$TMP/prefix") + if [ "$ACTUAL" != "$PREFIX" ]; then + emit_break prefix-changed "$SIZE" "$ACTUAL" + exit 0 + fi + if [ "$SIZE" -gt "$OFFSET" ]; then + tail -c "+$((OFFSET + 1))" "$TMP/source" | head -c "$MAX_BYTES" > "$TMP/chunk" || true + COMPLETE_BYTES=$(LC_ALL=C od -An -v -tu1 "$TMP/chunk" | awk ' + { for (i = 1; i <= NF; i++) { bytes++; if ($i == 10) complete=bytes } } + END { print complete + 0 } + ') + if [ "$COMPLETE_BYTES" -eq 0 ]; then : > "$TMP/payload"; else head -c "$COMPLETE_BYTES" "$TMP/chunk" > "$TMP/payload"; fi + BYTES=$(LC_ALL=C wc -c < "$TMP/payload" | tr -d ' ') + if [ "$BYTES" -gt 0 ]; then + TO=$((OFFSET + BYTES)) + copy_prefix "$TMP/source" "$TO" "$TMP/to-prefix" + TO_HASH=$(sha256_file "$TMP/to-prefix") + PAYLOAD_HASH=$(sha256_file "$TMP/payload") + printf 'schema=fm-remote-delta.v1\n' + printf 'status=delta\n' + printf 'path=%s\n' "$REL" + printf 'from_offset=%s\n' "$OFFSET" + printf 'to_offset=%s\n' "$TO" + printf 'from_prefix_sha256=%s\n' "$PREFIX" + printf 'to_prefix_sha256=%s\n' "$TO_HASH" + printf 'payload_sha256=%s\n' "$PAYLOAD_HASH" + printf 'payload_bytes=%s\n' "$BYTES" + printf 'reason=\n\n' + cat "$TMP/payload" + exit 0 + fi + if [ $((SIZE - OFFSET)) -ge "$MAX_BYTES" ]; then + emit_break line-exceeds-bound "$SIZE" "$ACTUAL" + exit 0 + fi + fi + elif [ "$OFFSET" -ne 0 ] || [ "$PREFIX" != "$EMPTY_HASH" ]; then + emit_break missing 0 "$EMPTY_HASH" + exit 0 + fi + NOW=$(date +%s) + [ $((NOW - START)) -lt "$WAIT" ] || exit 75 + sleep "$POLL_SECONDS" +done diff --git a/bin/fm-remote-doctor.sh b/bin/fm-remote-doctor.sh new file mode 100755 index 00000000000..aad9ce44aa7 --- /dev/null +++ b/bin/fm-remote-doctor.sh @@ -0,0 +1,800 @@ +#!/usr/bin/env bash +# Check, and optionally repair, one remote account's second-mate readiness. +# +# Usage: +# bin/fm-on.sh fm-remote-doctor.sh [--fix] +# +# Run it through fm-on.sh so the fixed entrypoint invokes this readiness owner +# over its plain SSH bootstrap. The command reports the same filesystem-composed +# PATH used by worker jobs while retaining authority to inspect and repair the +# worker itself. +# +# A remote second mate always runs on the Herdr backend in the dedicated +# fm-remote session. Its account therefore needs the Firstmate-owned Aqua Herdr +# agent plus the sibling dev.firstmate.remote-job worker that runs normal fm-on +# commands through the Aqua or Linux job-worker path. Doctor remains invokable +# over the plain-SSH bootstrap path to inspect and repair that worker. SSH cannot +# create an Aqua session, so a host with no GUI login is a human gap rather than +# something --fix attempts to bypass. +# +# Line protocol, one fact per line, stable for script consumers: +# mode=check|fix +# path= +# entrypoint=yes|no +# platform=darwin|linux||unknown +# required =|MISSING +# optional =|absent +# fix =applied: (--fix only) +# fix =failed: (--fix only) +# check =ok: +# check =skip: +# check =fixable: +# check =human: +# action: : +# Every check line is authoritative for the moment it printed: under --fix it is +# the state after the repair attempt, so a human gap is never presented as +# fixed. Any remaining fixable or human gap, and any missing required tool, +# exits non-zero. +# +# --fix is idempotent and closes only automatable gaps: it writes and reloads +# both Firstmate-owned Aqua agents, starts the Linux workers where no Aqua agent +# applies, recreates the entrypoint symlink, and may add an owned ~/.local/bin +# wrapper for a required tool it can discover under nvm, asdf, or mise. It never +# installs packages, creates a login session, writes an auto-login password, +# changes FileVault, stores an account password, or replaces a non-Firstmate +# wrapper; those remain reported gaps. +set -eu + +# Resolve this script's directory with builtins only: a host missing a required +# tool must still reach the report that names it, not die on a bare PATH. +SCRIPT_SELF=${BASH_SOURCE[0]} +SCRIPT_DIR=${SCRIPT_SELF%/*} +[ "$SCRIPT_DIR" != "$SCRIPT_SELF" ] || SCRIPT_DIR=. +SCRIPT_DIR=$(CDPATH='' cd -- "$SCRIPT_DIR" && pwd -P) +FM_ROOT="${FM_ROOT_OVERRIDE:-$(CDPATH='' cd "$SCRIPT_DIR/.." && pwd -P)}" +# shellcheck source=bin/fm-remote-job-lib.sh +. "$SCRIPT_DIR/fm-remote-job-lib.sh" +# shellcheck source=bin/fm-tasks-axi-lib.sh +. "$SCRIPT_DIR/fm-tasks-axi-lib.sh" +REQUIRED_TOOLS=(git jq herdr tasks-axi treehouse) +HARNESS_TOOLS=(claude codex opencode pi pi-signed grok kimi) +OPTIONAL_TOOLS=(tmux no-mistakes gh) +LAUNCH_AGENT_LABEL=dev.firstmate.herdr.fm-remote +# The dedicated remote-secondmate session. The user's interactive Herdr work +# remains in the separate default session, which this readiness check never +# requires or changes. +HERDR_SESSION_NAME=fm-remote +LAUNCH_AGENT_DIR="${HOME:-}/Library/LaunchAgents" +LAUNCH_AGENT_PLIST="$LAUNCH_AGENT_DIR/$LAUNCH_AGENT_LABEL.plist" +LAUNCH_AGENT_LOG_DIR="${HOME:-}/Library/Logs" +LAUNCH_AGENT_LOG="$LAUNCH_AGENT_LOG_DIR/$LAUNCH_AGENT_LABEL.log" +ENTRYPOINT_LINK="${HOME:-}/.local/bin/fm-remote-entrypoint.sh" + +usage() { sed -n '2,5p' "$0" | sed 's/^# \{0,1\}//'; exit 2; } + +MODE=check +case "${1:-}" in + '') ;; + --fix) MODE=fix; shift ;; + --worker-tool-probe) + [ "${FM_REMOTE_JOB_ACTIVE:-}" = 1 ] || { printf 'error: worker tool probe requires the remote job worker\n' >&2; exit 64; } + MODE='worker-tool-probe' + shift + ;; + *) usage ;; +esac +[ "$#" -eq 0 ] || usage + +PLATFORM=$(fm_remote_job_platform) +UID_NUM=$(id -u 2>/dev/null) || UID_NUM= + +CHECK_NAMES=() +CHECK_VALUES=() +CHECK_ACTIONS=() + +record() { # [operator-action] + CHECK_NAMES+=("$1") + CHECK_VALUES+=("$2") + CHECK_ACTIONS+=("${3:-}") +} + +check_value() { # ; prints the recorded value, empty when unrecorded + local i=0 + while [ "$i" -lt "${#CHECK_NAMES[@]}" ]; do + if [ "${CHECK_NAMES[$i]}" = "$1" ]; then + printf '%s' "${CHECK_VALUES[$i]}" + return 0 + fi + i=$((i + 1)) + done + return 1 +} + +check_is_ok() { # + case "$(check_value "$1" 2>/dev/null || true)" in ok:*) return 0 ;; esac + return 1 +} + +set_check() { # [operator-action] + local i=0 + while [ "$i" -lt "${#CHECK_NAMES[@]}" ]; do + if [ "${CHECK_NAMES[$i]}" = "$1" ]; then + CHECK_VALUES[i]=$2 + CHECK_ACTIONS[i]=${3:-} + return 0 + fi + i=$((i + 1)) + done + record "$@" +} + +herdr_cli_available() { + local herdr_bin jq_bin + herdr_bin=$(command -v herdr 2>/dev/null || true) + jq_bin=$(command -v jq 2>/dev/null || true) + [ -n "$herdr_bin" ] && [ -x "$herdr_bin" ] && [ -n "$jq_bin" ] && [ -x "$jq_bin" ] +} + +# The herdr adapter is the single owner of session-scoped herdr invocation and +# of starting a server, so read and start through it rather than restating +# either here. Sourced only when both tools resolve, so a bare host still +# reports its gaps instead of failing to load. +herdr_adapter_load() { + [ -z "${FM_REMOTE_DOCTOR_HERDR_LOADED:-}" ] || return 0 + herdr_cli_available || return 1 + [ -f "$SCRIPT_DIR/fm-backend.sh" ] && [ -f "$SCRIPT_DIR/backends/herdr.sh" ] || return 1 + # shellcheck source=bin/fm-backend.sh + . "$SCRIPT_DIR/fm-backend.sh" || return 1 + fm_backend_source herdr || return 1 + FM_REMOTE_DOCTOR_HERDR_LOADED=1 +} + +herdr_server_running() { + local running + herdr_adapter_load || return 1 + running=$(fm_backend_herdr_cli "$HERDR_SESSION_NAME" status --json 2>/dev/null \ + | jq -r '.server.running // false' 2>/dev/null) || return 1 + [ "$running" = true ] +} + +launch_agent_is_aqua() { + local stripped + [ -f "$LAUNCH_AGENT_PLIST" ] && [ ! -L "$LAUNCH_AGENT_PLIST" ] || return 1 + stripped=$(tr -d ' \t\r\n' < "$LAUNCH_AGENT_PLIST" 2>/dev/null) || return 1 + case "$stripped" in + *'LimitLoadToSessionTypeAqua'*) return 0 ;; + esac + return 1 +} + +render_launch_agent() { # + local herdr_bin=$1 + cat < + + + + Label + $LAUNCH_AGENT_LABEL + ProgramArguments + + $herdr_bin + server + --session + $HERDR_SESSION_NAME + + LimitLoadToSessionType + Aqua + RunAtLoad + + KeepAlive + + StandardOutPath + $LAUNCH_AGENT_LOG + StandardErrorPath + $LAUNCH_AGENT_LOG + + +XML +} + +launch_agent_contract_matches() { + local herdr_bin actual expected + [ -f "$LAUNCH_AGENT_PLIST" ] && [ ! -L "$LAUNCH_AGENT_PLIST" ] || return 1 + herdr_bin=$(command -v herdr 2>/dev/null) || return 1 + actual=$(tr -d ' \t\r\n' < "$LAUNCH_AGENT_PLIST" 2>/dev/null) || return 1 + expected=$(render_launch_agent "$herdr_bin" | tr -d ' \t\r\n') || return 1 + [ "$actual" = "$expected" ] +} + +launch_agent_loaded_contract_matches() { + local loaded herdr_bin herdr_compact plist_compact log_compact args + herdr_bin=$(command -v herdr 2>/dev/null) || return 1 + loaded=$(launchctl print "gui/$UID_NUM/$LAUNCH_AGENT_LABEL" 2>/dev/null) || return 1 + loaded=$(printf '%s' "$loaded" | tr -d ' \t\r\n') || return 1 + herdr_compact=$(printf '%s' "$herdr_bin" | tr -d ' \t\r\n') || return 1 + plist_compact=$(printf '%s' "$LAUNCH_AGENT_PLIST" | tr -d ' \t\r\n') || return 1 + log_compact=$(printf '%s' "$LAUNCH_AGENT_LOG" | tr -d ' \t\r\n') || return 1 + args="arguments={$herdr_compact"'server--session'"$HERDR_SESSION_NAME}" + [[ "$loaded" == *"path=$plist_compact"* ]] || return 1 + [[ "$loaded" == *"program=$herdr_compact"* ]] || return 1 + [[ "$loaded" == *"$args"* ]] || return 1 + [[ "$loaded" == *"stdoutpath=$log_compact"* ]] || return 1 + [[ "$loaded" == *"stderrpath=$log_compact"* ]] || return 1 + [[ "$loaded" == *'properties=keepalive|runatload'* ]] || return 1 +} + +# --- remote job and tool checks --------------------------------------------- + +remote_job_existing_state() { + local root + root=${FM_REMOTE_JOB_STATE_ROOT:-${HOME:-}/.firstmate/remote-job} + root=$(fm_remote_job_canonical_existing_dir "$root") || return 1 + fm_remote_job_canonical_existing_dir "$root/jobs" >/dev/null || return 1 + # shellcheck disable=SC2034 # The sourceable worker helpers consume the validated state root. + FM_REMOTE_JOB_STATE=$root +} + +remote_job_probe_ok() { + local ready mtime now + [ "${FM_REMOTE_JOB_ACTIVE:-}" = 1 ] && return 0 + remote_job_existing_state || return 1 + ready="$FM_REMOTE_JOB_STATE/worker.ready" + [ -f "$ready" ] && [ ! -L "$ready" ] || return 1 + mtime=$(fm_remote_job_path_mtime "$ready" 2>/dev/null || true) + case "$mtime" in ''|*[!0-9]*) return 1 ;; esac + now=$(date +%s) + [ $((now - mtime)) -le 10 ] +} + +remote_job_identity_ok() { + [ "${FM_REMOTE_JOB_ACTIVE:-}" = 1 ] && return 0 + remote_job_probe_ok || return 1 + fm_remote_job_worker_identity_matches "$FM_ROOT" "${HOME:-}" +} + +check_remote_job_worker() { + local worker + worker="$FM_ROOT/bin/fm-remote-job-worker.sh" + if [ ! -f "$worker" ] || [ -L "$worker" ] || [ ! -x "$worker" ]; then + record remote-job-worker "human: the configured Firstmate code root has no safe remote job worker" \ + "update the remote Firstmate checkout, then rerun this command with --fix" + record remote-job-worker-loaded "skip: no worker executable is available" + record remote-job-probe "skip: no worker executable is available" + return 0 + fi + if [ "$PLATFORM" = darwin ]; then + fm_remote_job_launchagent_paths "${HOME:-}" + if fm_remote_job_launchagent_contract_matches "$FM_ROOT" "${HOME:-}"; then + record remote-job-worker "ok: $FM_REMOTE_JOB_LAUNCH_AGENT_PLIST matches the Firstmate-owned Aqua worker contract" + else + record remote-job-worker "fixable: $FM_REMOTE_JOB_LAUNCH_AGENT_PLIST does not match the Firstmate-owned Aqua worker contract" \ + "rerun this command with --fix to write dev.firstmate.remote-job" + fi + if [ -z "$UID_NUM" ] || ! command -v launchctl >/dev/null 2>&1; then + record remote-job-worker-loaded "human: the remote job worker cannot be inspected without launchctl and an account uid" \ + "restore launchctl and a readable account uid, then rerun this command" + elif fm_remote_job_launchagent_loaded "$FM_ROOT" "${HOME:-}" "$UID_NUM"; then + record remote-job-worker-loaded "ok: $FM_REMOTE_JOB_LABEL is loaded in gui/$UID_NUM" + elif check_is_ok gui-session; then + record remote-job-worker-loaded "fixable: $FM_REMOTE_JOB_LABEL is not loaded in gui/$UID_NUM" \ + "rerun this command with --fix to bootstrap the worker" + else + record remote-job-worker-loaded "human: $FM_REMOTE_JOB_LABEL cannot be loaded because gui/$UID_NUM has no login session" \ + "close the login-session gap first; SSH cannot create an Aqua session" + fi + else + local pid + pid=$(cat "${FM_REMOTE_JOB_STATE_ROOT:-${HOME:-}/.firstmate/remote-job}/worker.pid" 2>/dev/null || true) + if [ "${FM_REMOTE_JOB_ACTIVE:-}" = 1 ] || + { remote_job_existing_state && case "$pid" in ''|*[!0-9]*) false ;; *) kill -0 "$pid" 2>/dev/null ;; esac; }; then + record remote-job-worker "ok: the Linux remote job worker is running" + record remote-job-worker-loaded "skip: Aqua launch agents do not apply on $PLATFORM" + else + record remote-job-worker "fixable: the Linux remote job worker is not running" \ + "rerun this command with --fix to start it" + record remote-job-worker-loaded "skip: Aqua launch agents do not apply on $PLATFORM" + fi + fi + if ! remote_job_probe_ok; then + record remote-job-probe "fixable: the remote job worker has not reported a fresh probe" \ + "rerun this command with --fix to restart the worker, then rerun through fm-on.sh" + elif ! remote_job_identity_ok; then + set_check remote-job-worker "fixable: the running remote job worker does not match the current Firstmate code" \ + "rerun this command with --fix to reload the current worker" + record remote-job-probe "fixable: the remote job worker identity is stale, so its runtime cannot be probed" \ + "rerun this command with --fix to reload the current worker" + else + record remote-job-probe "ok: the remote job worker published a fresh heartbeat" + fi +} + +report_required_tools() { + local tool resolved harness + MISSING=() + for tool in "${REQUIRED_TOOLS[@]}"; do + resolved=$(command -v "$tool" 2>/dev/null || true) + if [ -n "$resolved" ] && [ -x "$resolved" ]; then + if [ "$tool" = tasks-axi ] && ! fm_tasks_axi_compatible; then + printf 'required tasks-axi=MISSING (incompatible)\n' + MISSING+=(tasks-axi) + else + printf 'required %s=%s\n' "$tool" "$resolved" + fi + else + printf 'required %s=MISSING\n' "$tool" + MISSING+=("$tool") + fi + done + for harness in "${HARNESS_TOOLS[@]}"; do + resolved=$(command -v "$harness" 2>/dev/null || true) + if [ -n "$resolved" ] && [ -x "$resolved" ]; then + printf 'required harness=%s:%s\n' "$harness" "$resolved" + return 0 + fi + done + printf 'required harness=MISSING\n' + MISSING+=(harness) +} + +report_required_tools_from_worker() { + local job_id probe_stdout probe_stderr probe_exit line fact name value + local expected=6 count=0 valid=1 seen=' ' + if ! job_id=$(fm_remote_job_stage "${HOME:-}" "$FM_ROOT" "${FM_HOME:-}" \ + fm-remote-doctor.sh --worker-tool-probe /dev/null || true + set_check remote-job-probe "fixable: the remote job worker did not complete the required-tool probe" \ + "rerun this command with --fix to restart the worker" + report_required_tools + return 0 + fi + probe_stdout=$FM_REMOTE_JOB_STDOUT + probe_stderr=$FM_REMOTE_JOB_STDERR + probe_exit=$FM_REMOTE_JOB_EXIT + MISSING=() + while IFS= read -r line; do + case "$line" in required\ *=*) ;; *) valid=0; continue ;; esac + fact=${line#required } + name=${fact%%=*} + value=${fact#*=} + case "$name" in git|jq|herdr|tasks-axi|treehouse|harness) ;; *) valid=0; continue ;; esac + case "$seen" in *" $name "*) valid=0; continue ;; esac + seen="$seen$name " + count=$((count + 1)) + case "$value" in MISSING*) MISSING+=("$name") ;; '') valid=0 ;; esac + done < "$probe_stdout" + [ "$count" -eq "$expected" ] || valid=0 + [ ! -s "$probe_stderr" ] || valid=0 + case "$probe_exit:${#MISSING[@]}" in 0:0|1:[1-9]*) ;; *) valid=0 ;; esac + if [ "$valid" -eq 1 ]; then + cat "$probe_stdout" + set_check remote-job-probe "ok: the remote job worker completed the required-tool probe" + else + set_check remote-job-probe "fixable: the remote job worker returned an invalid required-tool probe result" \ + "rerun this command with --fix to restart the worker" + report_required_tools + fi + fm_remote_job_reap "${HOME:-}" "$job_id" 2>/dev/null || true +} + +wrapper_is_firstmate_owned() { # + local path=$1 first second + [ -f "$path" ] && [ ! -L "$path" ] || return 1 + IFS= read -r first < "$path" || return 1 + IFS= read -r second < <(tail -n +2 "$path") || return 1 + [ "$first" = '#!/usr/bin/env bash' ] && [ "$second" = '# Firstmate remote tool wrapper v1' ] +} + +repair_tool_wrapper() { # + local tool=$1 target wrapper tmp + local resolved + resolved=$(command -v "$tool" 2>/dev/null || true) + [ -n "$resolved" ] && [ -x "$resolved" ] && return 0 + target=$(fm_remote_job_manager_tool "${HOME:-}" "$tool" 2>/dev/null || true) + [ -n "$target" ] || return 1 + wrapper="${HOME:-}/.local/bin/$tool" + if [ -e "$wrapper" ] || [ -L "$wrapper" ]; then + if ! wrapper_is_firstmate_owned "$wrapper"; then + fix_report "required-$tool" failed "$wrapper exists and is not Firstmate-owned" + return 1 + fi + else + if ! mkdir -p "${HOME:-}/.local/bin" 2>/dev/null || [ -L "${HOME:-}/.local/bin" ]; then + fix_report "required-$tool" failed "cannot create ${HOME:-}/.local/bin" + return 1 + fi + fi + tmp="${HOME:-}/.local/bin/.$tool.tmp.$$" + { + printf '%s\n' '#!/usr/bin/env bash' + printf '%s\n' '# Firstmate remote tool wrapper v1' + printf 'exec %q "$@"\n' "$target" + } > "$tmp" || { rm -f -- "$tmp"; fix_report "required-$tool" failed "cannot write $wrapper"; return 1; } + if ! chmod 0700 "$tmp" || ! mv -f -- "$tmp" "$wrapper"; then + rm -f -- "$tmp" + fix_report "required-$tool" failed "cannot publish $wrapper" + return 1 + fi + fix_report "required-$tool" applied "linked the discoverable version-manager tool at $wrapper" +} + +repair_required_wrappers() { + local tool resolved + for tool in "${REQUIRED_TOOLS[@]}"; do + repair_tool_wrapper "$tool" || true + done + for tool in "${HARNESS_TOOLS[@]}"; do + resolved=$(command -v "$tool" 2>/dev/null || true) + [ -z "$resolved" ] || [ ! -x "$resolved" ] || return 0 + done + for tool in "${HARNESS_TOOLS[@]}"; do + fm_remote_job_manager_tool "${HOME:-}" "$tool" >/dev/null 2>&1 || continue + repair_tool_wrapper "$tool" && return 0 + done +} + +fix_remote_job_worker() { + if fm_remote_job_ensure_worker "$FM_ROOT" "${HOME:-}"; then + [ "$FM_REMOTE_JOB_REPAIRED" -eq 0 ] || fix_report remote-job-worker applied "installed or reloaded $FM_REMOTE_JOB_LABEL" + return 0 + fi + fix_report remote-job-worker failed "${FM_REMOTE_JOB_ERROR:-the remote job worker could not start}" + return 1 +} + +# --- checks ----------------------------------------------------------------- + +check_herdr() { + local resolved + if resolved=$(command -v herdr 2>/dev/null) && [ -x "$resolved" ]; then + record herdr "ok: $resolved" + return 0 + fi + record herdr "human: the herdr CLI does not resolve on the remote runtime PATH" \ + "install herdr from https://herdr.dev on that account, or add a ~/.local/bin wrapper for it; a remote second mate always runs on the Herdr backend" +} + +check_gui_session() { + if [ "$PLATFORM" != darwin ]; then + record gui-session "skip: no Aqua login session applies on $PLATFORM" + return 0 + fi + if [ -z "$UID_NUM" ]; then + record gui-session "human: the account uid could not be read, so its login session cannot be inspected" \ + "run 'id -u' on that account and report the failure; Firstmate cannot address gui/ without it" + return 0 + fi + if ! command -v launchctl >/dev/null 2>&1; then + record gui-session "human: launchctl does not resolve, so the login session cannot be inspected" \ + "restore /bin/launchctl on that macOS account; without it no launch agent can be inspected or loaded" + return 0 + fi + if launchctl print "gui/$UID_NUM" >/dev/null 2>&1; then + record gui-session "ok: gui/$UID_NUM" + return 0 + fi + record gui-session "human: no Aqua login session exists for uid $UID_NUM" \ + "log that account in once at the console, and enable automatic login in System Settings > Users & Groups if the machine runs headless; SSH cannot create a GUI session, and Firstmate never writes an auto-login password or changes FileVault" +} + +check_launch_agent() { + if [ "$PLATFORM" != darwin ]; then + record launchagent "skip: launch agents apply only on darwin" + record launchagent-scope "skip: launch agents apply only on darwin" + record launchagent-loaded "skip: launch agents apply only on darwin" + return 0 + fi + if [ -f "$LAUNCH_AGENT_PLIST" ] && [ ! -L "$LAUNCH_AGENT_PLIST" ]; then + if launch_agent_contract_matches; then + record launchagent "ok: $LAUNCH_AGENT_PLIST matches the Firstmate-owned contract" + else + record launchagent "fixable: $LAUNCH_AGENT_PLIST does not match the current Firstmate-owned contract" \ + "rerun this command with --fix to rewrite its label, program arguments, session scope, restart policy, and log paths" + fi + if launch_agent_is_aqua; then + record launchagent-scope "ok: LimitLoadToSessionType=Aqua" + else + record launchagent-scope "fixable: $LAUNCH_AGENT_PLIST is not scoped to the Aqua login session" \ + "rerun this command with --fix to rewrite it with LimitLoadToSessionType=Aqua" + fi + else + record launchagent "fixable: no Firstmate herdr launch agent at $LAUNCH_AGENT_PLIST" \ + "rerun this command with --fix to install it" + record launchagent-scope "skip: no launch agent is installed yet" + fi + check_launch_agent_loaded +} + +check_launch_agent_loaded() { + if [ -z "$UID_NUM" ] || ! command -v launchctl >/dev/null 2>&1; then + record launchagent-loaded "human: the launch agent domain gui/ cannot be inspected on this account" \ + "restore launchctl and a readable account uid, then rerun this command" + return 0 + fi + if launchctl print "gui/$UID_NUM/$LAUNCH_AGENT_LABEL" >/dev/null 2>&1; then + if launch_agent_loaded_contract_matches; then + record launchagent-loaded "ok: gui/$UID_NUM/$LAUNCH_AGENT_LABEL matches the effective contract" + else + record launchagent-loaded "fixable: gui/$UID_NUM/$LAUNCH_AGENT_LABEL does not match the effective Firstmate-owned contract" \ + "rerun this command with --fix to replace the loaded job with the current launch-agent contract" + fi + return 0 + fi + if check_is_ok gui-session; then + record launchagent-loaded "fixable: $LAUNCH_AGENT_LABEL is not loaded into gui/$UID_NUM" \ + "rerun this command with --fix to bootstrap and start it" + return 0 + fi + record launchagent-loaded "human: $LAUNCH_AGENT_LABEL cannot be loaded because gui/$UID_NUM has no login session" \ + "close the login-session gap first; a launch agent can only be bootstrapped into an existing GUI session" +} + +check_herdr_server() { + if ! herdr_cli_available; then + record herdr-server "human: herdr server status cannot be read without both herdr and jq on the runtime PATH" \ + "install the missing tool reported above, then rerun this command" + return 0 + fi + if herdr_server_running; then + record herdr-server "ok: session $HERDR_SESSION_NAME is running" + return 0 + fi + if [ "$PLATFORM" = darwin ] && ! check_is_ok gui-session; then + record herdr-server "human: the herdr server for session $HERDR_SESSION_NAME is not running and there is no GUI login session to start it in" \ + "close the login-session gap first; a server started over SSH would not belong to an Aqua session" + return 0 + fi + record herdr-server "fixable: the herdr server for session $HERDR_SESSION_NAME is not running" \ + "rerun this command with --fix to start it" +} + +check_entrypoint_link() { + local want + if [ -z "${FM_ROOT_OVERRIDE:-}" ]; then + record entrypoint-link "skip: this run did not come through the fixed remote entrypoint" + return 0 + fi + want="$FM_ROOT_OVERRIDE/bin/fm-remote-entrypoint.sh" + if [ -L "$ENTRYPOINT_LINK" ] && [ "$(readlink "$ENTRYPOINT_LINK")" = "$want" ]; then + record entrypoint-link "ok: $ENTRYPOINT_LINK" + return 0 + fi + if [ -e "$ENTRYPOINT_LINK" ] || [ -L "$ENTRYPOINT_LINK" ]; then + record entrypoint-link "human: $ENTRYPOINT_LINK exists but is not the symlink to $want" \ + "inspect that path yourself and replace it with 'ln -sfn $want $ENTRYPOINT_LINK' if it is stale; Firstmate never overwrites a file it did not create there" + return 0 + fi + record entrypoint-link "fixable: no entrypoint symlink at $ENTRYPOINT_LINK" \ + "rerun this command with --fix to create it" +} + +run_checks() { + CHECK_NAMES=() + CHECK_VALUES=() + CHECK_ACTIONS=() + check_herdr + check_gui_session + check_remote_job_worker + check_launch_agent + check_herdr_server + check_entrypoint_link +} + +# --- repairs ---------------------------------------------------------------- + +fix_report() { # applied|failed + printf 'fix %s=%s: %s\n' "$1" "$2" "$3" +} + +write_launch_agent() { + local herdr_bin tmp + if ! herdr_bin=$(command -v herdr 2>/dev/null); then + fix_report launchagent failed "herdr does not resolve, so no launch agent was written" + return 1 + fi + case "$herdr_bin" in + *'&'*|*'<'*|*'>'*|*'"'*|*"'"*) + fix_report launchagent failed "the resolved herdr path contains characters that cannot be embedded in a property list: $herdr_bin" + return 1 + ;; + esac + if ! mkdir -p "$LAUNCH_AGENT_DIR" 2>/dev/null; then + fix_report launchagent failed "cannot create $LAUNCH_AGENT_DIR" + return 1 + fi + mkdir -p "$LAUNCH_AGENT_LOG_DIR" 2>/dev/null || true + tmp="$LAUNCH_AGENT_DIR/.$LAUNCH_AGENT_LABEL.plist.tmp.$$" + render_launch_agent "$herdr_bin" > "$tmp" + chmod 0644 "$tmp" 2>/dev/null || true + if ! mv -f -- "$tmp" "$LAUNCH_AGENT_PLIST" 2>/dev/null; then + rm -f -- "$tmp" + fix_report launchagent failed "cannot publish $LAUNCH_AGENT_PLIST" + return 1 + fi + fix_report launchagent applied "wrote the Aqua-scoped $LAUNCH_AGENT_LABEL launch agent running $herdr_bin server" +} + +# Reload rather than plain bootstrap so a rewritten plist replaces a stale +# in-memory copy, and kickstart so the server is running now rather than at the +# next login. Both are safe to repeat. +reload_launch_agent() { # + local report=$1 out + [ -f "$LAUNCH_AGENT_PLIST" ] || { + fix_report "$report" failed "there is no launch agent to load at $LAUNCH_AGENT_PLIST" + return 1 + } + if [ -z "$UID_NUM" ] || ! command -v launchctl >/dev/null 2>&1; then + fix_report "$report" failed "launchctl or the account uid is unavailable" + return 1 + fi + launchctl bootout "gui/$UID_NUM/$LAUNCH_AGENT_LABEL" >/dev/null 2>&1 || true + if ! out=$(launchctl bootstrap "gui/$UID_NUM" "$LAUNCH_AGENT_PLIST" 2>&1); then + fix_report "$report" failed "launchctl bootstrap gui/$UID_NUM refused: ${out:-no diagnostic}" + return 1 + fi + if ! out=$(launchctl kickstart -k "gui/$UID_NUM/$LAUNCH_AGENT_LABEL" 2>&1); then + fix_report "$report" failed "launchctl kickstart gui/$UID_NUM/$LAUNCH_AGENT_LABEL refused: ${out:-no diagnostic}" + return 1 + fi + if ! wait_for_herdr_server; then + fix_report "$report" failed "the herdr server for session $HERDR_SESSION_NAME did not report running within 10s" + return 1 + fi + fix_report "$report" applied "bootstrapped and started $LAUNCH_AGENT_LABEL in gui/$UID_NUM" +} + +wait_for_herdr_server() { + local i=0 + while [ "$i" -lt 20 ]; do + herdr_server_running && return 0 + i=$((i + 1)) + sleep 0.5 + done + return 1 +} + +start_herdr_server() { + if ! herdr_adapter_load; then + fix_report herdr-server failed "herdr and jq must both resolve before the server can be started" + return 1 + fi + if fm_backend_herdr_server_ensure "$HERDR_SESSION_NAME" >/dev/null 2>&1; then + fix_report herdr-server applied "started the herdr server for session $HERDR_SESSION_NAME" + return 0 + fi + fix_report herdr-server failed "the herdr server for session $HERDR_SESSION_NAME did not come up" + return 1 +} + +link_entrypoint() { + local want="${FM_ROOT_OVERRIDE:-}/bin/fm-remote-entrypoint.sh" + if ! mkdir -p "$(dirname "$ENTRYPOINT_LINK")" 2>/dev/null; then + fix_report entrypoint-link failed "cannot create $(dirname "$ENTRYPOINT_LINK")" + return 1 + fi + if ! ln -s "$want" "$ENTRYPOINT_LINK" 2>/dev/null; then + fix_report entrypoint-link failed "cannot create the symlink at $ENTRYPOINT_LINK" + return 1 + fi + fix_report entrypoint-link applied "linked $ENTRYPOINT_LINK to $want" +} + +apply_fixes() { + local i name value launch_agent_written=0 launch_agent_reloaded=0 remote_job_fixed=0 + repair_required_wrappers + i=0 + while [ "$i" -lt "${#CHECK_NAMES[@]}" ]; do + name=${CHECK_NAMES[$i]} + value=${CHECK_VALUES[$i]} + i=$((i + 1)) + case "$value" in fixable:*) ;; *) continue ;; esac + case "$name" in + remote-job-worker|remote-job-worker-loaded|remote-job-probe) + [ "$remote_job_fixed" -eq 0 ] || continue + remote_job_fixed=1 + fix_remote_job_worker || true + ;; + launchagent|launchagent-scope) + [ "$launch_agent_written" -eq 0 ] || continue + launch_agent_written=1 + write_launch_agent || continue + # A freshly written plist runs nothing until it is (re)loaded, and only + # an existing GUI session can hold it. + check_is_ok gui-session || continue + launch_agent_reloaded=1 + reload_launch_agent launchagent-loaded || true + ;; + launchagent-loaded) + [ "$launch_agent_reloaded" -eq 0 ] || continue + launch_agent_reloaded=1 + reload_launch_agent launchagent-loaded || true + ;; + herdr-server) + # On darwin the launch agent owns the server, so restart it through + # launchd rather than starting a stray one outside the Aqua session. A + # reload earlier in this same pass has already done that. + if [ "$PLATFORM" = darwin ] && [ -f "$LAUNCH_AGENT_PLIST" ] && check_is_ok gui-session; then + [ "$launch_agent_reloaded" -eq 0 ] || continue + launch_agent_reloaded=1 + reload_launch_agent herdr-server || true + continue + fi + start_herdr_server || true + ;; + entrypoint-link) link_entrypoint || true ;; + esac + done +} + +# --- report ----------------------------------------------------------------- + +if [ "$MODE" = worker-tool-probe ]; then + report_required_tools + [ "${#MISSING[@]}" -eq 0 ] + exit +fi + +printf 'mode=%s\n' "$MODE" +printf 'path=%s\n' "${PATH:-}" +if [ -n "${FM_ROOT_OVERRIDE:-}" ] && [ "${PATH%%:*}" = "$FM_ROOT_OVERRIDE/bin" ]; then + printf 'entrypoint=yes\n' +else + printf 'entrypoint=no\n' + printf 'note: not launched through the fixed remote entrypoint; the reported PATH is this caller environment.\n' >&2 +fi +printf 'platform=%s\n' "$PLATFORM" + +run_checks +if [ "$MODE" = fix ]; then + apply_fixes + # Re-derive every check from the host itself, so what prints below is the + # state after repair rather than the intent of a repair. + run_checks +fi + +if [ "${FM_REMOTE_JOB_ACTIVE:-}" = 1 ] || ! remote_job_identity_ok; then + report_required_tools +else + report_required_tools_from_worker +fi +for tool in "${OPTIONAL_TOOLS[@]}"; do + if resolved=$(command -v "$tool" 2>/dev/null); then + printf 'optional %s=%s\n' "$tool" "$resolved" + else + printf 'optional %s=absent\n' "$tool" + fi +done + +GAPS=() +i=0 +while [ "$i" -lt "${#CHECK_NAMES[@]}" ]; do + printf 'check %s=%s\n' "${CHECK_NAMES[$i]}" "${CHECK_VALUES[$i]}" + case "${CHECK_VALUES[$i]}" in + fixable:*|human:*) GAPS+=("$i") ;; + esac + i=$((i + 1)) +done +for i in ${GAPS[@]+"${GAPS[@]}"}; do + [ -z "${CHECK_ACTIONS[$i]}" ] || printf 'action: %s: %s\n' "${CHECK_NAMES[$i]}" "${CHECK_ACTIONS[$i]}" +done + +if [ "${#MISSING[@]}" -gt 0 ]; then + printf 'error: required tools do not resolve on the remote runtime PATH: %s\n' "${MISSING[*]}" >&2 + printf 'fix: install each one where it resolves on the path reported above, or put a wrapper script for it in %s/.local/bin, which is always on that PATH.\n' "${HOME:-~}" >&2 + printf 'fix: tools in an unselected nvm version or outside the discovered asdf or mise paths need an absolute wrapper; see docs/remote-secondmates.md for the wrapper recipe.\n' >&2 +fi +if [ "${#MISSING[@]}" -gt 0 ] || [ "${#GAPS[@]}" -gt 0 ]; then + NAMES= + for i in ${GAPS[@]+"${GAPS[@]}"}; do + NAMES="${NAMES:+$NAMES }${CHECK_NAMES[$i]}" + done + printf 'error: this host is not ready for a remote second mate%s\n' "${NAMES:+; unresolved: $NAMES}" >&2 + exit 1 +fi +printf 'ok: remote second-mate readiness confirmed on this host\n' diff --git a/bin/fm-remote-entrypoint.sh b/bin/fm-remote-entrypoint.sh new file mode 100755 index 00000000000..eb42a4afd77 --- /dev/null +++ b/bin/fm-remote-entrypoint.sh @@ -0,0 +1,149 @@ +#!/usr/bin/env bash +# Fixed remote entrypoint for bin/fm-on.sh. +# +# Install this tracked file as fm-remote-entrypoint.sh on the remote account's +# non-interactive SSH PATH. It accepts protocol metadata plus a base64-encoded +# NUL argv stream, validates one genuine tracked executable in /bin/fm-*.sh, +# then stages it for the Firstmate-owned remote job worker. It never accepts a +# shell command string. +# +# The readiness-owning fm-remote-doctor.sh runs in this plain SSH bootstrap so +# check mode can inspect worker gaps without changing them and --fix can repair +# them. Every other command is staged after the worker is ready. On Darwin, a +# missing Aqua session fails before staging with the doctor-actionable +# console-login diagnostic. Linux uses the same queue and worker shape without +# an Aqua requirement. +# +# stdin is captured as bounded job input. The completed worker result is relayed +# with stdout and stderr kept separate and its exit status preserved. An SSH +# disconnect remains unknown completion to fm-on.sh, which preserves OpenSSH's +# exit 255 behavior. The shared library header owns job fields, bounds, PATH, +# LaunchAgent contract, and worker environment. +set -eu + +PROTOCOL=1 +DOCTOR_SHA256=7bb13d9fad8455978bf109d4681a3aa3cb170565c8a74be4ec7b520427db14c2 +SCRIPT_DIR=$(CDPATH='' cd "$(dirname "${BASH_SOURCE[0]}")" && pwd -P) + +# shellcheck source=bin/fm-remote-job-lib.sh +. "$SCRIPT_DIR/fm-remote-job-lib.sh" + +die() { printf 'error: %s\n' "$1" >&2; exit "${2:-64}"; } + +base64_decode_to() { # + local encoded=$1 destination=$2 + if printf '%s' "$encoded" | base64 --decode > "$destination" 2>/dev/null; then return 0; fi + if printf '%s' "$encoded" | base64 -D > "$destination" 2>/dev/null; then return 0; fi + return 1 +} + +decode_text() { #