diff --git a/.agents/skills/quota-array-dispatch/SKILL.md b/.agents/skills/quota-array-dispatch/SKILL.md index 11b84058125..24c0e44de57 100644 --- a/.agents/skills/quota-array-dispatch/SKILL.md +++ b/.agents/skills/quota-array-dispatch/SKILL.md @@ -2,7 +2,8 @@ name: quota-array-dispatch description: >- Agent-only decision procedure for resolving a matched crew-dispatch profile - array from current quota-axi output, including effective headroom and usable-runway evidence. + array from quota-axi's default TOON, ranking by spendPriority after three + orthogonal gates. Load when a dispatch rule or default resolves to more than one profile candidate. user-invocable: false metadata: @@ -14,43 +15,46 @@ metadata: This skill is the single owner of the completion-aware profile-array selection procedure. `AGENTS.md` section 4 owns the always-loaded intake boundary, load trigger, malformed-config refusal, every-candidate accounting, and strongest-reasoning/tie safety rules. `harness-adapters` owns harness verification, model/provider discovery, and effort fallback. -`quota-axi` remains data-only, reports whatever granularity the vendor supplies, and never recommends, selects, ranks, or infers a route. +`quota-axi` remains data-only: it publishes `spendPriority` as a comparable scalar and never recommends, selects, ranks, or infers a route. Do not add a daemon, opaque composite score, routing wrapper, hard-coded model-specific policy, or producer-side route recommendation. Deterministic shell owns only schema, configuration, and version validation plus concrete spawn safeguards; every model-to-provider, provider-to-credential, and quota-applicability relation is yours to establish transparently and to show your evidence for. -## Collect facts +## Read the default TOON -Run `quota-axi --json` once per intake and reuse that snapshot for every candidate. -Do not take a second snapshot to settle a candidate, and read `quota-axi auth --json` when a candidate's credential surface is in question. -For each candidate, preserve explicit `harness`, `model`, and `provider`; `harness-adapters` owns identity, and model/provider never infer harness: +Start each intake by running `quota-axi` once with no `--json`, and reuse that TOON for every candidate. +Post-consolidation quota-axi (the floor owned by `bin/fm-quota-axi-lib.sh`) puts `spendPriority` in the default `quota[]` block beside `effectivePercentRemaining`, `runway`, `confidence`, `limitedBy`, and `resetsAt`. +Sparse `exhaustion[]` carries finite-runway seconds only for `projected_exhaustion` and `exhausted_now`. +Sparse `attention[]` names auth, stale, and unmeasurable facts. +`spendPriority` is THE quota-perspective ranker. +It already computes the economics that older instructions reconstructed by hand from headroom, pace, reserve, and window-id lists; do not recompute those. +Do not read `--json` on the normal path, and do not reach for `--full` to rebuild that economics. -- task/profile fit and required reasoning class -- applicable effective headroom (`effectivePercentRemaining`) from the established provider/model scope -- usable runway status, `usableRunwaySeconds`, `projectedExhaustedAt`, `limitingWindowId`, `projectionConfidence`, `projectionBasis`, and any `unmeasurableWindowIds` -- the task-completion horizon and the evidence and confidence used to estimate it -- effective pace, signed reserve per window, and worst reserve (`worstReservePercentPoints` or minimum signed reserve) for later diagnostic tie-breaking -- schema notes when runway or pace fields are absent +After reading the TOON, fall back to one `quota-axi --json` call only when that TOON is genuinely ambiguous for the decision, or when the installed quota-axi is somehow below the floor so its TOON lacks `spendPriority`. +Ambiguous means a candidate's `spendPriority` is the literal `unknown` or unmeasurable, a real tie still needs extra evidence, or a candidate's eligibility is unclear from `quota[]` plus `attention[]`. +The fallback therefore has an explicit TOON-then-JSON call sequence; reuse its JSON result and do not take any further quota snapshots. +Below-floor is rare: bootstrap enforces `FM_QUOTA_AXI_MIN` and normally reports `MISSING` before dispatch; if an intake somehow reaches an older build whose TOON lacks `spendPriority`, use the defensive `--json` fallback rather than treating the missing scalar as healthy. +`--json` is a defensive belt, not a habit; never reach for it because it feels more complete. +Read `quota-axi auth --json` only when a candidate's credential surface is in question. -Stale raw windows are diagnostic, never headroom or fabricated runway. -Grok's `credits.remaining` is a prepaid balance unrelated to `percentRemaining`; never read it as exhaustion. -Read all windows named by `boundedBy`, `limitingWindowIds`, `aheadWindowIds`, `behindWindowIds`, `onPaceWindowIds`, `unknownWindowIds`, and `unmeasurableWindowIds`. -The compact default output intentionally omits numeric reserve, while `--json` and `--full` retain reserve diagnostics. +For each candidate, preserve explicit `harness`, `model`, and `provider`; `harness-adapters` owns identity, and model/provider never infer harness. -## Establish the provider relation before reading quota +## Three gates, then spendPriority + +Apply the three cheap orthogonal gates first. +`spendPriority` ranks only among candidates that pass all three. +It cannot override a hard-gate failure, and it is never hidden inside a new composite score. + +### 1. Eligibility Deterministic shell must never map a model to a provider, a provider to a credential store, or a name prefix to a family. You establish those relations yourself, in the open, from the candidate's own authoritative catalog (`harness-adapters` owns the per-harness discovery surface) plus the one intake snapshot. -Name the evidence for each relation you assert so the conclusion is inspectable. - -1. Confirm the catalog lists the candidate's model and record the provider family it reports. - A model the authoritative catalog does not list is concrete contradictory evidence: block that candidate and quote the catalog result. -2. Apply quota at the granularity the vendor actually supplies. - A provider-level or `all_models`/`all_products` scope bounds every model you established in that family, including one with no window of its own. - A named-model or named-product scope is an additional bound for that model alone and is irrelevant to every other model in the family. - Read `quotaSemantics.description`, which states the vendor's own bounding rule. -3. Record what remains unknown instead of converting it into a verdict. -## Authentication is scoped to the selected surface +Confirm the catalog lists the candidate's model and record the provider family it reports. +A model the catalog does not list is concrete contradictory evidence: block that candidate and quote the catalog result. +Apply quota at the granularity the vendor actually supplies. +A provider-level or `all_models`/`all_products` scope bounds every model you established in that family, including one with no window of its own. +A named-model or named-product scope is an additional bound for that model alone. +Match the candidate to its `quota[]` row by that established provider and scope; a stale, auth-required, or unmeasurable scope is named in `attention[]` instead of a fabricated number. A candidate authenticates through its own tuple's surface; another harness's CLI can never gate it, and `harness=pi` with `model=xai/grok-*` is Pi using xAI rather than the standalone Grok CLI. `quota-axi auth --json` lists each provider's credential sources independently, so read the one source the candidate actually uses rather than collapsing a provider to a single status. @@ -59,8 +63,8 @@ A Pi-hosted family may authenticate through the vendor's own store with no `pi:` Uncertainty and ineligibility are different findings: -- No model-level window, no matching auth source, an absent `state.authStatus`, an unmeasurable or `unknown` scope, or a surface quota-axi does not model at all is disclosed uncertainty. - Keep the candidate eligible, state the unknown, and prefer known sustainable evidence when otherwise comparable. +- No model-level window, no matching auth source, an unmeasurable or `unknown` scope, or a surface quota-axi does not model at all is disclosed uncertainty. + Keep the candidate eligible, state the unknown, and prefer known viable evidence when otherwise comparable. - An expired credential is a short-lived session token the owning vendor renews on next use, not a sign-out. - Only concrete contradictory evidence blocks: an authoritative catalog proving the model unsupported, or proof that the credential the candidate actually selects is unusable. - Reserve login wording for that proven-unusable case, and name the harness, model, surface, and evidence. @@ -69,45 +73,45 @@ When a credential's local classification is the only thing standing between a ca `bin/fm-vendor-auth-probe.sh` is the only approved vendor-credential probe; its `--help` owns the registered probes and mechanics. It takes no harness, model, or provider and returns a fact, not a route: only `authenticated` and `unauthenticated` are ground truth, while `indeterminate`, `timeout`, and `unavailable` establish nothing and must never be read as either outcome. Never launch a vendor CLI yourself, and never probe a credential store the candidate does not use. +Grok prepaid `credits` are unrelated to paid-window headroom; never read them as exhaustion. + +Malformed configuration is an actionable error, not a candidate to rank around. + +### 2. Reasoning-class fit + +Keep only candidates that meet the required reasoning class for this task (a simple bug fix versus very-difficult design). +Never use `spendPriority` or remaining quota to silently replace that class. +When every remaining candidate is tight, dispatch inside the strongest-reasoning class if one of those candidates can proceed, or stop and report that the strongest-class choice cannot proceed rather than downgrading it to spend or conserve quota. + +### 3. Runway feasibility floor + +Known runway that will not last until the inspectable likely-completion horizon fails this gate, even when that candidate has the highest `spendPriority`. +Read `runway` from the `quota[]` row: `through_reset` passes this generic feasibility floor because the window reaches its refill without exhausting; never compare its `resetsAt` with the completion horizon as though reset were an exhaustion deadline. +`exhausted_now` is zero, and `projected_exhaustion` uses the matching `exhaustion[]` row's `usableRunwaySeconds`. +A high `spendPriority` on a nearly empty window that will exhaust soon must not route into a mid-task stall. +Unknown or unmeasurable runway stays eligible with disclosed uncertainty and is never assumed to pass. +Do not invent a generic percentage floor, and honor an explicit captain floor for a candidate when one exists. + +## Rank by spendPriority + +Among candidates that pass all three gates, pick the highest known `spendPriority`. +A higher known scalar is better: positive means paid allowance is on track to reach reset unused, `0` is exact utilization, and negative means overdrawn against the reset clock. +Rank only from comparable known scalars. +Never treat absent, `unknown`, or unmeasurable `spendPriority` as zero or as healthy; `0` means exact utilization, a different claim from unknown. +An unknown `spendPriority` keeps the candidate eligible with disclosed uncertainty. +Prefer known viable evidence when otherwise comparable. +After the permitted TOON-to-JSON fallback, escalate to Firstmate instead of routing if no candidate can be ranked or runway uncertainty prevents proving the feasibility floor for any candidate that could be selected. +Never resolve that terminal uncertainty by treating unknown as healthy or by choosing arbitrarily. +Show the scalar or the literal `unknown` in the rationale; do not hide it in a score. + +Do not compare headroom against runway by hand. +Do not use pace or signed reserve as a later tie-break layer. +Do not read `aheadWindowIds`, `behindWindowIds`, `onPaceWindowIds`, `limitingWindowIds`, or other window-id lists to reconstruct what `spendPriority` already computed. + +Genuine ties: stop and report every tied candidate for captain choice. +Do not select by array order, harness name, or another arbitrary identity ordering. +Report duplicate concrete profiles as a configuration error. -## Pace semantics - -`reservePercentPoints = percentRemaining - timeRemainingPercent`. -Negative reserve means usage is ahead of reset pace and creates conservation pressure. -Positive reserve means usage is behind reset pace. -`on_pace` is neutral. -Conservation pressure is present for effective pace status `ahead`, effective pace status is `mixed` and any `aheadWindowIds` remain, or a bounding window is `ahead`. -`unknown` is valid explicit uncertainty from quota-axi, not parser failure or permission to assume health. - -## Selection order - -Apply only among candidates satisfying required fit and strongest reasoning class. -Never use headroom, runway, pace, or reserve to silently replace that reasoning class. - -1. Concrete contradictory evidence or malformed configuration: stop and report the tuple and that evidence. - Unmeasurable quota, a missing model-level window, an absent runway field, and a credential surface quota-axi does not model are uncertainty, never this rule. -2. Honor any explicit captain instruction that sets a floor for that candidate before the generic comparison. - Do not invent a generic percentage floor or treat a low percentage as an automatic failure. -3. Keep the strongest-reasoning class when every candidate is tight or completion evidence is poor. - Dispatch inside that class when a candidate can proceed, or report that its strongest-class choice cannot proceed rather than downgrading it to conserve quota. -4. Compare comparable-fit candidates on their applicable effective headroom and usable runway. - Eliminate a candidate only when another candidate Pareto-dominates it on both dimensions, with at least one dimension strictly better. - Establish dominance only from comparable known evidence, never by treating absent, `unknown`, or unmeasurable headroom or runway as zero or as a healthy value. -5. Prefer supported runway evidence that projects availability through the inspectable likely-completion horizon. - Known evidence that does not reach that horizon is inferior to known evidence that does, even when its signed reserve is less negative. - Preserve projection confidence and basis, the limiting window, and the horizon estimate in the rationale rather than hiding them in a score or model-specific heuristic. -6. Resolve remaining uncertainty explicitly. - An authenticated candidate with unknown or unmeasurable headroom or runway stays eligible and cannot be silently excluded or assumed sustainable. - Prefer known viable evidence when otherwise comparable, and report uncertainty or ask the captain when it still prevents a justified choice. -7. Use pace and signed reserve only as later diagnostic tie-break evidence among candidates still unresolved after headroom, runway, likely-completion viability, and uncertainty. - Pace and reserve never rescue a clearly inferior completion prospect. - Do not collapse these facts into an opaque composite score. -8. Older schemas or absent runway/pace fields: do not crash, fabricate runway or pace, treat absence as healthy, or silently exclude a candidate. - State which evidence is unavailable, retain the candidate, and apply only the comparisons the snapshot supports. -9. Genuine ties: stop and report every tied candidate for captain choice. - Do not select by array order, harness name, or another arbitrary identity ordering. - Report duplicate concrete profiles as a configuration error. - -Account for every candidate visibly before selecting or escalating, naming its catalog evidence, provider relation, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, effective headroom, usable runway, likely-completion reasoning, and later pace or reserve evidence when used. +Account for every candidate visibly before selecting or escalating, naming its catalog evidence, provider relation, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, `spendPriority`, and runway-versus-horizon result. A blocked credential report must name `harness`, `model`, authentication surface, and concrete failure evidence; never emit a bare `Grok unauthenticated` statement. Never conclude with an unexplained "best quota" label. diff --git a/.agents/skills/stuck-crewmate-recovery/SKILL.md b/.agents/skills/stuck-crewmate-recovery/SKILL.md index cf741b9d95f..b9b94b27d43 100644 --- a/.agents/skills/stuck-crewmate-recovery/SKILL.md +++ b/.agents/skills/stuck-crewmate-recovery/SKILL.md @@ -23,7 +23,7 @@ The target window's harness is recorded as `harness=` in `state/.meta`. This procedure covers ordinary `kind=ship` and `kind=scout` direct reports. Load `secondmate-provisioning` instead for `kind=secondmate` recovery. -For a REMOTE secondmate, `fm-crew-state`'s `unknown`/`worktree gone` and `fm-send`'s `remote send failed`/`delivery unconfirmed` verdicts are unreliable and routinely false-negative; do not conclude the mate is dead or the send failed from those alone, confirm against the actual remote pane first. +For a REMOTE secondmate, `fm-crew-state` and `fm-peek` read the actual remote endpoint over `fm-on.sh`, and `fm-send` reports a delivered-with-pending-confirmation steer as delivered (their headers own the contracts); an `unknown-remote` read or unreachable-host failure means the remote state could not be read, never that the mate is dead or the send failed. Recover a genuinely stuck remote mate only through `bin/fm-spawn.sh --secondmate`, never raw herdr pane close/kill surgery, which strands the endpoint binding. Treat the digest's endpoint result as a presence signal, not proof that the task's work or validation run is gone. diff --git a/AGENTS.md b/AGENTS.md index 0f0a8e75c41..4aff87c66cd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Hard rules, in priority order: If work failed, say so plainly with the evidence. You may maintain this repo's private operational state directly. -Shared tracked material is `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `fork-divergences.json`, `.github/workflows/`, `bin/`, `.agents/skills/`, and public `skills/`. +Shared tracked material is `AGENTS.md`, `GROK_BOT.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `fork-divergences.json`, `.github/workflows/`, `bin/`, `.agents/skills/`, and public `skills/`. When any crewmate is live, delegate changes to shared tracked material rather than competing with supervision; when the fleet is empty, firstmate may change it directly. This repo is a shared template, while `.env`, `data/`, `state/`, `config/`, `projects/`, and `.no-mistakes/` are captain-private and gitignored. Ship shared tracked changes through this repo's no-mistakes pipeline and PR path, with the same merge authority as any other project. @@ -102,8 +102,8 @@ If static `config/crew-harness` or `config/secondmate-harness` names an unverifi `docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation. When dispatch profiles exist, consult them at every crewmate or scout intake and pass the resolved concrete profile required by `fm-spawn`. Routing precedence is an explicit per-task captain override, then the best-fit configured rule, then the configured default, then the static crewmate harness. -Firstmate alone resolves a matched profile array: run `quota-axi --json` at that intake, evaluate every configured candidate against that current output, and choose with inspectable effective headroom and usable runway, using pace and reserve only later when needed. -Account for every candidate with the catalog evidence, provider relationship, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, and the headroom, runway, and later pace or reserve evidence used in selection; never omit a candidate, guess, fall back silently, or call the result quota-informed without them. +Firstmate alone resolves a matched profile array: begin with `quota-axi`'s default TOON at that intake, using the skill's narrow TOON-then-`--json` fallback only for genuine ambiguity, evaluate every configured candidate against that current output, and choose with inspectable `spendPriority` as the one quota-perspective ranker after the skill's eligibility, reasoning-class, and runway-feasibility gates. +Account for every candidate with the catalog evidence, provider relationship, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, and the spendPriority and runway evidence used in selection; never omit a candidate, guess, fall back silently, or call the result quota-informed without them. Establish model support and provider family from that harness's own authoritative catalog, then read `quota-axi` at the granularity the vendor actually supplies: provider-level or all-model evidence applies to every model established in that family, and a named-model window bounds only that model. Missing model-level quota, a missing authentication source, unmeasurable headroom, or unmodeled authentication is disclosed uncertainty that keeps a candidate eligible, never a credential or login escalation. Only concrete contradictory evidence blocks a candidate, such as an authoritative catalog proving the model unsupported or proof that the credential selected for that surface is unusable; never infer a credential store, provider family, or quota mapping from a harness, model, or source name, and never launch another harness's CLI to judge a candidate. @@ -111,7 +111,7 @@ Preserve malformed profile configuration as an actionable error rather than sele When every candidate is tight, preserve the captain's strongest-reasoning class rather than silently downgrading it solely to conserve quota; stop and report the tight choice if that class cannot proceed. Break genuine evidence ties without array-order or harness bias. `quota-axi` owns how model or product windows relate to bounding account windows and remains data-only. -Load `quota-array-dispatch` before choosing among a matched profile array; that skill is the single owner of the completion-aware selection procedure. +Load `quota-array-dispatch` before choosing among a matched profile array; that skill is the single owner of the TOON-first spendPriority selection procedure. The generic effort fallback and its precedence are owned by `harness-adapters`: explicit captain and standing configured effort win; otherwise use low for well-understood explicit work, xhigh for ambiguous investigation or design, intermediate levels proportionally, and never max without explicit captain preference. Do not add model-specific versions of that policy. @@ -440,7 +440,7 @@ These skills are not captain-invocable; load them only at their precise triggers - `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `NETWORK_CHECKS:`, `PR_CHECK_MIGRATION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`); silence and `BOOTSTRAP_INFO:` need no load. - `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report. - `ask-user-authority` - load before deciding any ask-user finding, regardless of the project's `yolo` posture. -- `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi output. +- `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi default TOON. - `harness-adapters` - load before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter. - `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata. - `project-management` - load before adding, creating, removing, or initializing a project. diff --git a/GROK_BOT.md b/GROK_BOT.md new file mode 100644 index 00000000000..e875c3546a6 --- /dev/null +++ b/GROK_BOT.md @@ -0,0 +1,47 @@ +You are Firstmate: the single agent the captain talks to. +They bring you everything; you make sure it gets done. + +Other bots are your crewmates: persistent and role-based, each holding a stable charter - e.g. one for the inbox, one for documents like PDFs and decks, one for research. +Before signing on a new crewmate, check whether an existing one already covers a related charter: if a charter matches or highly overlaps, reuse that crewmate; if the overlap is only limited, sign on the new crewmate and clarify the distinction in both crewmates' charters. +Sign on a genuinely new crewmate only when no existing one fits. +When you sign one on, write into its charter that it reports its outcomes and blockers back to you (Firstmate), never to the captain directly - the captain only ever talks to you. +Delegate by messaging a crewmate; it wakes, does the work, and messages you back. + +Default to handing work off. +If a job is more than one tool call, especially computer or browser work or anything that will take minutes, give it to the crewmate whose charter fits. +Do not keep that grind in this chat because you already have a login, a token, or an open page. +The computer is shared across the crew. +Browser logins persist for every bot. +A login on your screen is not a reason to do the work yourself. +Secrets are per-bot. +They do not propagate to the crew. +If a crewmate needs a credential, tell the crewmate to request it and then tell the captain to give that secret to that bot on a secure card. +Do not keep the secret and do the work yourself. +Do not paste or forward secrets in chat. +After the captain has given the secret to that bot, hand the task off and wait for the outcome. + +Software and code go through a crewmate, never through you directly: sign on a crewmate per project or project area - once the captain has expressed how its charter should be set - and let that crewmate drive the code work with cursor cloud agents. +You never call a cursor cloud agent yourself. + +Don't reach for subagents. +Needing one means the work is substantial, which means it belongs with a crewmate, not with you. +Subagents are a tool for crewmates to break down their own work. + +Mark every task you hand off as coming from you, with a short task id, and ask for the outcome back against that id - so the crewmate routes its result and any blockers to you rather than just handling them in its own chat, and you can match a reply to the right task. +The marker is visible in the chat; that's fine. + +Work asynchronously. +Delegating doesn't block you - a crewmate replies on a later turn and shows up in this chat. +So hand off, tell the captain what's under way, and relay each result as it lands. +Reserve a priority send for when something must interrupt a crewmate's current task. + +When you notice crewmates making mistakes or working inefficiently, update their description to refine their behavior so your crew does better next time. + +How you talk. +Address the captain as "captain" at least once in every reply - always, even when the news is bad ("Captain, that didn't work..."). +Let light nautical seasoning land only when it fits naturally - an occasional "aye", "on deck", "shipshape", "under way", "ahoy" - never letting it crowd out the substance, and drop it entirely for bad news or serious findings. +Speak in outcomes and consequences, not internal mechanics. + +Keep it simple for the captain. +Focus on communicating outcomes, not mechanics. +They scale by talking only to you; protect that. diff --git a/bin/fm-crew-state.sh b/bin/fm-crew-state.sh index 2cb290373cb..df627b487f2 100755 --- a/bin/fm-crew-state.sh +++ b/bin/fm-crew-state.sh @@ -16,10 +16,17 @@ # fixed mapping logic, no heuristics and no LLM. Output is one stable, parseable, # token-tight line firstmate can read every heartbeat: # -# state: · source: · +# state: · source: · # # Logic, in order: -# 1. Resolve worktree + backend target + kind from state/.meta. +# 1. Resolve worktree + backend target + kind from state/.meta. A meta +# recording remote_host= is a remote secondmate: its worktree and endpoint +# live on that host, so the local worktree and pane reads are skipped and +# the remote host is asked for the endpoint's recovery-grade state +# (fm-on.sh + fm-remote-secondmate-control.sh state). alive falls through +# to the routed status log; dead/missing report the remote verdict; an +# unreachable or unreadable remote reports unknown-remote, never a false +# gone/dead. # 2. Matching no-mistakes run for this crew's branch AND current code identity, # active or terminal (from `axi status`, or the coarse `no-mistakes runs` # fallback)? Branch name alone is not enough: a historical run on a reused @@ -101,10 +108,13 @@ meta_value() { # WT=$(meta_value worktree) KIND=$(meta_value kind) HARNESS=$(meta_value harness) +REMOTE_HOST=$(meta_value remote_host) [ -n "$KIND" ] || KIND=ship -# A torn-down (or never-created) worktree has no current state to read. -if [ -z "$WT" ] || [ ! -d "$WT" ]; then +# A torn-down (or never-created) worktree has no current state to read. A +# remote secondmate's recorded worktree is a path on ITS host, so the local +# probe proves nothing for it - the remote arm below reads the true source. +if [ -z "$REMOTE_HOST" ] && { [ -z "$WT" ] || [ ! -d "$WT" ]; }; then emit unknown none "worktree gone (torn down?)" fi @@ -138,6 +148,45 @@ map_log_state() { # LOG_LINE=$(log_last_line || true) LOG_VERB=$(status_line_verb "$LOG_LINE") +# --- remote secondmate: the true source is the remote endpoint --------------- +# A remote mate's recorded worktree and backend target live on its own host, so +# the local worktree probe above and the local pane reads below would misreport +# a healthy remote mate as gone or dead. Ask the remote host for the endpoint's +# recovery-grade state over the same fm-on.sh transport fm-send uses, then read +# current activity from the routed status log exactly as for a local +# secondmate (an idle endpoint is healthy for a secondmate either way). An +# unreachable host or unreadable endpoint is reported as unknown-remote - +# explicitly NOT proof of death - so a transport blip never reads as a torn +# down or dead mate; only the remote host's own dead/missing verdict may say +# the endpoint is actually gone. +if [ -n "$REMOTE_HOST" ]; then + if ! REMOTE_STATE=$(FM_HOME="$FM_HOME" "$SCRIPT_DIR/fm-on.sh" "$ID" \ + fm-remote-secondmate-control.sh state "$ID" < /dev/null 2>/dev/null); then + REMOTE_STATE= + fi + REMOTE_STATE=$(printf '%s\n' "$REMOTE_STATE" | tail -1) + case "$REMOTE_STATE" in + alive) + if [ -n "$LOG_VERB" ]; then + LOG_STATE=$(map_log_state "$LOG_LINE") + if [ "$LOG_STATE" != unknown ]; then + emit "$LOG_STATE" status-log "$(status_line_note "$LOG_LINE")${SEP}remote endpoint alive on $REMOTE_HOST" + fi + fi + emit unknown remote-endpoint "alive on $REMOTE_HOST (an idle secondmate is healthy)" + ;; + dead|missing) + emit unknown remote-endpoint "remote endpoint $REMOTE_STATE on $REMOTE_HOST" + ;; + '') + emit unknown remote-endpoint "unknown-remote: $REMOTE_HOST unreachable or endpoint unreadable (not proof of death)" + ;; + *) + emit unknown remote-endpoint "unknown-remote: endpoint state '$REMOTE_STATE' on $REMOTE_HOST (not proof of death)" + ;; + esac +fi + # pane_readable is consulted ONLY in the no-run fallback below. The run-step path # stays authoritative regardless of pane liveness - judge by the run-step, not the # shell - so a finished crew whose endpoint has closed still reports its run-step diff --git a/bin/fm-inactive-reconcile.sh b/bin/fm-inactive-reconcile.sh index 79ece97a0a7..30d451db5ae 100755 --- a/bin/fm-inactive-reconcile.sh +++ b/bin/fm-inactive-reconcile.sh @@ -9,9 +9,17 @@ # not a watcher, daemon, PR poll, or forge client of its own. # `scan` evaluates at most once per FM_INACTIVE_RECONCILE_SECS (default 900, # valid 60..1800) per home, except that --startup performs the same cheap scan -# immediately during a locked session start. Each scan has an aggregate -# FM_INACTIVE_RECONCILE_BUDGET_SECS bound (default 10, valid 1..30) and resumes -# after its last visited child on the next scan. +# immediately during a locked session start. Each scan uses an aggregate +# FM_INACTIVE_RECONCILE_BUDGET_SECS deadline (default 10, valid 1..30) and +# resumes after its last visited child on the next scan. +# The scan enforces that budget itself through a whole-second deadline, and the +# first due child of every scan is always visited with at least a one-second +# state-read bound: whole-second arithmetic can otherwise round a small budget +# to zero mid-scan, and an invocation that exits having visited nothing would +# advance the durable cursor past a child it never examined. A process-group +# kill one second after the budget remains as a backstop for a scan wedged in +# an unbounded wait (for example a live-held wake-queue lock), so the clean +# deadline path is not racing its own backstop. # # It considers only a direct ordinary crewmate whose newest meta, status, or # turn-ended mtime is older than that interval and whose last status is not @@ -377,8 +385,13 @@ reconcile_direct_child() { # return "$rc" } +# SCAN_FIRST_VISIT_PENDING is armed by scan() before its passes. The deadline +# below is whole-second arithmetic, so a small budget can quantize to zero +# between the deadline computation and these checks; without the guaranteed +# first visit, such a scan would return 3 having examined no child at all while +# write_scan_marker had already advanced the cursor past the skipped child. scan_pass() { # - local cursor=$1 range=$2 deadline=$3 self=${4:-} meta id remaining rc + local cursor=$1 range=$2 deadline=$3 self=${4:-} meta id remaining rc first for meta in "$STATE"/*.meta; do [ -f "$meta" ] || continue id=$(basename "$meta" .meta) @@ -387,9 +400,19 @@ scan_pass() { # after) [ -z "$cursor" ] || [[ "$id" > "$cursor" ]] || continue ;; through) [ -n "$cursor" ] && [[ "$id" > "$cursor" ]] && continue ;; esac - [ "$(date +%s)" -lt "$deadline" ] || return 3 + first=0 + if [ "${SCAN_FIRST_VISIT_PENDING:-0}" -eq 1 ]; then + first=1 + SCAN_FIRST_VISIT_PENDING=0 + fi + if [ "$first" -eq 0 ]; then + [ "$(date +%s)" -lt "$deadline" ] || return 3 + fi write_scan_marker "$id" || return 1 remaining=$((deadline - $(date +%s))) + if [ "$first" -eq 1 ] && [ "$remaining" -lt 1 ]; then + remaining=1 + fi [ "$remaining" -gt 0 ] || return 3 reconcile_direct_child "$id" "$meta" "$self" "$remaining" || { rc=$? @@ -420,6 +443,7 @@ scan() { fi fi deadline=$(( $(date +%s) + FM_INACTIVE_RECONCILE_BUDGET_SECS )) + SCAN_FIRST_VISIT_PENDING=1 scan_pass "$cursor" after "$deadline" "$self" || rc=$? if [ "$rc" -eq 0 ] && [ -n "$cursor" ]; then scan_pass "$cursor" through "$deadline" "$self" || rc=$? @@ -461,7 +485,11 @@ case "$mode" in --startup) startup=1 ;; *) printf 'usage: fm-inactive-reconcile.sh scan [--startup]\n' >&2; exit 2 ;; esac - if fm_run_timed "$FM_INACTIVE_RECONCILE_BUDGET_SECS" "$0" _scan-locked "$startup"; then + # The scan's own whole-second deadline enforces the budget; this outer + # process-group kill is only the backstop for a scan wedged outside every + # bounded section (an unbounded lock wait), so it fires one second after + # the deadline instead of racing the clean bounded exit it exists to guard. + if fm_run_timed $((FM_INACTIVE_RECONCILE_BUDGET_SECS + 1)) "$0" _scan-locked "$startup"; then : elif [ "$?" -ne 124 ]; then exit 1 diff --git a/bin/fm-peek.sh b/bin/fm-peek.sh index 97d2ffe2d25..e3156f66ed4 100755 --- a/bin/fm-peek.sh +++ b/bin/fm-peek.sh @@ -3,6 +3,11 @@ # Usage: fm-peek.sh [lines=40] # may be an exact task id, a legacy fm- task label resolved # through this home's state/.meta, or an explicit backend target. +# A selector whose meta records remote_host= is a remote secondmate: its pane +# lives on that host, so the capture routes over fm-on.sh to the host-local +# capture (fm-remote-secondmate-control.sh), clamped to that command's +# 100-line cap. An unreachable host or unreadable endpoint fails loudly naming +# the host; the local backend adapters are never asked to read a remote target. set -eu SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -16,9 +21,25 @@ STATE="${FM_STATE_OVERRIDE:-$FM_HOME/state}" "$SCRIPT_DIR/fm-guard.sh" || true RAW_TARGET=$1 -T=$(fm_backend_resolve_selector "$RAW_TARGET" "$STATE") N=${2:-40} +REMOTE_META=$(fm_backend_meta_for_selector "$RAW_TARGET" "$STATE" 2>/dev/null || true) +if [ -n "$REMOTE_META" ] && [ -n "$(fm_meta_get "$REMOTE_META" remote_host)" ]; then + REMOTE_ID=${REMOTE_META##*/} + REMOTE_ID=${REMOTE_ID%.meta} + REMOTE_HOST=$(fm_meta_get "$REMOTE_META" remote_host) + case "$N" in ''|*[!0-9]*|0) N=40 ;; esac + [ "$N" -le 100 ] || N=100 + if ! FM_HOME="$FM_HOME" "$SCRIPT_DIR/fm-on.sh" "$REMOTE_ID" \ + fm-remote-secondmate-control.sh capture "$REMOTE_ID" "$N" < /dev/null; then + echo "error: could not read the remote pane of $REMOTE_ID on $REMOTE_HOST (host unreachable or endpoint unreadable; the mate is not thereby dead)" >&2 + exit 1 + fi + exit 0 +fi + +T=$(fm_backend_resolve_selector "$RAW_TARGET" "$STATE") + BACKEND=$(fm_backend_of_selector "$RAW_TARGET" "$T" "$STATE") EXPECTED_LABEL=$(fm_backend_expected_label_of_selector "$RAW_TARGET" "$STATE") diff --git a/bin/fm-quota-axi-lib.sh b/bin/fm-quota-axi-lib.sh index 7be4c99614c..1f59be67920 100644 --- a/bin/fm-quota-axi-lib.sh +++ b/bin/fm-quota-axi-lib.sh @@ -9,7 +9,7 @@ # turns a failing check into the operator-facing MISSING diagnostic, which is # what keeps an older build from reaching a dispatch intake at all. -FM_QUOTA_AXI_MIN=0.1.25 +FM_QUOTA_AXI_MIN=0.1.29 fm_quota_axi_compatible() { local timeout=${1:-} output parts major minor patch extra diff --git a/bin/fm-remote-secondmate-control.sh b/bin/fm-remote-secondmate-control.sh index 566ae095306..14576547894 100755 --- a/bin/fm-remote-secondmate-control.sh +++ b/bin/fm-remote-secondmate-control.sh @@ -188,6 +188,12 @@ cmd_send() { validate_id "$id" validate_home "$id" remote_endpoint_require "$id" + # fm-send's exit status is the delivery verdict the parent home acts on + # (0 = confirmed, 3 = delivered with the submit read-back unconfirmed, other + # nonzero = failed; see bin/fm-send.sh's header). The job worker, entrypoint, + # and ssh all preserve it, so no mapping may happen here: flattening exit 3 + # into a generic failure is exactly the false-negative the parent's remote + # send path exists to avoid. FM_HOME="$TARGET_HOME" FM_ROOT_OVERRIDE="$FM_ROOT" FM_STATE_OVERRIDE="$TARGET_HOME/state" \ "$SCRIPT_DIR/fm-send.sh" "$REMOTE_ENDPOINT_TARGET" "$message" } diff --git a/bin/fm-send.sh b/bin/fm-send.sh index c46c55a340f..1da45d86f46 100755 --- a/bin/fm-send.sh +++ b/bin/fm-send.sh @@ -15,6 +15,12 @@ # submit or reports an inconclusive send. If a swallowed Enter is positively # confirmed, fm-send exits NON-ZERO so the caller knows the steer did not land # instead of silently leaving an unsubmitted instruction. +# Exit status contract: 0 = submit confirmed (or, for a remote secondmate +# target, delivered with confirmation pending - see the remote paragraph); +# 3 = the text was typed into the live endpoint and Enter was sent, but the +# submit read-back stayed unconfirmed (verify the pane before any resend, and +# never re-type blindly); any other nonzero = the send failed and nothing may +# be assumed delivered. # Submission dispatches through the target's recorded backend; the tmux adapter # shares its composer/submit core with the away-mode daemon via bin/fm-tmux-lib.sh. # Tune with FM_SEND_RETRIES (default 3) / FM_SEND_SLEEP (0.4). @@ -37,6 +43,20 @@ # re-sending a recovery request for an already-open expectation so a second # record is not created. Direct unmarked captain input never creates one. # +# Remote secondmate delivery: the send crosses fm-on.sh to a host-local leg +# (bin/fm-remote-secondmate-control.sh cmd_send) that runs this same verified +# submit against the recorded remote Herdr pane and relays its exit status +# unchanged. A leg that delivered the text into the live verified pane but +# could not synchronously confirm the submit (exit 3 - typically a busy mate +# whose harness queues the steer and keeps rendering it) is reported here as +# DELIVERED with confirmation pending: fm-send prints a non-error notice, +# exits 0, marks the pending-reply expectation delivered, and closes any +# --resolve-key decisions. Empirically that pattern is a delivered steer, a +# resend duplicates the instruction, and the parent's pending-reply +# recovery/escalation still surfaces the rare genuinely lost request. Transport +# loss (ssh exit 255, completion unknown) and every real remote failure keep +# failing loudly with the remote leg's own stderr attached. +# # Decision closure (answerer-closes): pass --resolve-key (repeatable, # before the message) when this send answers an open keyed needs-decision: or # blocked: record in the target task's state/.status. After the submit is @@ -63,7 +83,9 @@ # in this home's status log per status_open_decisions (bin/fm-classify-lib.sh), or # an active captain hold for the target task. A key in neither is refused before # sending, so a mistyped key cannot deliver an answer while silently orphaning the -# decision. A failed or unconfirmed send never closes a key; a +# decision. A failed or unconfirmed send never closes a key (a remote +# delivered-with-pending-confirmation outcome counts as delivered - see the +# remote paragraph above); a # delivered answer whose closing append fails exits nonzero with the exact # manual close command, leaving the decision open to re-surface (the safe # direction). A send without the flag never closes anything: a routine steer, @@ -537,12 +559,27 @@ else # Type once, submit, verify. Only exact empty confirms delivery; every other # verdict preserves the loud refusal boundary. send_rc=0 + REMOTE_DELIVERY_NOTICE=0 if [ "$TARGET_BACKEND" = remote ]; then - if "$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send "$TARGET_REMOTE_ID" "$MESSAGE" < /dev/null >/dev/null; then + # The remote leg is this same script running host-locally against the + # recorded Herdr pane (cmd_send in fm-remote-secondmate-control.sh), so its + # submit verification IS the local one, and fm-on/the remote worker relay + # its exit status unchanged. Exit 3 is the delivered-unconfirmed contract + # (see this script's header) crossing the ssh boundary: the text reached + # the live verified pane and Enter was sent; only the synchronous read-back + # stayed unconfirmed. The remote stderr is held back and replayed only for + # a real failure, so a delivered outcome does not surface the inner leg's + # diagnostics as alarm. + remote_err=$("$SCRIPT_DIR/fm-on.sh" "$TARGET_REMOTE_ID" fm-remote-secondmate-control.sh send "$TARGET_REMOTE_ID" "$MESSAGE" < /dev/null 2>&1 >/dev/null) || send_rc=$? + if [ "$send_rc" -eq 0 ]; then + verdict=empty + elif [ "$send_rc" -eq 3 ]; then verdict=empty + send_rc=0 + REMOTE_DELIVERY_NOTICE=1 else - send_rc=$? verdict=send-failed + [ -z "$remote_err" ] || printf '%s\n' "$remote_err" >&2 fi elif verdict=$(fm_backend_send_text_submit "$TARGET_BACKEND" "$T" "$MESSAGE" "$retries" "$sleep_s" "$settle" "$EXPECTED_LABEL"); then : @@ -571,6 +608,19 @@ else echo "error: text not sent to $T ($TARGET_BACKEND send failed; tried $RESOLUTION_TRIED)" >&2 exit 1 ;; + pending) + # The text was typed into the live target and Enter was sent; only the + # submit read-back stayed unconfirmed (e.g. a busy harness queues the + # steer and keeps rendering it). That is not a proven failure, so never + # re-type the message: verify the pane instead. Exit 3 is the documented + # delivered-unconfirmed status, and the remote send leg above depends on + # it crossing the ssh boundary intact. + if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then + fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true + fi + echo "fm-send: text delivered to $T but submission is unconfirmed (verdict=pending; tried $RESOLUTION_TRIED); do not retype or blindly resend - verify with fm-peek.sh, then re-send '--key Enter' only if the composer still holds the text" >&2 + exit 3 + ;; *) if [ "$PENDING_REPLY_CREATED" = 1 ] && [ -n "$PENDING_REPLY_CORR" ]; then fm_pending_reply_discard_undelivered "$STATE" "$PENDING_REPLY_CORR" || true @@ -600,6 +650,12 @@ else fm_send_close_resolved_keys "$RESOLVE_ANSWER_TEXT" || exit 1 fm_send_feed_resolved_holds "$RESOLVE_ANSWER_TEXT" || exit 1 fi + # Remote delivered-with-pending-confirmation: the outcome above is treated as + # delivered (expectation marked, keys closed), and this one non-error notice + # carries the remaining nuance so nobody re-sends the steer. + if [ "$REMOTE_DELIVERY_NOTICE" = 1 ]; then + echo "fm-send: delivered to remote secondmate $TARGET_REMOTE_ID; the remote pane accepted the text and Enter, and only the synchronous submit confirmation is still pending. This is not a failure - do not resend; the pending-reply expectation stays armed." >&2 + fi # Submit landed with exact empty. Confirmation only proves the text was # accepted; the harness still needs a beat to spin up the # turn before its busy footer shows. Pause so an immediate peek catches the diff --git a/docs/configuration.md b/docs/configuration.md index 25c2e871a02..fe3f3c65e24 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -628,7 +628,7 @@ FM_POLL=15 # seconds between watcher poll cycles FM_HEARTBEAT=600 # base seconds between heartbeat scans; no-change heartbeats are absorbed while idle FM_HEARTBEAT_MAX=7200 # heartbeat backoff cap FM_INACTIVE_RECONCILE_SECS=900 # 60..1800-second watcher cadence and inactivity threshold; locked session start also scans immediately -FM_INACTIVE_RECONCILE_BUDGET_SECS=10 # 1..30-second aggregate bound per inactive-outcome scan +FM_INACTIVE_RECONCILE_BUDGET_SECS=10 # 1..30-second scan deadline; wedged-scan kill backstop follows one second later FM_CHECK_INTERVAL=300 # seconds between slow checks (authenticated merge polls, custom checks, or Relay dispatch) FM_CHECK_TIMEOUT=30 # seconds allowed per slow check script FM_PROCEVENT_MAX_OUTPUT_BYTES=1048576 # bound on one captured process-to-event result diff --git a/docs/documentation-audiences.json b/docs/documentation-audiences.json index 034831a5be2..848022b2a48 100644 --- a/docs/documentation-audiences.json +++ b/docs/documentation-audiences.json @@ -201,6 +201,10 @@ "path": "CONTRIBUTING.md", "audience": "maintainer-architecture" }, + { + "path": "GROK_BOT.md", + "audience": "public-product" + }, { "path": "README.md", "audience": "public-product" diff --git a/docs/remote-secondmates.md b/docs/remote-secondmates.md index 9b2cd38be46..b7020b57c15 100644 --- a/docs/remote-secondmates.md +++ b/docs/remote-secondmates.md @@ -172,6 +172,11 @@ Send routed requests normally: FM_HOME= bin/fm-send.sh fm- '' ``` +The [`fm-send.sh` header](../bin/fm-send.sh) owns the exact delivery-status contract. +When the verified remote endpoint accepts the text and Enter but synchronous submit confirmation remains pending, the primary reports the request as delivered rather than failed; do not resend it, because its pending-reply expectation remains armed. +`fm-peek.sh` and `fm-crew-state.sh` route remote-secondmate reads to the endpoint's host instead of consulting local worktree or backend state. +An unreachable or unreadable remote read is unknown, not evidence that the endpoint is dead. + Marked requests keep the existing correlation contract. The remote charter appends replies to `state/parent-replies.status` in the remote home. A process-event source performs a non-destructive, cursor-anchored delta read, fetches only referenced `data/*.md` documents through the confined reader, mirrors every content-bearing line at most once into the primary status channel, and does not carry blank separators. @@ -236,6 +241,9 @@ The lifecycle test covers seeding a registered project that this machine has nev ```sh bin/fm-test-run.sh tests/fm-on.test.sh +bin/fm-test-run.sh tests/fm-send-remote-delivery.test.sh +bin/fm-test-run.sh tests/fm-peek-remote.test.sh +bin/fm-test-run.sh tests/fm-crew-state.test.sh bin/fm-test-run.sh tests/fm-remote-job.test.sh bin/fm-test-run.sh tests/fm-remote-doctor.test.sh bin/fm-test-run.sh tests/fm-project-origin.test.sh diff --git a/docs/tmux-backend.md b/docs/tmux-backend.md index 4d8c3e75feb..c2acead0c2f 100644 --- a/docs/tmux-backend.md +++ b/docs/tmux-backend.md @@ -85,7 +85,7 @@ The supervisor guard selects only the detected primary harness's signature rathe It types a message once and retries Enter only until the composer clears. Only a proven empty composer is a positive delivery acknowledgement. Text left in established structure remains `pending`, text in ambiguous structure remains unproven, and unreadable or unsafe state remains unknown. -`fm-send.sh` reports every unconfirmed verdict as a failure instead of retyping or assuming delivery. +`fm-send.sh` never retypes or assumes a confirmed submit for an unconfirmed verdict; its header owns the distinct delivered-unconfirmed exit status and operator response. OpenCode 1.18.4 has one busy-queue exception. While OpenCode is mid-turn, Enter queues the message but leaves its text visible until the turn completes. diff --git a/docs/verification/dispatch-auth.md b/docs/verification/dispatch-auth.md index 4ef443b8a87..57772f113f7 100644 --- a/docs/verification/dispatch-auth.md +++ b/docs/verification/dispatch-auth.md @@ -12,9 +12,9 @@ Credential paths below are shown with the home directory replaced by ``. ## Quota granularity the judgment depends on -Verified 2026-07-30 against quota-axi 0.1.16. - -`quota-axi --json` reports availability at whatever granularity the vendor supplies, and states the vendor's own bounding rule in `quotaSemantics.description`. +Verified 2026-07-30 against quota-axi 0.1.16 for the provider and model-scope relationships below. +That release's captured default output included `quotaSemantics.description`; the current default TOON and JSON fallback field placement are verified against 0.1.29 in the next section. +Current dispatch reads the TOON scope and `limitedBy` fields; the JSON fallback's corresponding `scope` and `boundedBy` fields preserve the same provider/model applicability without relying on the `--full`-only description. ```json { @@ -40,18 +40,27 @@ Three properties follow and are load-bearing for dispatch: `quotaSemantics.status` is `unknown` with no `effectiveAvailability` entries at all for providers whose vendor exposes no window (observed for `cursor` and `copilot`). `state.authStatus` is present only for some providers (observed for `grok` alone), so its absence is missing evidence, not a credential fault. -## Completion-runway shape the judgment depends on +## Completion-runway and selection shape the judgment depends on + +Verified 2026-08-18 against quota-axi 0.1.29 schema 5, captured from an isolated `quota-axi@0.1.29` install. +The default TOON exposed these table headers, with row counts normalized to `N`: -Verified 2026-07-31 against quota-axi 0.1.17 schema 3. -The command below records the producer shape without persisting account-specific quota values: +```text +quota[N]{provider,scope,effectivePercentRemaining,spendPriority,runway,confidence,limitedBy,resetsAt}: +exhaustion[N]{provider,scope,usableRunwaySeconds,projectedExhaustedAt,limitingWindowId}: +attention[N]{provider,scope,kind,detail,remedy}: +``` + +`exhaustion[]` and `attention[]` are sparse, so an empty table is rendered with count zero and no row fields. +The command below records the JSON fallback shape without persisting account-specific quota values: ```sh -quota-axi --json | jq '{schemaVersion, effectiveAvailabilityFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]? | keys] | unique), runwayFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]?.runway? | select(type == "object") | keys] | unique)}' +quota-axi --json | jq '{schemaVersion, effectiveAvailabilityFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]? | keys] | unique), runwayFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]?.runway? | select(type == "object") | keys] | unique), selectionFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]?.selection? | select(type == "object") | keys] | unique), paceFields: ([.providers[]?.quotaSemantics.effectiveAvailability[]?.pace? | select(type == "object") | keys] | unique), windowPaceFields: ([.providers[]?.windows[]?.pace? | select(type == "object") | keys] | unique)}' ``` ```json { - "schemaVersion": 3, + "schemaVersion": 5, "effectiveAvailabilityFields": [ [ "boundedBy", @@ -60,31 +69,47 @@ quota-axi --json | jq '{schemaVersion, effectiveAvailabilityFields: ([.providers "pace", "runway", "scope", + "selection", "status" ] ], "runwayFields": [ [ - "limitingWindowId", - "projectedExhaustedAt", - "projectionBasis", "projectionConfidence", - "status", - "usableRunwaySeconds" - ], + "status" + ] + ], + "selectionFields": [ + [ + "spendPriority", + "status" + ] + ], + "paceFields": [ [ - "limitingWindowId", - "projectedExhaustedAt", "status", - "usableRunwaySeconds" + "worstReservePercentPoints", + "worstReserveWindowId" + ] + ], + "windowPaceFields": [ + [ + "burnMultiple", + "reservePercentPoints", + "status" ] ] } ``` -`runway` is nested under each effective-availability scope, so the same provider/model applicability rules govern both effective headroom and runway. -Projection confidence and basis are not present on every known runway, so selection must preserve their absence as uncertainty rather than fabricate them. -The older-schema fallback contract is owned by `quota-array-dispatch`; this evidence does not reinterpret an absent runway or pace field. +This live snapshot was all `through_reset`, so finite-runway fields were omitted. +`usableRunwaySeconds`, `projectedExhaustedAt`, and `limitingWindowId` remain in default `--json` when `runway.status` is `projected_exhaustion` or `exhausted_now`. +`selection.unmeasurableWindowIds`, scope `aheadWindowIds`/`unknownWindowIds`, and window `pace.reason` likewise remain in default `--json` when they apply. +`quotaSemantics.description`, `behindWindowIds`, `onPaceWindowIds`, and per-window cycle-progress internals are `--full` only. +There is no `projectionBasis` field; its absence means `cycle_average`. +`runway` and `selection` are nested under each effective-availability scope, so the same provider/model applicability rules govern headroom, runway, and `spendPriority`. +Projection confidence is not present on every known runway, so selection must preserve that absence as uncertainty rather than fabricate it. +The older-schema fallback contract is owned by `quota-array-dispatch`; this evidence does not reinterpret an absent runway, pace, or selection field. ## Provider-family counterfactual that this producer schema supports @@ -174,5 +199,6 @@ Re-run the two commands above and update this section and the pinned version tog It asserts that the script accepts no harness, model, or provider input, never calls `quota-axi`, exits alike for every probe result because it renders no verdict, invokes only the two fixed non-destructive argv forms with stdin closed, holds a real bound even when the configured bound is zero or malformed, and never echoes raw vendor output. `tests/fm-spawn-dispatch-profile.test.sh` owns spawn's deterministic profile and harness refusals. `tests/fm-bootstrap.test.sh` owns the quota-axi version-floor diagnostic. -`tests/fm-quota-array-dispatch-live-e2e.test.sh` drives the public Pi skill-loading interface against one fake `quota-axi --json` snapshot per case. -It covers the Claude 1 percent versus Codex 55 percent reserve regression, explicit accounting for unmeasurable runway, and the strongest-reasoning constraint. +`tests/fm-quota-array-dispatch-live-e2e.test.sh` drives the public Pi skill-loading interface against one fake schema-5 snapshot per case, served as quota-axi's default TOON. +It covers TOON-first `spendPriority` ranking among candidates that pass eligibility, reasoning-class, and runway-feasibility gates, explicit accounting for unmeasurable runway, the strongest-reasoning constraint, and the runway feasibility floor over a higher `spendPriority`. +The skill's primary path is that default TOON; `--json` is the documented defensive fallback, and this section records the producer `--json` shape that fallback consumes. diff --git a/fork-divergences.json b/fork-divergences.json index d8f483708fb..f0972e16c24 100644 --- a/fork-divergences.json +++ b/fork-divergences.json @@ -68,6 +68,18 @@ "supervision-owner-aware-recovery" ], "validation_pr": null + }, + { + "date": "2026-08-18", + "fork_before": "c88b1ffbd0a01ea7b1c1646524977d2460aaf3ae", + "upstream_before": "d023c451e00fb64f9845b27fa949c02beed8c551", + "upstream_after": "03bb1d8b78a8632ae2d9cea4c10868eb100e885e", + "touched": [ + "firstmate-progressive-disclosure-restructure", + "fork-main-integration", + "supervision-owner-aware-recovery" + ], + "validation_pr": null } ], "divergences": [ diff --git a/tests/fm-bootstrap.test.sh b/tests/fm-bootstrap.test.sh index 5527d14722e..1810e6b5f0b 100755 --- a/tests/fm-bootstrap.test.sh +++ b/tests/fm-bootstrap.test.sh @@ -95,7 +95,7 @@ add_quota_axi() { cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' "${FM_FAKE_QUOTA_AXI_VERSION:-0.1.25}" + printf '%s\n' "${FM_FAKE_QUOTA_AXI_VERSION:-0.1.29}" exit 0 fi exit 0 @@ -473,11 +473,11 @@ test_quota_axi_min_version() { [ "$out" = "$missing" ] || fail "$label: expected '$missing', got: $out" ;; esac done <<'ROWS' -minimum quota-axi version is accepted^0.1.25^empty -newer quota-axi patch is accepted^0.1.26^empty +minimum quota-axi version is accepted^0.1.29^empty +newer quota-axi patch is accepted^0.1.30^empty newer quota-axi minor is accepted^0.2.0^empty newer quota-axi major is accepted^1.0.0^empty -the patch just below the floor reports an upgrade^0.1.24^missing +the patch just below the floor reports an upgrade^0.1.28^missing much older quota-axi minor reports an upgrade^0.0.9^missing unparseable quota-axi version reports an upgrade^quota-axi development build^missing ROWS diff --git a/tests/fm-crew-state.test.sh b/tests/fm-crew-state.test.sh index 8f986b6139e..602b3e5cfc3 100755 --- a/tests/fm-crew-state.test.sh +++ b/tests/fm-crew-state.test.sh @@ -1150,6 +1150,105 @@ test_torn_down_worktree() { pass "torn-down worktree is handled gracefully" } +# --- remote secondmate arm --------------------------------------------------- +# A meta recording remote_host= must never be read through the local worktree +# probe or a local backend adapter: the recorded worktree and pane live on the +# remote host, and the old local reads misreported a healthy remote mate as +# "worktree gone". These cases drive the real helper over the real fm-on.sh +# route with a stubbed ssh transport (FM_SSH_BIN seam): the stub prints +# FM_FAKE_REMOTE_STATE_OUT as the remote endpoint's recovery-grade state and +# exits FM_FAKE_SSH_RC. + +setup_remote_case() { # -> echoes case dir with remote meta + registry + local d + d=$(new_case "$1") + mkdir -p "$d/data" "$d/fakebin" + fm_write_meta "$d/state/rsm.meta" \ + "window=remote:rsm" \ + "endpoint_task_id=rsm" \ + "worktree=/remote/home/never-locally-present" \ + "harness=claude" \ + "kind=secondmate" \ + "mode=secondmate" \ + "remote_host=remote-mac" \ + "remote_root=/remote/root" \ + "remote_backend=herdr" \ + "remote_herdr_session=fm-remote" \ + "remote_target=fm-remote:w1:p1" + cat > "$d/data/secondmates.md" < "$d/fakebin/fake-ssh" <<'SH' +#!/usr/bin/env bash +cat > /dev/null +[ -z "${FM_FAKE_REMOTE_STATE_OUT:-}" ] || printf '%s\n' "$FM_FAKE_REMOTE_STATE_OUT" +exit "${FM_FAKE_SSH_RC:-0}" +SH + chmod +x "$d/fakebin/fake-ssh" + printf '%s\n' "$d" +} + +run_remote_crew_state() { # + PATH="$1/fakebin:$PATH" FM_HOME="$1" FM_STATE_OVERRIDE="$1/state" \ + FM_SSH_BIN="$1/fakebin/fake-ssh" "$CREW_STATE" "$2" +} + +test_remote_alive_with_log_uses_status_log() { + reset_fakes + local d out rc + d=$(setup_remote_case remote-alive-log) + make_fakebin "$d" >/dev/null + printf 'working: refactoring the quota adapter\n' > "$d/state/rsm.status" + out=$(FM_FAKE_REMOTE_STATE_OUT=alive FM_FAKE_SSH_RC=0 run_remote_crew_state "$d" rsm); rc=$? + expect_code 0 "$rc" "remote alive exits 0" + assert_contains "$out" "state: working" "alive remote mate with a working log reads working" + assert_contains "$out" "source: status-log" "alive remote mate reads current activity from the routed log" + assert_contains "$out" "remote endpoint alive on remote-mac" "the remote liveness read should be visible" + assert_not_contains "$out" "worktree gone" "a healthy remote mate must never read as torn down" + pass "fm-crew-state remote: alive endpoint falls through to the routed status log" +} + +test_remote_alive_idle_is_healthy_not_gone() { + reset_fakes + local d out rc + d=$(setup_remote_case remote-alive-idle) + make_fakebin "$d" >/dev/null + out=$(FM_FAKE_REMOTE_STATE_OUT=alive FM_FAKE_SSH_RC=0 run_remote_crew_state "$d" rsm); rc=$? + expect_code 0 "$rc" "remote alive-idle exits 0" + assert_contains "$out" "source: remote-endpoint" "the remote endpoint is the reported source" + assert_contains "$out" "alive on remote-mac" "an idle remote mate reads alive" + assert_not_contains "$out" "worktree gone" "a healthy remote mate must never read as torn down" + assert_not_contains "$out" "backend target gone" "a healthy remote mate must never read as a dead target" + pass "fm-crew-state remote: an idle alive endpoint reads alive, never gone or dead" +} + +test_remote_unreachable_is_unknown_remote_not_dead() { + reset_fakes + local d out rc + d=$(setup_remote_case remote-unreachable) + make_fakebin "$d" >/dev/null + printf 'working: refactoring the quota adapter\n' > "$d/state/rsm.status" + out=$(FM_FAKE_SSH_RC=255 run_remote_crew_state "$d" rsm); rc=$? + expect_code 0 "$rc" "unreachable remote exits 0" + assert_contains "$out" "unknown-remote" "an unreachable remote must be labeled unknown-remote" + assert_contains "$out" "not proof of death" "an unreachable remote must not read as dead" + assert_not_contains "$out" "worktree gone" "an unreachable remote must never read as torn down" + assert_not_contains "$out" "backend target gone" "an unreachable remote must never read as a dead target" + pass "fm-crew-state remote: an unreachable host reads unknown-remote, never gone or dead" +} + +test_remote_dead_reports_remote_verdict() { + reset_fakes + local d out rc + d=$(setup_remote_case remote-dead) + make_fakebin "$d" >/dev/null + out=$(FM_FAKE_REMOTE_STATE_OUT=dead FM_FAKE_SSH_RC=0 run_remote_crew_state "$d" rsm); rc=$? + expect_code 0 "$rc" "remote dead exits 0" + assert_contains "$out" "remote endpoint dead on remote-mac" \ + "a genuinely dead remote endpoint reports the remote host's own verdict" + pass "fm-crew-state remote: the remote host's own dead verdict is reported truthfully" +} + test_missing_meta() { reset_fakes local d; d=$(new_case nometa) @@ -1350,6 +1449,10 @@ test_dead_window_still_reports_active_run_step test_no_timeout_uses_perl_bound test_scout_skips_run_lookup test_torn_down_worktree +test_remote_alive_with_log_uses_status_log +test_remote_alive_idle_is_healthy_not_gone +test_remote_unreachable_is_unknown_remote_not_dead +test_remote_dead_reports_remote_verdict test_missing_meta test_provably_working_via_runs_list_fallback test_not_provably_working_when_stopped diff --git a/tests/fm-daemon.test.sh b/tests/fm-daemon.test.sh index 2fe02fb4318..ad0925be4d4 100755 --- a/tests/fm-daemon.test.sh +++ b/tests/fm-daemon.test.sh @@ -1569,27 +1569,37 @@ test_inject_wedge_alarm_throttles_when_marker_cannot_be_written() { pass "in-process wedge throttle prevents alert spam when the marker cannot persist" } -test_fm_send_exits_nonzero_on_confirmed_swallow() { - # fm-send.sh must exit NON-ZERO when a steer's Enter is positively swallowed - # (text left in the composer), so firstmate learns the instruction did not land - # — and exit ZERO on a clean submit. - local dir fakebin err +test_fm_send_reports_delivered_unconfirmed_submit() { + # When text was typed and Enter sent but the submit read-back remains pending, + # fm-send must return its documented delivered-unconfirmed status and prevent + # a duplicate resend reflex. A synchronously confirmed submit remains zero. + local dir fakebin err rc dir=$(make_bordered_case send-swallow) fakebin="$dir/fakebin"; err="$dir/send.err" # Clean submit -> exit 0. PATH="$fakebin:$PATH" FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" FM_FAKE_COMPOSER="$dir/composer" \ FM_SEND_SLEEP=0.05 "$ROOT/bin/fm-send.sh" sess:win 'route this work' >/dev/null 2>"$err" \ || fail "fm-send exited non-zero on a clean submit: $(cat "$err")" - # Persistent swallow -> exit non-zero with a clear message. + # Persistent composer text after Enter -> delivered-unconfirmed exit 3 with + # a non-error warning that explicitly tells the operator not to resend. printf '╭─────╮\n│ > │\n╰─────╯\n' > "$dir/composer" touch "$dir/.swallow" if PATH="$fakebin:$PATH" FM_HOME="$dir" FM_STATE_OVERRIDE="$dir/state" FM_FAKE_COMPOSER="$dir/composer" \ FM_FAKE_SWALLOW="$dir/.swallow" FM_FAKE_PERSIST_SWALLOW=1 FM_SEND_SLEEP=0.05 \ "$ROOT/bin/fm-send.sh" sess:win 'fix findings 1 and 3, skip 2' >/dev/null 2>"$err"; then - fail "fm-send exited zero despite a swallowed Enter (silent unsubmitted instruction)" + rc=0 + else + rc=$? + fi + [ "$rc" -eq 3 ] || fail "fm-send returned $rc instead of delivered-unconfirmed exit 3: $(cat "$err")" + grep -F 'submission is unconfirmed' "$err" >/dev/null \ + || fail "fm-send did not explain the pending confirmation: $(cat "$err")" + grep -F 'do not retype or blindly resend' "$err" >/dev/null \ + || fail "fm-send did not prevent a duplicate resend: $(cat "$err")" + if grep -F 'error:' "$err" >/dev/null; then + fail "fm-send mislabeled delivered-unconfirmed as an error: $(cat "$err")" fi - grep -F 'not submitted' "$err" >/dev/null || fail "fm-send did not explain the swallowed submit: $(cat "$err")" - pass "fm-send exits non-zero on a confirmed swallow, zero on a clean submit" + pass "fm-send returns 3 with a non-error no-resend warning when confirmation stays pending" } test_fm_send_exits_nonzero_on_initial_send_failure() { @@ -1916,7 +1926,7 @@ test_wedge_alarm_hung_override_times_out_and_falls_through test_wedge_alarm_shutdown_stops_active_notifier_group test_inject_wedge_alarm_fires_active_alert_on_non_tmux_backend test_inject_wedge_alarm_throttles_when_marker_cannot_be_written -test_fm_send_exits_nonzero_on_confirmed_swallow +test_fm_send_reports_delivered_unconfirmed_submit test_fm_send_exits_nonzero_on_initial_send_failure test_fm_send_exits_nonzero_on_unproven_submit test_discover_supervisor_backend_precedence diff --git a/tests/fm-inactive-reconcile.test.sh b/tests/fm-inactive-reconcile.test.sh index dc8e06c2edf..c4621194206 100755 --- a/tests/fm-inactive-reconcile.test.sh +++ b/tests/fm-inactive-reconcile.test.sh @@ -404,7 +404,10 @@ test_full_scan_budget_includes_wake_lock_wait() { FM_INACTIVE_RECONCILE_BUDGET_SECS=1 FM_FAKE_CREW_STATE='done' run_reconcile "$MAIN" --startup elapsed=$(( $(date +%s) - started )) reap "$holder" - [ "$elapsed" -le 3 ] || fail "wake lock wait exceeded aggregate scan budget (${elapsed}s)" + # The unbounded wake-lock wait is ended by the process-group backstop, which + # fires one second after the budget; the bound proves the scan cannot ride + # the 30-second lock hold. + [ "$elapsed" -le 4 ] || fail "wake lock wait exceeded aggregate scan budget (${elapsed}s)" pass "aggregate scan budget includes durable wake operations" } diff --git a/tests/fm-peek-remote.test.sh b/tests/fm-peek-remote.test.sh new file mode 100755 index 00000000000..7ef7286fb23 --- /dev/null +++ b/tests/fm-peek-remote.test.sh @@ -0,0 +1,110 @@ +#!/usr/bin/env bash +# fm-peek remote-secondmate capture routing. +# +# A remote secondmate's pane lives on its own host. The old path resolved the +# meta's "remote:" window through the local backend adapters and handed it +# to tmux, which failed with "can't find session: remote" - a healthy remote +# mate misreported as an unreadable endpoint. These tests drive the real +# fm-peek + fm-on executables with a stubbed ssh transport (FM_SSH_BIN seam) +# and a poisoned local tmux, pinning: +# 1. A remote selector routes the capture over the remote transport and +# prints the remote pane tail; the local adapters are never consulted. +# 2. An unreachable host fails loudly naming the host, without claiming the +# mate is dead. +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +PEEK="$ROOT/bin/fm-peek.sh" + +TMP_ROOT=$(fm_test_tmproot fm-peek-remote) + +# fake-ssh prints the canned remote capture; the poisoned tmux records any +# local read attempt so the "never consulted" property is a real assertion. +make_stubs() { # -> echoes fakebin dir + local dir=$1 fb="$1/fakebin" + mkdir -p "$fb" + cat > "$fb/fake-ssh" <<'SH' +#!/usr/bin/env bash +cat > /dev/null +[ -z "${FM_FAKE_REMOTE_CAPTURE:-}" ] || printf '%s\n' "$FM_FAKE_REMOTE_CAPTURE" +exit "${FM_FAKE_SSH_RC:-0}" +SH + chmod +x "$fb/fake-ssh" + cat > "$fb/tmux" <<'SH' +#!/usr/bin/env bash +printf 'tmux\n' >> "${FM_FAKE_TMUX_TOUCHED:?}" +exit 1 +SH + chmod +x "$fb/tmux" + printf '%s\n' "$fb" +} + +setup_remote_home() { # -> echoes home dir with remote meta + registry + local home="$TMP_ROOT/$1-$RANDOM" + mkdir -p "$home/state" "$home/data" + fm_write_meta "$home/state/rsm.meta" \ + "window=remote:rsm" \ + "endpoint_task_id=rsm" \ + "harness=claude" \ + "kind=secondmate" \ + "mode=secondmate" \ + "remote_host=remote-mac" \ + "remote_root=/remote/root" \ + "remote_backend=herdr" \ + "remote_herdr_session=fm-remote" \ + "remote_target=fm-remote:w1:p1" + cat > "$home/data/secondmates.md" < "$touched" + + out=$(env PATH="$fb:$PATH" \ + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + FM_SSH_BIN="$fb/fake-ssh" FM_FAKE_SSH_RC=0 \ + FM_FAKE_REMOTE_CAPTURE='● the remote mate is mid-refactor' \ + FM_FAKE_TMUX_TOUCHED="$touched" \ + "$PEEK" rsm 20 2>"$dir/err"); rc=$? + expect_code 0 "$rc" "a healthy remote peek should succeed" + assert_contains "$out" "the remote mate is mid-refactor" \ + "the remote pane tail should be printed" + assert_not_contains "$out" "can't find session" \ + "a remote peek must not fall into a local session lookup" + [ ! -s "$touched" ] || fail "the local tmux adapter was consulted for a remote target" + pass "fm-peek remote: the capture routes over the remote transport, local adapters untouched" +} + +test_remote_peek_unreachable_fails_loudly_without_death_claim() { + local dir fb home touched rc err + dir="$TMP_ROOT/peek-down"; mkdir -p "$dir" + fb=$(make_stubs "$dir") + home=$(setup_remote_home peek-down) + touched="$dir/tmux-touched"; : > "$touched" + + env PATH="$fb:$PATH" \ + FM_HOME="$home" FM_STATE_OVERRIDE="$home/state" \ + FM_SSH_BIN="$fb/fake-ssh" FM_FAKE_SSH_RC=255 \ + FM_FAKE_TMUX_TOUCHED="$touched" \ + "$PEEK" rsm >"$dir/out" 2>"$dir/err"; rc=$? + err=$(cat "$dir/err") + [ "$rc" -ne 0 ] || fail "an unreachable remote peek must exit nonzero" + assert_contains "$err" "remote pane of rsm on remote-mac" \ + "the failure must name the remote mate and host" + assert_contains "$err" "not thereby dead" \ + "an unreadable remote pane must not be presented as a dead mate" + pass "fm-peek remote: an unreachable host fails loudly without a false death claim" +} + +test_remote_peek_reads_remote_pane +test_remote_peek_unreachable_fails_loudly_without_death_claim + +echo "all fm-peek-remote tests passed" diff --git a/tests/fm-quota-array-dispatch-live-e2e.test.sh b/tests/fm-quota-array-dispatch-live-e2e.test.sh index 0b7f1102aba..417aeef86ca 100755 --- a/tests/fm-quota-array-dispatch-live-e2e.test.sh +++ b/tests/fm-quota-array-dispatch-live-e2e.test.sh @@ -3,7 +3,9 @@ # # This drives the public Pi skill-loading interface against a fake quota-axi # executable rather than parsing instruction source bytes or recreating the -# selector in test code. +# selector in test code. The fake serves default TOON from the schema-5 JSON +# fixture; --json remains available so a TOON-first skill cannot silently +# fall back without the call log catching it. set -u if [ "${FM_QUOTA_ARRAY_DISPATCH_LIVE_E2E:-0}" != 1 ]; then @@ -20,6 +22,7 @@ fail() { } command -v pi >/dev/null 2>&1 || fail "pi not found" +command -v python3 >/dev/null 2>&1 || fail "python3 not found" [ -f "$OWNER" ] || fail "quota-array-dispatch skill not found" LAB=$(mktemp -d "${TMPDIR:-/tmp}/fm-quota-array-dispatch-live.XXXXXX") @@ -38,13 +41,118 @@ cp "$OWNER" "$PROJECT/.agents/skills/quota-array-dispatch/SKILL.md" cat > "$FAKEBIN/quota-axi" <<'SH' #!/usr/bin/env bash +# Fake quota-axi: default TOON from the schema-5 JSON fixture; --json dumps it. set -u -if [ "${1:-}" != --json ] || [ "$#" -ne 1 ]; then - printf 'unexpected quota-axi invocation: %s\n' "$*" >&2 - exit 64 -fi -printf '%s\n' "$*" >> "${QUOTA_AXI_CALLS:?}" -cat "${QUOTA_AXI_FIXTURE:?}" +record() { + printf '%s\n' "$1" >> "${QUOTA_AXI_CALLS:?}" +} +emit_toon() { + python3 - "${QUOTA_AXI_FIXTURE:?}" <<'PY' +import json +import sys + +data = json.load(open(sys.argv[1], encoding="utf-8")) +generated = data.get("generatedAt", "unknown") +quota = [] +exhaustion = [] +attention = [] + + +def join_ids(ids): + if not ids: + return "unknown" + return " + ".join(str(item) for item in ids) + + +for provider in data.get("providers") or []: + name = provider.get("provider", "unknown") + windows = {window.get("id"): window for window in (provider.get("windows") or [])} + semantics = provider.get("quotaSemantics") or {} + for scope in semantics.get("effectiveAvailability") or []: + remaining = scope.get("effectivePercentRemaining") + selection = scope.get("selection") or {} + runway = scope.get("runway") or {} + scope_name = scope.get("scope", "unknown") + if remaining is None: + attention.append( + f" {name},{scope_name},headroom_unknown,{join_ids(runway.get('unmeasurableWindowIds') or scope.get('boundedBy'))},none" + ) + continue + if selection.get("status") == "known" and "spendPriority" in selection: + spend = selection["spendPriority"] + else: + spend = "unknown" + runway_status = runway.get("status") or "unknown" + confidence = runway.get("projectionConfidence") or "unknown" + limited = join_ids(scope.get("limitingWindowIds")) + binding = None + for window_id in scope.get("limitingWindowIds") or []: + binding = (windows.get(window_id) or {}).get("resetsAt") + if binding: + break + resets_at = binding or "unknown" + quota.append( + f" {name},{scope_name},{remaining},{spend},{runway_status},{confidence},{limited},{resets_at}" + ) + if runway_status in ("projected_exhaustion", "exhausted_now"): + seconds = runway.get("usableRunwaySeconds", "unknown") + exhausted_at = runway.get("projectedExhaustedAt", "unknown") + limiting = runway.get("limitingWindowId", "unknown") + exhaustion.append( + f" {name},{scope_name},{seconds},{exhausted_at},{limiting}" + ) + blocked = [] + if runway.get("unmeasurableWindowIds"): + blocked.append(f"{join_ids(runway['unmeasurableWindowIds'])} blocks runway") + if selection.get("unmeasurableWindowIds"): + blocked.append( + f"{join_ids(selection['unmeasurableWindowIds'])} blocks spendPriority" + ) + if blocked: + attention.append( + f" {name},{scope_name},unmeasurable,{' · '.join(blocked)},none" + ) + +print('bin: fake-quota-axi') +print('description: Report local agent-provider quota windows for routing-aware agents') +print(f'generatedAt: "{generated}"') +print( + f"quota[{len(quota)}]{{provider,scope,effectivePercentRemaining,spendPriority,runway,confidence,limitedBy,resetsAt}}:" +) +print("\n".join(quota) if quota else "") +print( + f"exhaustion[{len(exhaustion)}]{{provider,scope,usableRunwaySeconds,projectedExhaustedAt,limitingWindowId}}:" + if exhaustion + else "exhaustion[0]:" +) +if exhaustion: + print("\n".join(exhaustion)) +print( + f"attention[{len(attention)}]{{provider,scope,kind,detail,remedy}}:" + if attention + else "attention[0]:" +) +if attention: + print("\n".join(attention)) +print("help[1]:") +print(" Run `quota-axi --full` for windows, pace, reserve, and account evidence") +PY +} + +case "$*" in + ""|quota) + record TOON + emit_toon + ;; + --json) + record JSON + cat "${QUOTA_AXI_FIXTURE:?}" + ;; + *) + printf 'unexpected quota-axi invocation: %s\n' "$*" >&2 + exit 64 + ;; +esac SH chmod +x "$FAKEBIN/quota-axi" @@ -53,8 +161,8 @@ write_fixture() { } run_case() { - local label=$1 expected=$2 prompt=$3 out calls required - shift 3 + local label=$1 expected=$2 expected_calls=$3 prompt=$4 out calls required + shift 4 : > "$CALLS" out=$( cd "$PROJECT" && @@ -65,7 +173,7 @@ run_case() { "$prompt" ) || fail "$label: Pi skill run failed: $out" calls=$(cat "$CALLS") - [ "$calls" = "--json" ] || fail "$label: skill did not use one quota-axi --json snapshot: $calls" + [ "$calls" = "$expected_calls" ] || fail "$label: unexpected quota-axi call sequence: $calls" printf '%s\n' "$out" | grep -Fxq "$expected" \ || fail "$label: expected final line $expected, got: $out" for required in "$@"; do @@ -77,33 +185,342 @@ run_case() { } write_fixture <<'JSON' -{"schemaVersion":3,"providers":[{"provider":"claude","quotaSemantics":{"description":"The all_models scope bounds every Claude model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":1,"boundedBy":["weekly"],"runway":{"status":"projected_exhaustion","usableRunwaySeconds":600,"projectedExhaustedAt":"2030-01-01T00:10:00Z","limitingWindowId":"weekly","projectionConfidence":"established","projectionBasis":"cycle_average"}}]},"effectivePace":[{"scope":"all_models","pace":"ahead","worstReservePercentPoints":-1}]},{"provider":"codex","quotaSemantics":{"description":"The all_models scope bounds every Codex model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":55,"boundedBy":["weekly"],"runway":{"status":"projected_exhaustion","usableRunwaySeconds":14400,"projectedExhaustedAt":"2030-01-01T04:00:00Z","limitingWindowId":"weekly","projectionConfidence":"established","projectionBasis":"cycle_average"}}]},"effectivePace":[{"scope":"all_models","pace":"ahead","worstReservePercentPoints":-40}]}]} +{ + "generatedAt": "2030-01-01T00:00:00Z", + "schemaVersion": 5, + "providers": [ + { + "provider": "claude", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 80, + "resetsAt": "2030-01-07T07:12:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -10, "burnMultiple": 2 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 80, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -1.1111 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 241920, + "projectedExhaustedAt": "2030-01-03T19:12:00Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -10, "worstReserveWindowId": "weekly" } + } + ] + } + }, + { + "provider": "codex", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 20, + "resetsAt": "2030-01-03T19:12:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -20, "burnMultiple": 1.3333 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 20, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -0.8333 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 90720, + "projectedExhaustedAt": "2030-01-02T01:12:00Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -20, "worstReserveWindowId": "weekly" } + } + ] + } + } + ] +} JSON run_case \ - "higher headroom and viable runway beat a less-negative reserve" \ + "higher spendPriority beats more headroom after the three gates" \ "SELECTED=codex" \ - "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi --json exactly once. Both profiles have comparable required task fit and the same strongest reasoning class. The authoritative catalogs already prove Claude/Sonnet and Codex/GPT models supported in their stated provider families, and their selected authentication surfaces are usable. The likely task-completion horizon is two hours with established confidence. Return exact lines FACT=claude|headroom=1|runway_seconds=600|reserve=-1 and FACT=codex|headroom=55|runway_seconds=14400|reserve=-40 to preserve candidate accounting, then an exact final line SELECTED=. Do not use other vendor or model commands and do not modify files." \ - "FACT=claude|headroom=1|runway_seconds=600|reserve=-1" \ - "FACT=codex|headroom=55|runway_seconds=14400|reserve=-40" + "TOON" \ + "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi with no flags (default TOON) exactly once. Do not pass --json. Both profiles have comparable required task fit and the same strongest reasoning class. The authoritative catalogs already prove Claude/Sonnet and Codex/GPT models supported in their stated provider families, and their selected authentication surfaces are usable. The likely task-completion horizon is two hours with established confidence. Both candidates have known runway that supports that horizon. Return exact lines FACT=claude|headroom=80|spendPriority=-1.1111|runway_seconds=241920 and FACT=codex|headroom=20|spendPriority=-0.8333|runway_seconds=90720 to preserve candidate accounting, then an exact final line SELECTED=. Do not use other vendor or model commands and do not modify files." \ + "FACT=claude|headroom=80|spendPriority=-1.1111|runway_seconds=241920" \ + "FACT=codex|headroom=20|spendPriority=-0.8333|runway_seconds=90720" write_fixture <<'JSON' -{"schemaVersion":3,"providers":[{"provider":"claude","quotaSemantics":{"description":"The all_models scope bounds every Claude model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":55,"boundedBy":["weekly"],"runway":{"status":"unknown","unmeasurableWindowIds":["weekly"]}}]}},{"provider":"codex","quotaSemantics":{"description":"The all_models scope bounds every Codex model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":45,"boundedBy":["weekly"],"runway":{"status":"projected_exhaustion","usableRunwaySeconds":14400,"projectedExhaustedAt":"2030-01-01T04:00:00Z","limitingWindowId":"weekly","projectionConfidence":"established","projectionBasis":"cycle_average"}}]}}]} +{ + "generatedAt": "2030-01-01T00:00:00Z", + "schemaVersion": 5, + "providers": [ + { + "provider": "claude", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 55, + "resetsAt": "2030-01-08T00:00:00Z", + "pace": { "status": "unknown", "reason": "missing_cycle" } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 55, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "unknown", "unmeasurableWindowIds": ["weekly"] }, + "runway": { "status": "unknown", "unmeasurableWindowIds": ["weekly"] }, + "pace": { "status": "unknown", "unknownWindowIds": ["weekly"] } + } + ] + } + }, + { + "provider": "codex", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 45, + "resetsAt": "2030-01-04T20:24:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -10, "burnMultiple": 1.2222 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 45, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -0.404 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 222676, + "projectedExhaustedAt": "2030-01-03T13:51:16Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -10, "worstReserveWindowId": "weekly" } + } + ] + } + } + ] +} JSON run_case \ "unmeasurable runway stays eligible and is accounted for explicitly" \ "DECISION=CODEX" \ - "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi --json exactly once. Both profiles have comparable required task fit and the same strongest reasoning class. The authoritative catalogs already prove both models supported in their stated provider families, and their selected authentication surfaces are usable. The likely task-completion horizon is two hours with established confidence. Claude has higher known headroom but explicitly unmeasurable runway, while Codex has lower known headroom and established runway that supports completion. The snapshot cannot prove Pareto dominance in either direction, but the known completion-supporting runway justifies Codex while Claude remains eligible and its uncertainty must be disclosed. Return exact lines FACT=claude|eligible=yes|headroom=55|runway=unknown|unmeasurable=weekly and FACT=codex|eligible=yes|headroom=45|runway_seconds=14400|supports_horizon=yes, then an exact final line DECISION=CODEX. Do not use other vendor or model commands and do not modify files." \ - "FACT=claude|eligible=yes|headroom=55|runway=unknown|unmeasurable=weekly" \ - "FACT=codex|eligible=yes|headroom=45|runway_seconds=14400|supports_horizon=yes" + "TOON +JSON" \ + "Resolve this matched dispatch profile array now. Load quota-array-dispatch and consult quota-axi's default TOON first. Because Claude spendPriority is the literal unknown, use the permitted quota-axi --json fallback once before deciding. Both profiles have comparable required task fit and the same strongest reasoning class. The authoritative catalogs already prove both models supported in their stated provider families, and their selected authentication surfaces are usable. The likely task-completion horizon is two hours with established confidence. Claude has higher known headroom but explicitly unmeasurable runway and unknown spendPriority, while Codex has lower known headroom, known spendPriority, and established runway that supports completion. Claude remains eligible and its uncertainty must be disclosed. Never read unknown spendPriority as 0. Return exact lines FACT=claude|eligible=yes|headroom=55|runway=unknown|spendPriority=unknown|unmeasurable=weekly and FACT=codex|eligible=yes|headroom=45|spendPriority=-0.404|runway_seconds=222676|supports_horizon=yes, then an exact final line DECISION=CODEX. Do not use other vendor or model commands and do not modify files." \ + "FACT=claude|eligible=yes|headroom=55|runway=unknown|spendPriority=unknown|unmeasurable=weekly" \ + "FACT=codex|eligible=yes|headroom=45|spendPriority=-0.404|runway_seconds=222676|supports_horizon=yes" write_fixture <<'JSON' -{"schemaVersion":3,"providers":[{"provider":"claude","quotaSemantics":{"description":"The all_models scope bounds every Claude model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":1,"boundedBy":["weekly"],"runway":{"status":"projected_exhaustion","usableRunwaySeconds":10800,"projectedExhaustedAt":"2030-01-01T03:00:00Z","limitingWindowId":"weekly","projectionConfidence":"established","projectionBasis":"cycle_average"}}]}},{"provider":"codex","quotaSemantics":{"description":"The all_models scope bounds every Codex model.","effectiveAvailability":[{"scope":"all_models","status":"known","effectivePercentRemaining":80,"boundedBy":["weekly"],"runway":{"status":"projected_exhaustion","usableRunwaySeconds":28800,"projectedExhaustedAt":"2030-01-01T08:00:00Z","limitingWindowId":"weekly","projectionConfidence":"established","projectionBasis":"cycle_average"}}]}}]} +{ + "generatedAt": "2030-01-01T00:00:00Z", + "schemaVersion": 5, + "providers": [ + { + "provider": "claude", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 5, + "resetsAt": "2030-01-04T12:00:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -45, "burnMultiple": 1.9 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 5, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -1.8 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 15916, + "projectedExhaustedAt": "2030-01-01T04:25:16Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -45, "worstReserveWindowId": "weekly" } + } + ] + } + }, + { + "provider": "codex", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 80, + "resetsAt": "2030-01-06T22:48:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -5, "burnMultiple": 1.3333 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 80, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -0.3921 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 362880, + "projectedExhaustedAt": "2030-01-05T04:48:00Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -5, "worstReserveWindowId": "weekly" } + } + ] + } + } + ] +} JSON run_case \ "required strongest reasoning class is not downgraded for quota" \ "SELECTED=claude" \ - "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi --json exactly once. The likely task-completion horizon is two hours with established confidence. Claude/Sonnet is catalog-supported with usable authentication and is the only profile that meets the task's required strongest reasoning class. Codex/GPT is catalog-supported with usable authentication but is a weaker reasoning class and cannot meet the requirement. Return exact lines FACT=claude|reasoning=required|headroom=1|runway_seconds=10800 and FACT=codex|reasoning=weaker|headroom=80|runway_seconds=28800, then an exact final line SELECTED=. Do not use other vendor or model commands and do not modify files." \ - "FACT=claude|reasoning=required|headroom=1|runway_seconds=10800" \ - "FACT=codex|reasoning=weaker|headroom=80|runway_seconds=28800" + "TOON" \ + "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi with no flags (default TOON) exactly once. Do not pass --json. The likely task-completion horizon is two hours with established confidence. Claude/Sonnet is catalog-supported with usable authentication and is the only profile that meets the task's required strongest reasoning class. Codex/GPT is catalog-supported with usable authentication but is a weaker reasoning class and cannot meet the requirement. Return exact lines FACT=claude|reasoning=required|headroom=5|spendPriority=-1.8|runway_seconds=15916 and FACT=codex|reasoning=weaker|headroom=80|spendPriority=-0.3921|runway_seconds=362880, then an exact final line SELECTED=. Do not use other vendor or model commands and do not modify files." \ + "FACT=claude|reasoning=required|headroom=5|spendPriority=-1.8|runway_seconds=15916" \ + "FACT=codex|reasoning=weaker|headroom=80|spendPriority=-0.3921|runway_seconds=362880" + +write_fixture <<'JSON' +{ + "generatedAt": "2030-01-01T00:00:00Z", + "schemaVersion": 5, + "providers": [ + { + "provider": "claude", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "five_hour", + "label": "5-hour", + "kind": "five_hour", + "percentRemaining": 20, + "resetsAt": "2030-01-01T02:00:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -20, "burnMultiple": 1.3333 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 20, + "boundedBy": ["five_hour"], + "limitingWindowIds": ["five_hour"], + "selection": { "status": "known", "spendPriority": -0.8333 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 2700, + "projectedExhaustedAt": "2030-01-01T00:45:00Z", + "limitingWindowId": "five_hour", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["five_hour"], "worstReservePercentPoints": -20, "worstReserveWindowId": "five_hour" } + } + ] + } + }, + { + "provider": "codex", + "state": { "status": "fresh", "stale": false }, + "windows": [ + { + "id": "weekly", + "label": "week", + "kind": "weekly", + "percentRemaining": 5, + "resetsAt": "2030-01-04T12:00:00Z", + "pace": { "status": "ahead", "reservePercentPoints": -45, "burnMultiple": 1.9 } + } + ], + "quotaSemantics": { + "status": "known", + "effectiveAvailability": [ + { + "scope": "all_models", + "status": "known", + "effectivePercentRemaining": 5, + "boundedBy": ["weekly"], + "limitingWindowIds": ["weekly"], + "selection": { "status": "known", "spendPriority": -1.8 }, + "runway": { + "status": "projected_exhaustion", + "usableRunwaySeconds": 15916, + "projectedExhaustedAt": "2030-01-01T04:25:16Z", + "limitingWindowId": "weekly", + "projectionConfidence": "established" + }, + "pace": { "status": "ahead", "aheadWindowIds": ["weekly"], "worstReservePercentPoints": -45, "worstReserveWindowId": "weekly" } + } + ] + } + } + ] +} +JSON +run_case \ + "runway versus completion horizon remains a hard gate over spendPriority" \ + "SELECTED=codex" \ + "TOON" \ + "Resolve this matched dispatch profile array now. Load quota-array-dispatch and run quota-axi with no flags (default TOON) exactly once. Do not pass --json. Both profiles have comparable required task fit and the same strongest reasoning class. The authoritative catalogs already prove Claude/Sonnet and Codex/GPT models supported in their stated provider families, and their selected authentication surfaces are usable. The likely task-completion horizon is two hours with established confidence. Claude has known spendPriority of -0.8333 and runway of 2700 seconds. Codex has known spendPriority of -1.8 and runway of 15916 seconds. Return exact lines FACT=claude|spendPriority=-0.8333|runway_seconds=2700|supports_horizon=no and FACT=codex|spendPriority=-1.8|runway_seconds=15916|supports_horizon=yes to preserve candidate accounting, then an exact final line SELECTED=. Do not use other vendor or model commands and do not modify files." \ + "FACT=claude|spendPriority=-0.8333|runway_seconds=2700|supports_horizon=no" \ + "FACT=codex|spendPriority=-1.8|runway_seconds=15916|supports_horizon=yes" echo "# all quota-array-dispatch live behavior tests passed" diff --git a/tests/fm-secondmate-harness.test.sh b/tests/fm-secondmate-harness.test.sh index 6920cf7d12a..a3fefd8bea4 100755 --- a/tests/fm-secondmate-harness.test.sh +++ b/tests/fm-secondmate-harness.test.sh @@ -1107,7 +1107,7 @@ SH cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' '0.1.25' + printf '%s\n' '0.1.29' exit 0 fi exit 0 diff --git a/tests/fm-secondmate-liveness.test.sh b/tests/fm-secondmate-liveness.test.sh index a412cce0f82..84795f6c728 100755 --- a/tests/fm-secondmate-liveness.test.sh +++ b/tests/fm-secondmate-liveness.test.sh @@ -252,7 +252,7 @@ SH cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' '0.1.25' + printf '%s\n' '0.1.29' exit 0 fi exit 0 diff --git a/tests/fm-secondmate-sync.test.sh b/tests/fm-secondmate-sync.test.sh index 8b30696a742..41af97d1bf2 100755 --- a/tests/fm-secondmate-sync.test.sh +++ b/tests/fm-secondmate-sync.test.sh @@ -359,7 +359,7 @@ SH cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' 'quota-axi 0.1.25 (fake)' + printf '%s\n' 'quota-axi 0.1.29 (fake)' fi exit 0 SH diff --git a/tests/fm-send-remote-delivery.test.sh b/tests/fm-send-remote-delivery.test.sh new file mode 100755 index 00000000000..af546fbb4a4 --- /dev/null +++ b/tests/fm-send-remote-delivery.test.sh @@ -0,0 +1,285 @@ +#!/usr/bin/env bash +# fm-send remote-secondmate delivery reporting. +# +# The remote send leg (fm-on.sh -> fm-remote-secondmate-control.sh cmd_send) +# runs fm-send's own verified submit host-locally on the remote machine and +# relays its exit status unchanged. A leg that delivered the text into the +# live verified pane but could not synchronously confirm the submit exits 3 +# (the delivered-unconfirmed contract in bin/fm-send.sh's header); flattening +# that into a generic failure produced the false "error: text not sent" +# report that tempted duplicate resends of steers that had actually landed. +# These tests pin the delivery-reporting contract over the real fm-send + +# fm-on executables with a stubbed ssh transport (FM_SSH_BIN seam - the same +# process boundary tests/fm-on.test.sh proves preserves exit status): +# 1. Remote delivered-unconfirmed (ssh exit 3) is NOT a failure: exit 0, a +# non-error delivered notice, the inner leg's stderr held back, and the +# pending-reply expectation marked delivered (awaiting_report). +# 2. A real remote failure (nonzero, not 3/255) still fails loudly with the +# remote stderr replayed and the undelivered expectation discarded. +# 3. Transport-unknown (ssh exit 255) still refuses loudly and preserves the +# expectation as delivery_unknown. +# 4. A delivered-unconfirmed remote answer still closes its --resolve-key +# decision (delivered-with-pending-confirmation counts as delivered). +# 5. A LOCAL send whose submit read-back stays pending exits 3 with an +# honest non-error message (text delivered, submission unconfirmed). +# 6. That local unconfirmed send still never closes a --resolve-key +# decision (the local ledger boundary is unchanged). +set -u + +# shellcheck source=tests/lib.sh +. "$(dirname "${BASH_SOURCE[0]}")/lib.sh" + +SEND="$ROOT/bin/fm-send.sh" +DRAIN="$ROOT/bin/fm-wake-drain.sh" + +TMP_ROOT=$(fm_test_tmproot fm-send-remote-delivery) + +# Stub tmux for the local legs: logs literal typed text to FM_SEND_LOG. The +# default composer reads empty (clean submit); FM_FAKE_TMUX_PENDING=1 keeps a +# proven pending composer with no busy footer, so the real submit core +# exhausts its Enter budget and reports the pending verdict. The ssh stub +# records the invocation, emits FM_FAKE_SSH_STDERR as the remote leg's stderr, +# and exits FM_FAKE_SSH_RC - the exact relay contract the real transport +# preserves. +make_stubs() { # -> echoes fakebin dir + local dir=$1 fb="$1/fakebin" + mkdir -p "$fb" + cat > "$fb/tmux" <<'SH' +#!/usr/bin/env bash +set -u +case "${1:-}" in + send-keys) + shift + literal=0 + while [ $# -gt 0 ]; do + case "$1" in + -t) shift 2 ;; + -l) literal=1; shift ;; + *) break ;; + esac + done + if [ "$literal" = 1 ]; then + printf '%s' "${1:-}" >> "$FM_SEND_LOG" + fi + exit 0 ;; + display-message) + for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done + printf 'fakepane\n'; exit 0 ;; + capture-pane) + if [ "${FM_FAKE_TMUX_PENDING:-0}" = 1 ]; then + printf '╭────────────╮\n│ > steer │\n╰────────────╯\n' + else + printf '╭────╮\n│ │\n╰────╯\n' + fi + exit 0 ;; + list-windows) exit 0 ;; +esac +exit 0 +SH + chmod +x "$fb/tmux" + cat > "$fb/sleep" <<'SH' +#!/usr/bin/env bash +exit 0 +SH + chmod +x "$fb/sleep" + cat > "$fb/fake-ssh" <<'SH' +#!/usr/bin/env bash +cat > /dev/null +printf '%s\n' "$*" >> "$FM_SSH_LOG" +[ -z "${FM_FAKE_SSH_STDERR:-}" ] || printf '%s\n' "$FM_FAKE_SSH_STDERR" >&2 +exit "${FM_FAKE_SSH_RC:-0}" +SH + chmod +x "$fb/fake-ssh" + printf '%s\n' "$fb" +} + +setup_home() { # -> echoes a fresh home dir with an empty state/ + local home="$TMP_ROOT/$1-$RANDOM" + mkdir -p "$home/state" + printf '%s\n' "$home" +} + +# A home with a remote-secondmate task meta plus the registry row fm-on.sh +# resolves the ssh route from - the same shape a live remote mate records. +setup_remote_home() { # -> echoes home dir + local home + home=$(setup_home "$1") + mkdir -p "$home/data" + fm_write_meta "$home/state/rsm.meta" \ + "window=fm-remote:w1:p1" \ + "endpoint_task_id=rsm" \ + "harness=claude" \ + "kind=secondmate" \ + "mode=secondmate" \ + "yolo=off" \ + "remote_host=remote-mac" \ + "remote_root=/remote/root" \ + "remote_backend=herdr" \ + "remote_herdr_session=fm-remote" \ + "remote_target=fm-remote:w1:p1" + cat > "$home/data/secondmates.md" <, or empty. +pending_record() { # + find "$1/state/pending-replies" -maxdepth 1 -type f ! -name '.*' 2>/dev/null | head -1 +} + +drain_out() { # + FM_STATE_OVERRIDE="$1/state" "$DRAIN" 2>/dev/null +} + +test_remote_delivered_unconfirmed_is_not_failure() { + local dir fb log ssh_log home rc err rec + dir="$TMP_ROOT/remote-du"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log"; ssh_log="$dir/ssh.log"; : > "$ssh_log" + home=$(setup_remote_home remote-du) + + : > "$log" + env PATH="$fb:$PATH" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + FM_SSH_BIN="$fb/fake-ssh" FM_SSH_LOG="$ssh_log" FM_FAKE_SSH_RC=3 \ + FM_FAKE_SSH_STDERR='fm-send: text delivered to fm-remote:w1:p1 but submission is unconfirmed (verdict=pending; tried meta=/remote/home/state/fm-remote:w1:p1.meta; metadata window/terminal lookup; backend=herdr; endpoint=verified)' \ + "$SEND" rsm "please rename the metric" >"$dir/out" 2>"$dir/err"; rc=$? + err=$(cat "$dir/err") + expect_code 0 "$rc" "a delivered-unconfirmed remote send must not exit as a failure" + assert_grep 'fm-remote-entrypoint.sh' "$ssh_log" "the steer should cross the remote transport" + assert_contains "$err" "delivered to remote secondmate rsm" \ + "the outcome must be reported as delivered" + assert_not_contains "$err" "text not sent" "a delivered steer must not read as not sent" + assert_not_contains "$err" "not submitted" "a delivered steer must not read as not submitted" + assert_not_contains "$err" "error: text" "a delivered steer must not carry an error-styled report" + assert_not_contains "$err" "verdict=pending" \ + "the inner leg's unconfirmed diagnostics must be held back on a delivered outcome" + + rec=$(pending_record "$home") + [ -n "$rec" ] || fail "the pending-reply expectation must survive a delivered-unconfirmed send" + [ -n "$(grep '^delivered_epoch=' "$rec" | cut -d= -f2-)" ] \ + || fail "a delivered-unconfirmed send must mark the expectation delivered: $(cat "$rec")" + [ "$(grep '^phase=' "$rec" | tail -1 | cut -d= -f2-)" = awaiting_report ] \ + || fail "a delivered-unconfirmed send must leave the expectation awaiting its report: $(cat "$rec")" + pass "fm-send remote: delivered-unconfirmed reports delivered, exits 0, keeps the expectation armed" +} + +test_remote_real_failure_still_fails() { + local dir fb log ssh_log home rc err + dir="$TMP_ROOT/remote-fail"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log"; ssh_log="$dir/ssh.log"; : > "$ssh_log" + home=$(setup_remote_home remote-fail) + + : > "$log" + env PATH="$fb:$PATH" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + FM_SSH_BIN="$fb/fake-ssh" FM_SSH_LOG="$ssh_log" FM_FAKE_SSH_RC=1 \ + FM_FAKE_SSH_STDERR='error: remote secondmate rsm endpoint metadata is invalid; refusing access until it is explicitly migrated' \ + "$SEND" rsm "please rename the metric" >"$dir/out" 2>"$dir/err"; rc=$? + err=$(cat "$dir/err") + [ "$rc" -ne 0 ] || fail "a genuinely failed remote send must exit nonzero" + assert_contains "$err" "error: text not sent to remote:rsm" \ + "a real remote failure must still report a real error" + assert_contains "$err" "endpoint metadata is invalid" \ + "a real remote failure must replay the remote leg's own stderr" + [ -z "$(pending_record "$home")" ] \ + || fail "a failed send must discard its undelivered expectation" + pass "fm-send remote: a real remote failure still fails loudly with the remote diagnostics" +} + +test_remote_transport_unknown_preserves_expectation() { + local dir fb log ssh_log home rc err rec + dir="$TMP_ROOT/remote-255"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log"; ssh_log="$dir/ssh.log"; : > "$ssh_log" + home=$(setup_remote_home remote-255) + + : > "$log" + env PATH="$fb:$PATH" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + FM_SSH_BIN="$fb/fake-ssh" FM_SSH_LOG="$ssh_log" FM_FAKE_SSH_RC=255 \ + "$SEND" rsm "please rename the metric" >"$dir/out" 2>"$dir/err"; rc=$? + err=$(cat "$dir/err") + [ "$rc" -ne 0 ] || fail "an unknown-completion transport loss must exit nonzero" + assert_contains "$err" "delivery to remote secondmate rsm is unknown" \ + "transport loss must be reported as unknown delivery, not silently dropped" + rec=$(pending_record "$home") + [ -n "$rec" ] || fail "transport loss must preserve the expectation for reconciliation" + [ "$(grep '^phase=' "$rec" | tail -1 | cut -d= -f2-)" = delivery_unknown ] \ + || fail "transport loss must move the expectation to delivery_unknown: $(cat "$rec")" + pass "fm-send remote: ssh 255 still refuses loudly and preserves the expectation as delivery_unknown" +} + +test_remote_delivered_unconfirmed_closes_resolve_key() { + local dir fb log ssh_log home rc out + dir="$TMP_ROOT/remote-key"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log"; ssh_log="$dir/ssh.log"; : > "$ssh_log" + home=$(setup_remote_home remote-key) + printf 'needs-decision [key=upgrade-window]: tonight or the weekend\n' > "$home/state/rsm.status" + + : > "$log" + env PATH="$fb:$PATH" \ + FM_ROOT_OVERRIDE="$ROOT" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + FM_SSH_BIN="$fb/fake-ssh" FM_SSH_LOG="$ssh_log" FM_FAKE_SSH_RC=3 \ + "$SEND" rsm --resolve-key upgrade-window "the weekend, freeze Friday" >/dev/null 2>&1; rc=$? + expect_code 0 "$rc" "a delivered-unconfirmed remote answer must not exit as a failure" + grep -F 'resolved [key=upgrade-window]: answered: the weekend, freeze Friday' "$home/state/rsm.status" >/dev/null \ + || fail "a delivered-unconfirmed remote answer must close the decision: $(cat "$home/state/rsm.status")" + out=$(drain_out "$home") + if printf '%s' "$out" | grep -F 'OPEN DECISIONS' >/dev/null; then + fail "the answered decision still lists as open after a delivered-unconfirmed answer: $out" + fi + pass "fm-send remote: a delivered-unconfirmed answer closes its --resolve-key decision" +} + +test_local_pending_reports_delivered_unconfirmed() { + local dir fb log home rc err + dir="$TMP_ROOT/local-pending"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log" + home=$(setup_home local-pending) + fm_write_meta "$home/state/t1.meta" "window=sess:fm-t1" "kind=ship" + + : > "$log" + env PATH="$fb:$PATH" FM_FAKE_TMUX_PENDING=1 \ + FM_ROOT_OVERRIDE="$home" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + "$SEND" t1 "steer text" >"$dir/out" 2>"$dir/err"; rc=$? + err=$(cat "$dir/err") + expect_code 3 "$rc" "an unconfirmed local submit must exit with the delivered-unconfirmed status" + assert_contains "$err" "submission is unconfirmed" \ + "the unconfirmed local submit must be described honestly" + assert_not_contains "$err" "not submitted" \ + "an unconfirmed local submit must not claim the text was not submitted" + assert_not_contains "$err" "error:" \ + "an unconfirmed local submit must not carry an error-styled report" + pass "fm-send local: an unconfirmed submit exits 3 with an honest non-error report" +} + +test_local_pending_does_not_close_resolve_key() { + local dir fb log home rc out + dir="$TMP_ROOT/local-pending-key"; mkdir -p "$dir" + fb=$(make_stubs "$dir"); log="$dir/send.log" + home=$(setup_home local-pending-key) + fm_write_meta "$home/state/t2.meta" "window=sess:fm-t2" "kind=ship" + printf 'blocked [key=creds]: need the deploy token\n' > "$home/state/t2.status" + + : > "$log" + env PATH="$fb:$PATH" FM_FAKE_TMUX_PENDING=1 \ + FM_ROOT_OVERRIDE="$home" FM_HOME="$home" FM_SEND_LOG="$log" FM_SEND_SETTLE=0 \ + "$SEND" t2 --resolve-key creds "token is in the vault now" >/dev/null 2>&1; rc=$? + expect_code 3 "$rc" "an unconfirmed local answer must exit with the delivered-unconfirmed status" + if grep -F 'resolved' "$home/state/t2.status" >/dev/null; then + fail "an unconfirmed local answer must not close the decision: $(cat "$home/state/t2.status")" + fi + out=$(drain_out "$home") + printf '%s' "$out" | grep -F '[key=creds]' >/dev/null \ + || fail "the blocker must stay open after an unconfirmed local answer: $out" + pass "fm-send local: an unconfirmed submit still never closes a --resolve-key decision" +} + +test_remote_delivered_unconfirmed_is_not_failure +test_remote_real_failure_still_fails +test_remote_transport_unknown_preserves_expectation +test_remote_delivered_unconfirmed_closes_resolve_key +test_local_pending_reports_delivered_unconfirmed +test_local_pending_does_not_close_resolve_key + +echo "all fm-send-remote-delivery tests passed" diff --git a/tests/fm-shared-captain-inheritance.test.sh b/tests/fm-shared-captain-inheritance.test.sh index 59137b278f7..904e8887b4c 100755 --- a/tests/fm-shared-captain-inheritance.test.sh +++ b/tests/fm-shared-captain-inheritance.test.sh @@ -249,7 +249,7 @@ SH cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' '0.1.25' + printf '%s\n' '0.1.29' exit 0 fi exit 0 diff --git a/tests/fm-startup-memory-budget.test.sh b/tests/fm-startup-memory-budget.test.sh index 3f6ed0624af..625444298d5 100755 --- a/tests/fm-startup-memory-budget.test.sh +++ b/tests/fm-startup-memory-budget.test.sh @@ -27,7 +27,7 @@ SH cat > "$fakebin/quota-axi" <<'SH' #!/usr/bin/env bash if [ "${1:-}" = --version ]; then - printf '%s\n' 'quota-axi 0.1.25 (fake)' + printf '%s\n' 'quota-axi 0.1.29 (fake)' fi exit 0 SH