Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 9 additions & 6 deletions .agents/skills/harness-adapters/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,8 @@ Use that value for interrupt, exit, resume, and skill-invocation facts.

The primary integrations for `claude`, `codex`, `opencode`, `pi`, `pi-signed`, and `grok` have empirically validated hook paths for the "no turn ends blind" guard.
`claude` and `codex` block directly through Stop hooks that preserve exit status 2 and stderr from `bin/fm-turnend-guard.sh`.
`opencode`, `pi`, `pi-signed`, and `grok` expose passive lifecycle callbacks for this purpose, so their tracked primary adapters force one bounded follow-up or resume when the shared predicate blocks.
`opencode`, `pi`, and `pi-signed` expose passive lifecycle callbacks and force one bounded follow-up when the shared predicate blocks.
Grok selects native blocking or its pre-native bounded resume fallback from the exact running Stop payload; [`docs/turnend-guard.md`](../../../docs/turnend-guard.md) owns that contract.
Kimi is outside the primary turn-end guard scope, while `docs/turnend-guard.md` owns its separate guarded global hook for crew wake signals.
The exact hook files, commands, scoping rules, and fail-open tradeoffs are owned by `docs/turnend-guard.md`.
`docs/verification/supervision.md` "Turn-end guard" owns active validation evidence.
Expand Down Expand Up @@ -126,6 +127,8 @@ The supported launch-profile flags below are verified locally; each row records
| opencode | `--model <provider/model>` | none for firstmate's interactive launch | Verified on opencode 1.17.6. `opencode run` has `--variant`, but firstmate launches the interactive `opencode --prompt` path, which has no verified effort flag. |
| kimi | `--model <model>` | none | Verified 2026-07-25 on Kimi Code CLI 0.29.1. |

The concrete `harness` field owns adapter identity independently of the model provider: `harness=pi` with `model=xai/grok-*` is Pi using xAI, not `harness=grok`, and does not require Grok CLI login; `harness=grok` remains the standalone Grok Build CLI adapter.

### Model support discovery

Treat model and provider knowledge as current source-of-truth discovery, not as a permanent namespace or provider mapping.
Expand Down Expand Up @@ -343,13 +346,13 @@ This keeps the hook outside the worktree, needs no trust grant, and writes only
`fm-teardown` removes the worktree pointer before returning a pooled worktree.
Secondmate spawns skip the pointer (idle panes are healthy, no stale-pane detection for them).

**Primary-session guard fact (verified 2026-07-08, Grok 0.2.91).**
**Primary-session guard fact (verified 2026-07-28, Grok 0.2.112 and 0.2.73).**
The firstmate PRIMARY's own `.grok/hooks/fm-primary-turnend-guard.json` invokes `bin/fm-turnend-guard-grok.sh`.
Grok Stop hooks are passive for this purpose: exit 2 does not make the model continue.
The adapter therefore runs the shared predicate and, when it returns 2, forces one same-session follow-up with `grok --resume <sessionId> -p <guard-reason>` while setting `GROK_TURNEND_GUARD_ACTIVE=1` so the nested Stop hook does not recurse.
It does not pass `--permission-mode`, so the passive hook cannot escalate the primary session's tool permissions.
Grok 0.2.112 exposes native same-process Stop continuation in its running payload, while the genuine pre-native 0.2.73 payload omits that capability and still needs one guarded `grok --resume`.
The exact adaptive and malformed-input contract is owned by `docs/turnend-guard.md`.
The tracked Claude Stop hooks skip themselves under `GROK_AGENT`, because Grok also loads Claude-compatible project settings and otherwise creates a second blocking path.
Project-local Grok hooks require folder trust, verified with launch-time `--trust`; if the primary firstmate checkout is not trusted for Grok hooks, this primary guard fails open and `fm-guard.sh` remains the next-command alarm.
Grok's primary watcher protocol is Claude-shaped background-notify around `bin/fm-watch-arm.sh`; the passive Stop hook is only a backstop for blind turn ends.
Grok's primary watcher protocol remains background-notify around `bin/fm-watch-arm.sh`; native Stop continuation does not provide Pi-like extension ownership.

## kimi (VERIFIED 2026-07-25, kimi 0.29.1)

Expand Down
175 changes: 35 additions & 140 deletions .agents/skills/quota-array-dispatch/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,159 +12,54 @@ metadata:
# quota-array-dispatch

This skill is the single owner of the pace-aware profile-array selection procedure.
The concise always-loaded intake boundary remains in `AGENTS.md` section 4.
`docs/configuration.md` owns the `config/crew-dispatch.json` schema only.
`AGENTS.md` section 4 owns the always-loaded intake boundary, load trigger, malformed-config refusal, every-candidate accounting, and strongest-reasoning/tie safety rules.
`harness-adapters` owns harness verification, model/provider discovery, and effort fallback.
`quota-axi` remains data-only and never recommends a route.
Firstmate owns the judgment.
Do not add a daemon, opaque composite score, routing wrapper, hard-coded model-specific policy, or producer-side route recommendation.

## When to load
## Collect facts

Load this skill whenever a matched dispatch rule or the configured default resolves to a profile array (more than one candidate), before choosing the concrete `--harness`, `--model`, and `--effort` passed to `fm-spawn`.
Keep using `harness-adapters` for harness verification, model/provider discovery, and effort fallback.
Run `quota-axi --json` once per intake and reuse that snapshot for every candidate.
For each candidate, preserve explicit `harness`, `model`, and `provider`; `harness-adapters` owns identity, and model/provider never infer harness:

## Intake boundary this skill does not relax
- task/profile fit and required reasoning class
- raw applicable headroom (`effectivePercentRemaining` or tightest applicable percentage)
- effective pace, signed reserve per window, and worst reserve (`worstReservePercentPoints` or minimum signed reserve)
- whether applicable windows/summary are ahead, or pace is `unknown`
- schema note when pace fields are absent

1. Explicit per-task captain overrides still win over configured profiles.
2. Configured profile matching precedence is unchanged: best-fit rule, then configured default, then static crewmate harness.
3. Malformed `config/crew-dispatch.json` remains an actionable error; never select around it.
4. Every configured candidate in the matched array must be accounted for.
5. If any harness/model/provider relationship, applicable quota data, or interpretation cannot be established, stop and report that candidate instead of omitting it, guessing, falling back, or calling the result quota-informed.
6. When every candidate is tight, preserve the captain's strongest-reasoning class rather than silently downgrading it solely to conserve quota; stop and report the tight choice if that class cannot proceed.
7. Genuine ties must remain free of array-order or harness bias.
Stale raw windows are diagnostic, never headroom.
Read all windows named by `boundedBy`, `limitingWindowIds`, `aheadWindowIds`, `behindWindowIds`, `onPaceWindowIds`, and `unknownWindowIds`.

## Collect inspectable facts for every candidate
## Pace semantics

For each candidate profile:
`reservePercentPoints = percentRemaining - timeRemainingPercent`.
Negative reserve means usage is ahead of reset pace and creates conservation pressure.
Positive reserve means usage is behind reset pace.
`on_pace` is neutral.
Conservation pressure is present for effective pace status `ahead`, effective pace status is `mixed` and any `aheadWindowIds` remain, or a bounding window is `ahead`.
`unknown` is valid explicit uncertainty from quota-axi, not parser failure or permission to assume health.

1. Establish the harness/model/provider relationship from current authoritative discovery owned by `harness-adapters`.
Fail loudly on an unresolved relationship.
2. Run `quota-axi --json` once per intake and reuse that snapshot for every candidate.
3. Require a current provider report with known quota semantics and a known applicable effective-availability record for that candidate's provider and model scope.
Stale raw windows remain diagnostic evidence only and are never current headroom.
4. Read every bounding window relevant to that candidate, including windows named by `boundedBy`, `limitingWindowIds`, `aheadWindowIds`, `behindWindowIds`, `onPaceWindowIds`, and `unknownWindowIds` on the effective record.
5. Record these inspectable facts, never a hidden score:
- task/profile fit
- reasoning class required by the captain request or task ambiguity
- raw applicable headroom (`effectivePercentRemaining` or the tightest applicable remaining percentage)
- effective pace status when present
- signed reserve for each applicable window and the effective worst reserve when present
- whether any applicable window or effective summary is ahead of reset
- whether any applicable pace is `unknown`
- schema compatibility note when pace fields are absent
## Selection order

## Pace signals
Apply only among candidates satisfying required fit and strongest reasoning class.
Never use pace or raw headroom to silently replace that reasoning class.

quota-axi `schemaVersion` 3 window pace uses:

- `reservePercentPoints = percentRemaining - timeRemainingPercent`
- Negative reserve means usage is ahead of reset pace and creates conservation pressure.
- Positive reserve means usage is behind reset pace.
- `on_pace` is neutral.

Effective-availability pace summaries may report `ahead`, `behind`, `on_pace`, `mixed`, or `unknown`.

Treat conservation pressure as present when:

- effective pace status is `ahead`, or
- effective pace status is `mixed` and any `aheadWindowIds` remain, or
- any applicable bounding window itself has pace status `ahead`.

An effective `mixed` result is never healthy merely because one window is behind.
Any remaining `aheadWindowIds` keep conservation pressure.

Signed reserve comparison uses the worst applicable reserve, preferring the producer field `worstReservePercentPoints` when present and otherwise the minimum signed reserve across applicable bounding windows.

## Selection procedure

Apply these steps only among candidates that already satisfy required task/profile fit and the strongest reasoning class the request genuinely needs.
Never use pace or raw headroom to silently replace that reasoning class with a weaker one.

1. **Unresolved relationship or quota data**
Stop and report the blocked candidate.
2. **Strongest-reasoning / all-tight**
If every remaining candidate is tight, keep the strongest-reasoning class and either dispatch inside that class or stop and report that the tight choice cannot proceed.
Do not conserve quota through an unapproved downgrade.
3. **Conservation pressure vs sustainable pace**
When fit and reasoning class are comparable, prefer a candidate without ahead-of-reset conservation pressure over one with conservation pressure, even when the pressured candidate has somewhat higher raw remaining percentage.
4. **Among pressured candidates**
Prefer the least-negative worst applicable reserve.
Example: worst reserve `-4` is safer than `-18` when other inspectable facts are comparable.
5. **Among sustainable candidates**
Use known behind/on-pace evidence plus raw headroom transparently.
1. Unresolved relationship or quota: stop and report the tuple and concrete evidence.
2. All-tight: keep strongest reasoning; dispatch inside it or report if blocked.
3. Comparable fit/reasoning: prefer no ahead pressure over pressure, even with higher raw headroom.
4. Among pressured candidates, prefer the least-negative worst applicable reserve.
5. Sustainable candidates: use known pace plus raw headroom.
Prefer known sustainable evidence over `unknown` when comparable.
Do not collapse those facts into an opaque composite score.
Prefer known sustainable evidence over `unknown` pace when otherwise comparable.
Between known sustainable candidates, prefer the clearly better inspectable pair of pace reserve and raw headroom; state both facts in the choice rationale.
6. **Unknown pace**
`unknown` is valid explicit uncertainty from quota-axi, not a parser failure and not permission to assume the window is healthy or exhausted.
Inspect `unknownWindowIds` and each window's pace `reason` so the rationale preserves the producer's stated uncertainty.
Prefer known sustainable evidence when otherwise comparable.
If the dispatch choice materially hinges on unresolved pace, report the uncertainty rather than inventing a conclusion.
7. **Absent pace / older schema**
`schemaVersion` 2 payloads or missing pace fields must degrade explicitly and safely.
Do not crash, fabricate pace, or silently reinterpret absence as healthy/`on_pace`.
Compare raw applicable headroom only, using known effective availability rather than stale or isolated window percentages, state that pace is unavailable, and keep every other safety rule above.
8. **Genuine ties**
If every inspectable selection fact is equal, stop and report every tied candidate for captain choice.
6. If unresolved pace changes the choice, report uncertainty.
7. Absent pace or older schema: do not crash, fabricate pace, or treat absence as healthy/`on_pace`.
Compare raw headroom only, state pace is unavailable, and keep safety rules.
8. Genuine ties: stop and report every tied candidate for captain choice.
Do not select by array order, harness name, or another arbitrary identity ordering.
Report duplicate concrete profiles as a configuration error.

The intake rationale must name the inspectable facts used for every candidate.
Name the inspectable facts used for every candidate.
After selecting, check auth only through that tuple's surface; another harness CLI cannot block it.
A blocked credential report must name `harness`, `model`, authentication surface, and concrete failure evidence; never emit a bare `Grok unauthenticated` statement.
Never conclude with an unexplained "best quota" label.

## Acceptance scenarios

These scenarios are normative examples of the procedure above.

### Higher raw quota but materially ahead vs lower raw quota on/behind pace

Candidate A has higher `effectivePercentRemaining` but conservation pressure from an ahead bounding window.
Candidate B has lower raw headroom, no conservation pressure, and known behind or on-pace evidence.
Choose B when fit and reasoning class are comparable.

### Mixed effective pace with an ahead bound

Effective pace status is `mixed` and `aheadWindowIds` is non-empty.
Treat the candidate as conservation-pressured even if another window is behind or on pace.

### Both candidates ahead with different worst reserves

Both candidates have conservation pressure.
Choose the least-negative worst applicable reserve when fit and reasoning class are comparable.

### Known sustainable versus unknown

Candidate A has known behind or on-pace evidence.
Candidate B has comparable fit, reasoning class, and raw headroom but `unknown` pace.
Prefer A.
If the only way to prefer one side depends on unresolved pace and no known sustainable candidate remains, report the uncertainty.

### Every candidate tight while strongest-reasoning applies

All candidates are tight on real headroom.
Keep the strongest reasoning class required by the request.
Do not pick a weaker class only to save quota.
Dispatch inside that class or stop and report that the tight strongest-class choice cannot proceed.

### Genuine tie without array-order or harness bias

Two candidates match on fit, reasoning class, conservation pressure, worst reserve, pace class, raw headroom, and unknown flags.
Choosing either array order or a standing harness preference is forbidden.
Stop and report both tied candidates for captain choice.

### schemaVersion 2 or absent-pace compatibility

Older quota-axi output or missing pace fields still allow array resolution.
Compare raw headroom only, state that pace is unavailable, and do not invent ahead/behind/on_pace.

## Sanitized producer shape

Validate consumers against a sanitized `schemaVersion` 3 shape derived from quota-axi 0.1.15:

- top level: `schemaVersion`, `generatedAt`, `providers[]`
- each provider: `provider`, `state`, `windows[]`, and optional `quotaSemantics` with `status` and `effectiveAvailability[]`
- each window: `id`, `label`, `kind`, and optional `percentRemaining` and `pace`; pace has `status` plus optional `reason`, `timeRemainingPercent`, and `reservePercentPoints`
- each effective-availability entry: `scope`, `status`, `boundedBy`, optional `effectivePercentRemaining`, optional `limitingWindowIds`, and optional pace summary
- each effective pace summary: `status` plus optional `aheadWindowIds`, `behindWindowIds`, `onPaceWindowIds`, `unknownWindowIds`, `worstReservePercentPoints`, and `worstReserveWindowId`

Never persist live provider balances, reset timestamps, account identifiers, or other private account details in tracked fixtures.
Loading
Loading