Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
140 changes: 72 additions & 68 deletions .agents/skills/quota-array-dispatch/SKILL.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/skills/stuck-crewmate-recovery/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ The target window's harness is recorded as `harness=` in `state/<id>.meta`.
This procedure covers ordinary `kind=ship` and `kind=scout` direct reports.
Load `secondmate-provisioning` instead for `kind=secondmate` recovery.

For a REMOTE secondmate, `fm-crew-state`'s `unknown`/`worktree gone` and `fm-send`'s `remote send failed`/`delivery unconfirmed` verdicts are unreliable and routinely false-negative; do not conclude the mate is dead or the send failed from those alone, confirm against the actual remote pane first.
For a REMOTE secondmate, `fm-crew-state` and `fm-peek` read the actual remote endpoint over `fm-on.sh`, and `fm-send` reports a delivered-with-pending-confirmation steer as delivered (their headers own the contracts); an `unknown-remote` read or unreachable-host failure means the remote state could not be read, never that the mate is dead or the send failed.
Recover a genuinely stuck remote mate only through `bin/fm-spawn.sh <id> --secondmate`, never raw herdr pane close/kill surgery, which strands the endpoint binding.

Treat the digest's endpoint result as a presence signal, not proof that the task's work or validation run is gone.
Expand Down
10 changes: 5 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Hard rules, in priority order:
If work failed, say so plainly with the evidence.

You may maintain this repo's private operational state directly.
Shared tracked material is `AGENTS.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `fork-divergences.json`, `.github/workflows/`, `bin/`, `.agents/skills/`, and public `skills/`.
Shared tracked material is `AGENTS.md`, `GROK_BOT.md`, `README.md`, `CONTRIBUTING.md`, `.tasks.toml`, `fork-divergences.json`, `.github/workflows/`, `bin/`, `.agents/skills/`, and public `skills/`.
When any crewmate is live, delegate changes to shared tracked material rather than competing with supervision; when the fleet is empty, firstmate may change it directly.
This repo is a shared template, while `.env`, `data/`, `state/`, `config/`, `projects/`, and `.no-mistakes/` are captain-private and gitignored.
Ship shared tracked changes through this repo's no-mistakes pipeline and PR path, with the same merge authority as any other project.
Expand Down Expand Up @@ -102,16 +102,16 @@ If static `config/crew-harness` or `config/secondmate-harness` names an unverifi
`docs/configuration.md` owns dispatch-profile and runtime-backend schemas, `bin/fm-harness.sh` owns static resolution, and `bin/fm-spawn.sh` owns launch flags and fail-closed validation.
When dispatch profiles exist, consult them at every crewmate or scout intake and pass the resolved concrete profile required by `fm-spawn`.
Routing precedence is an explicit per-task captain override, then the best-fit configured rule, then the configured default, then the static crewmate harness.
Firstmate alone resolves a matched profile array: run `quota-axi --json` at that intake, evaluate every configured candidate against that current output, and choose with inspectable effective headroom and usable runway, using pace and reserve only later when needed.
Account for every candidate with the catalog evidence, provider relationship, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, and the headroom, runway, and later pace or reserve evidence used in selection; never omit a candidate, guess, fall back silently, or call the result quota-informed without them.
Firstmate alone resolves a matched profile array: begin with `quota-axi`'s default TOON at that intake, using the skill's narrow TOON-then-`--json` fallback only for genuine ambiguity, evaluate every configured candidate against that current output, and choose with inspectable `spendPriority` as the one quota-perspective ranker after the skill's eligibility, reasoning-class, and runway-feasibility gates.
Account for every candidate with the catalog evidence, provider relationship, applicable quota and authentication facts, remaining uncertainty, fit and reasoning class, and the spendPriority and runway evidence used in selection; never omit a candidate, guess, fall back silently, or call the result quota-informed without them.
Establish model support and provider family from that harness's own authoritative catalog, then read `quota-axi` at the granularity the vendor actually supplies: provider-level or all-model evidence applies to every model established in that family, and a named-model window bounds only that model.
Missing model-level quota, a missing authentication source, unmeasurable headroom, or unmodeled authentication is disclosed uncertainty that keeps a candidate eligible, never a credential or login escalation.
Only concrete contradictory evidence blocks a candidate, such as an authoritative catalog proving the model unsupported or proof that the credential selected for that surface is unusable; never infer a credential store, provider family, or quota mapping from a harness, model, or source name, and never launch another harness's CLI to judge a candidate.
Preserve malformed profile configuration as an actionable error rather than selecting around it.
When every candidate is tight, preserve the captain's strongest-reasoning class rather than silently downgrading it solely to conserve quota; stop and report the tight choice if that class cannot proceed.
Break genuine evidence ties without array-order or harness bias.
`quota-axi` owns how model or product windows relate to bounding account windows and remains data-only.
Load `quota-array-dispatch` before choosing among a matched profile array; that skill is the single owner of the completion-aware selection procedure.
Load `quota-array-dispatch` before choosing among a matched profile array; that skill is the single owner of the TOON-first spendPriority selection procedure.
The generic effort fallback and its precedence are owned by `harness-adapters`: explicit captain and standing configured effort win; otherwise use low for well-understood explicit work, xhigh for ambiguous investigation or design, intermediate levels proportionally, and never max without explicit captain preference.
Do not add model-specific versions of that policy.

Expand Down Expand Up @@ -440,7 +440,7 @@ These skills are not captain-invocable; load them only at their precise triggers
- `bootstrap-diagnostics` - load whenever the session-start digest's bootstrap or network-checks section prints an actionable diagnostic line (`MISSING:`, `MISSING_MANUAL:`, `BACKEND_INVALID:`, `NEEDS_GH_AUTH`, `TANGLE:`, `STARTUP_MEMORY_BUDGET:`, `CREW_DISPATCH: invalid`, `FLEET_SYNC:`, `NETWORK_CHECKS:`, `PR_CHECK_MIGRATION:`, `SECONDMATE_SYNC:`, `SECONDMATE_LIVENESS:`, `SECONDMATE_HANDOFF:`, `NUDGE_SECONDMATES:`, or `FMX:`); silence and `BOOTSTRAP_INFO:` need no load.
- `diagnostic-reasoning` - load before scoping a reported bug and before acting on a diagnostic report.
- `ask-user-authority` - load before deciding any ask-user finding, regardless of the project's `yolo` posture.
- `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi output.
- `quota-array-dispatch` - load before choosing among a matched crew-dispatch profile array from current quota-axi default TOON.
- `harness-adapters` - load before spawning or recovering a crewmate or secondmate, handling a trust dialog, sending a harness-specific skill invocation, interrupting or exiting an agent, resuming an exited agent, or verifying a new harness adapter.
- `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata.
- `project-management` - load before adding, creating, removing, or initializing a project.
Expand Down
47 changes: 47 additions & 0 deletions GROK_BOT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
You are Firstmate: the single agent the captain talks to.
They bring you everything; you make sure it gets done.

Other bots are your crewmates: persistent and role-based, each holding a stable charter - e.g. one for the inbox, one for documents like PDFs and decks, one for research.
Before signing on a new crewmate, check whether an existing one already covers a related charter: if a charter matches or highly overlaps, reuse that crewmate; if the overlap is only limited, sign on the new crewmate and clarify the distinction in both crewmates' charters.
Sign on a genuinely new crewmate only when no existing one fits.
When you sign one on, write into its charter that it reports its outcomes and blockers back to you (Firstmate), never to the captain directly - the captain only ever talks to you.
Delegate by messaging a crewmate; it wakes, does the work, and messages you back.

Default to handing work off.
If a job is more than one tool call, especially computer or browser work or anything that will take minutes, give it to the crewmate whose charter fits.
Do not keep that grind in this chat because you already have a login, a token, or an open page.
The computer is shared across the crew.
Browser logins persist for every bot.
A login on your screen is not a reason to do the work yourself.
Secrets are per-bot.
They do not propagate to the crew.
If a crewmate needs a credential, tell the crewmate to request it and then tell the captain to give that secret to that bot on a secure card.
Do not keep the secret and do the work yourself.
Do not paste or forward secrets in chat.
After the captain has given the secret to that bot, hand the task off and wait for the outcome.

Software and code go through a crewmate, never through you directly: sign on a crewmate per project or project area - once the captain has expressed how its charter should be set - and let that crewmate drive the code work with cursor cloud agents.
You never call a cursor cloud agent yourself.

Don't reach for subagents.
Needing one means the work is substantial, which means it belongs with a crewmate, not with you.
Subagents are a tool for crewmates to break down their own work.

Mark every task you hand off as coming from you, with a short task id, and ask for the outcome back against that id - so the crewmate routes its result and any blockers to you rather than just handling them in its own chat, and you can match a reply to the right task.
The marker is visible in the chat; that's fine.

Work asynchronously.
Delegating doesn't block you - a crewmate replies on a later turn and shows up in this chat.
So hand off, tell the captain what's under way, and relay each result as it lands.
Reserve a priority send for when something must interrupt a crewmate's current task.

When you notice crewmates making mistakes or working inefficiently, update their description to refine their behavior so your crew does better next time.

How you talk.
Address the captain as "captain" at least once in every reply - always, even when the news is bad ("Captain, that didn't work...").
Let light nautical seasoning land only when it fits naturally - an occasional "aye", "on deck", "shipshape", "under way", "ahoy" - never letting it crowd out the substance, and drop it entirely for bad news or serious findings.
Speak in outcomes and consequences, not internal mechanics.

Keep it simple for the captain.
Focus on communicating outcomes, not mechanics.
They scale by talking only to you; protect that.
57 changes: 53 additions & 4 deletions bin/fm-crew-state.sh
Original file line number Diff line number Diff line change
Expand Up @@ -16,10 +16,17 @@
# fixed mapping logic, no heuristics and no LLM. Output is one stable, parseable,
# token-tight line firstmate can read every heartbeat:
#
# state: <working|parked|done|blocked|paused|failed|unknown> · source: <run-step|pane|status-log|none> · <detail>
# state: <working|parked|done|blocked|paused|failed|unknown> · source: <run-step|pane|status-log|remote-endpoint|none> · <detail>
#
# Logic, in order:
# 1. Resolve worktree + backend target + kind from state/<id>.meta.
# 1. Resolve worktree + backend target + kind from state/<id>.meta. A meta
# recording remote_host= is a remote secondmate: its worktree and endpoint
# live on that host, so the local worktree and pane reads are skipped and
# the remote host is asked for the endpoint's recovery-grade state
# (fm-on.sh + fm-remote-secondmate-control.sh state). alive falls through
# to the routed status log; dead/missing report the remote verdict; an
# unreachable or unreadable remote reports unknown-remote, never a false
# gone/dead.
# 2. Matching no-mistakes run for this crew's branch AND current code identity,
# active or terminal (from `axi status`, or the coarse `no-mistakes runs`
# fallback)? Branch name alone is not enough: a historical run on a reused
Expand Down Expand Up @@ -101,10 +108,13 @@ meta_value() { # <key>
WT=$(meta_value worktree)
KIND=$(meta_value kind)
HARNESS=$(meta_value harness)
REMOTE_HOST=$(meta_value remote_host)
[ -n "$KIND" ] || KIND=ship

# A torn-down (or never-created) worktree has no current state to read.
if [ -z "$WT" ] || [ ! -d "$WT" ]; then
# A torn-down (or never-created) worktree has no current state to read. A
# remote secondmate's recorded worktree is a path on ITS host, so the local
# probe proves nothing for it - the remote arm below reads the true source.
if [ -z "$REMOTE_HOST" ] && { [ -z "$WT" ] || [ ! -d "$WT" ]; }; then
emit unknown none "worktree gone (torn down?)"
fi

Expand Down Expand Up @@ -138,6 +148,45 @@ map_log_state() { # <line>
LOG_LINE=$(log_last_line || true)
LOG_VERB=$(status_line_verb "$LOG_LINE")

# --- remote secondmate: the true source is the remote endpoint ---------------
# A remote mate's recorded worktree and backend target live on its own host, so
# the local worktree probe above and the local pane reads below would misreport
# a healthy remote mate as gone or dead. Ask the remote host for the endpoint's
# recovery-grade state over the same fm-on.sh transport fm-send uses, then read
# current activity from the routed status log exactly as for a local
# secondmate (an idle endpoint is healthy for a secondmate either way). An
# unreachable host or unreadable endpoint is reported as unknown-remote -
# explicitly NOT proof of death - so a transport blip never reads as a torn
# down or dead mate; only the remote host's own dead/missing verdict may say
# the endpoint is actually gone.
if [ -n "$REMOTE_HOST" ]; then
if ! REMOTE_STATE=$(FM_HOME="$FM_HOME" "$SCRIPT_DIR/fm-on.sh" "$ID" \
fm-remote-secondmate-control.sh state "$ID" < /dev/null 2>/dev/null); then
REMOTE_STATE=
fi
REMOTE_STATE=$(printf '%s\n' "$REMOTE_STATE" | tail -1)
case "$REMOTE_STATE" in
alive)
if [ -n "$LOG_VERB" ]; then
LOG_STATE=$(map_log_state "$LOG_LINE")
if [ "$LOG_STATE" != unknown ]; then
emit "$LOG_STATE" status-log "$(status_line_note "$LOG_LINE")${SEP}remote endpoint alive on $REMOTE_HOST"
fi
fi
emit unknown remote-endpoint "alive on $REMOTE_HOST (an idle secondmate is healthy)"
;;
dead|missing)
emit unknown remote-endpoint "remote endpoint $REMOTE_STATE on $REMOTE_HOST"
;;
'')
emit unknown remote-endpoint "unknown-remote: $REMOTE_HOST unreachable or endpoint unreadable (not proof of death)"
;;
*)
emit unknown remote-endpoint "unknown-remote: endpoint state '$REMOTE_STATE' on $REMOTE_HOST (not proof of death)"
;;
esac
fi

# pane_readable is consulted ONLY in the no-run fallback below. The run-step path
# stays authoritative regardless of pane liveness - judge by the run-step, not the
# shell - so a finished crew whose endpoint has closed still reports its run-step
Expand Down
40 changes: 34 additions & 6 deletions bin/fm-inactive-reconcile.sh
Original file line number Diff line number Diff line change
Expand Up @@ -9,9 +9,17 @@
# not a watcher, daemon, PR poll, or forge client of its own.
# `scan` evaluates at most once per FM_INACTIVE_RECONCILE_SECS (default 900,
# valid 60..1800) per home, except that --startup performs the same cheap scan
# immediately during a locked session start. Each scan has an aggregate
# FM_INACTIVE_RECONCILE_BUDGET_SECS bound (default 10, valid 1..30) and resumes
# after its last visited child on the next scan.
# immediately during a locked session start. Each scan uses an aggregate
# FM_INACTIVE_RECONCILE_BUDGET_SECS deadline (default 10, valid 1..30) and
# resumes after its last visited child on the next scan.
# The scan enforces that budget itself through a whole-second deadline, and the
# first due child of every scan is always visited with at least a one-second
# state-read bound: whole-second arithmetic can otherwise round a small budget
# to zero mid-scan, and an invocation that exits having visited nothing would
# advance the durable cursor past a child it never examined. A process-group
# kill one second after the budget remains as a backstop for a scan wedged in
# an unbounded wait (for example a live-held wake-queue lock), so the clean
# deadline path is not racing its own backstop.
#
# It considers only a direct ordinary crewmate whose newest meta, status, or
# turn-ended mtime is older than that interval and whose last status is not
Expand Down Expand Up @@ -377,8 +385,13 @@ reconcile_direct_child() { # <id> <meta> <secondmate-id-or-empty> <timeout>
return "$rc"
}

# SCAN_FIRST_VISIT_PENDING is armed by scan() before its passes. The deadline
# below is whole-second arithmetic, so a small budget can quantize to zero
# between the deadline computation and these checks; without the guaranteed
# first visit, such a scan would return 3 having examined no child at all while
# write_scan_marker had already advanced the cursor past the skipped child.
scan_pass() { # <cursor> <after|through> <deadline> <secondmate-id-or-empty>
local cursor=$1 range=$2 deadline=$3 self=${4:-} meta id remaining rc
local cursor=$1 range=$2 deadline=$3 self=${4:-} meta id remaining rc first
for meta in "$STATE"/*.meta; do
[ -f "$meta" ] || continue
id=$(basename "$meta" .meta)
Expand All @@ -387,9 +400,19 @@ scan_pass() { # <cursor> <after|through> <deadline> <secondmate-id-or-empty>
after) [ -z "$cursor" ] || [[ "$id" > "$cursor" ]] || continue ;;
through) [ -n "$cursor" ] && [[ "$id" > "$cursor" ]] && continue ;;
esac
[ "$(date +%s)" -lt "$deadline" ] || return 3
first=0
if [ "${SCAN_FIRST_VISIT_PENDING:-0}" -eq 1 ]; then
first=1
SCAN_FIRST_VISIT_PENDING=0
fi
if [ "$first" -eq 0 ]; then
[ "$(date +%s)" -lt "$deadline" ] || return 3
fi
write_scan_marker "$id" || return 1
remaining=$((deadline - $(date +%s)))
if [ "$first" -eq 1 ] && [ "$remaining" -lt 1 ]; then
remaining=1
fi
[ "$remaining" -gt 0 ] || return 3
reconcile_direct_child "$id" "$meta" "$self" "$remaining" || {
rc=$?
Expand Down Expand Up @@ -420,6 +443,7 @@ scan() {
fi
fi
deadline=$(( $(date +%s) + FM_INACTIVE_RECONCILE_BUDGET_SECS ))
SCAN_FIRST_VISIT_PENDING=1
scan_pass "$cursor" after "$deadline" "$self" || rc=$?
if [ "$rc" -eq 0 ] && [ -n "$cursor" ]; then
scan_pass "$cursor" through "$deadline" "$self" || rc=$?
Expand Down Expand Up @@ -461,7 +485,11 @@ case "$mode" in
--startup) startup=1 ;;
*) printf 'usage: fm-inactive-reconcile.sh scan [--startup]\n' >&2; exit 2 ;;
esac
if fm_run_timed "$FM_INACTIVE_RECONCILE_BUDGET_SECS" "$0" _scan-locked "$startup"; then
# The scan's own whole-second deadline enforces the budget; this outer
# process-group kill is only the backstop for a scan wedged outside every
# bounded section (an unbounded lock wait), so it fires one second after
# the deadline instead of racing the clean bounded exit it exists to guard.
if fm_run_timed $((FM_INACTIVE_RECONCILE_BUDGET_SECS + 1)) "$0" _scan-locked "$startup"; then
:
elif [ "$?" -ne 124 ]; then
exit 1
Expand Down
Loading
Loading