Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions .agents/skills/stuck-crewmate-recovery/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ name: stuck-crewmate-recovery
description: >-
Agent-only playbook for stuck or missing ordinary Firstmate direct reports.
Use when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer.
Also use on the inverse case: a live crewmate reporting the no-mistakes pipeline dead, unreachable, or timed out.
Reconciles recorded work before escalating from targeted inspection through safe relaunch or failure.
user-invocable: false
metadata:
Expand Down Expand Up @@ -39,6 +40,25 @@ Preserve its uncommitted changes and commits, keep the same task identity, and r
Do not use a fresh generic spawn while the recorded worktree is unaccounted for, because allocating another worktree can split one task across two copies.
If the worktree or ownership cannot be reconciled safely, leave all state intact and report the task failed or blocked with the conflicting evidence.

## A live crewmate claiming the pipeline is dead

This is the inverse of the dead-endpoint case above: the worker is alive and the pipeline it declares dead usually is too.
A drive call blocks until the next gate or outcome, far longer than a harness lets one command run, and the daemon accepts a response immediately and runs the round in the background.
So a crewmate's timed-out, killed, or errored drive call leaves it waiting on a read it never got, and the "the daemon is gone" conclusion it draws from that is a guess, not evidence.

Read the two authoritative sources yourself before believing the claim:

1. `no-mistakes daemon status` for the socket.
2. `no-mistakes axi status --run <id>` for the run, or `bin/fm-crew-state.sh <id>`, which already folds this contradiction in and reports a non-socket daemon-or-timeout `blocked:` line over a running or fixing run with fresh activity as superseded because the run is alive.

A refused connection or missing socket from `daemon status` is positive daemon-down evidence and must be escalated even if the persisted run record still says running or fixing; that record can be stale after the daemon exits.
Otherwise, if the run is still running or fixing with recent activity, the claim is wrong: steer the crewmate to reattach with `no-mistakes axi run` from its own worktree, which is safe and idempotent while the run still matches its `HEAD`, and tell it a timeout is not daemon death.
Nothing reaches the captain in that case.

Never restart, stop, or update the shared daemon on a crewmate's claim.
It is one instance serving every lane and home, so a restart kills other lanes' in-flight runs.
Only positive socket refusal or absence is a daemon-down finding; escalate that finding, or a failed run record that names a daemon error, to the captain.

## Live-endpoint escalation

Escalate in order:
Expand Down
4 changes: 2 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -450,7 +450,7 @@ The skill owns the daemon procedure; these safety facts remain inline:

### Stuck-worker trigger

Load `stuck-crewmate-recovery` after a stale wake, looping or confused pane, answered-by-brief question, unresponsive worker, or failed steer.
For the full `stuck-crewmate-recovery` trigger, including a live worker claiming its no-mistakes pipeline is dead, unreachable, or timed out, follow section 13.

## 9. Escalation and captain etiquette

Expand Down Expand Up @@ -553,7 +553,7 @@ These skills are not captain-invocable; load them only at their precise triggers
- `firstmate-orca` - load before switching to Orca, spawning or supervising Orca-backed work, smoke-testing Orca backend behavior, debugging Orca task state, or reconciling Orca-backed task metadata.
- `project-management` - load before adding, creating, removing, or initializing a project.
Cloning or registering a project is add intake and uses the same trigger.
- `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, or after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer.
- `stuck-crewmate-recovery` - load when the session-start digest reports an ordinary direct report's endpoint dead or its metadata has no window, after a stale wake, looping pane, repeated confusion, an answered-by-brief question, an unresponsive crewmate, or a failed steer, and whenever a live worker reports its no-mistakes pipeline dead, unreachable, or timed out.
- `secondmate-provisioning` - load before creating, seeding, validating, launching, handing backlog to, recovering, pushing inherited local material into, or retiring a secondmate home, and before editing `data/secondmates.md`.
- `captain-hold-lifecycle` - load before treating an investigation or visual review as complete, before ending a visual review that exposed a captain decision, when recording or routing the captain's answer, and on any `RECORD DIVERGENCE` line from the wake drain.
- `process-event-sources` - load before arming a long-polling source, before registering a deterministic condition->action watch (do X as soon as Y is true), and on any `procevent <adapter> <source-id> <sequence>` check wake.
Expand Down
26 changes: 22 additions & 4 deletions bin/fm-brief.sh
Original file line number Diff line number Diff line change
Expand Up @@ -387,8 +387,17 @@ The report is the only thing that survives, so anything worth keeping must be in
A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work.
Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=<slug>]\` if you opened it with one) as you resume.
7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving
every lane/home, so restarting it kills other lanes' in-flight pipeline runs. On ANY no-mistakes
daemon error, append \`blocked: {the daemon error}\` and stop; only firstmate manages the daemon.
every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate
manages the daemon.
Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and
\`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append
\`blocked: {the daemon error}\` and stop even when the local run record still says running or
fixing, because that record can be stale after the daemon exits. A run record failed with a
daemon error is also a real block.
Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep
going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error:
the daemon accepts \`respond\` immediately and runs the round in the background, so a killed or
timed-out call was only waiting for a read while the run kept working.

$INBOX_SECTION

Expand Down Expand Up @@ -469,8 +478,17 @@ $ASK_USER_BLOCK
A decision or blocker you opened stays open until a \`resolved\` line carrying its exact key lands; a later \`done:\` or \`working:\` line never closes it, even when the answer is what started that work.
Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved: {how it cleared}\` yourself (same \`[key=<slug>]\` if you opened it with one) as you resume.
7. Never stop, restart, or update the shared \`no-mistakes\` daemon - it is one instance serving
every lane/home, so restarting it kills other lanes' in-flight pipeline runs. On ANY no-mistakes
daemon error, append \`blocked: {the daemon error}\` and stop; only firstmate manages the daemon.
every lane/home, so restarting it kills other lanes' in-flight pipeline runs; only firstmate
manages the daemon.
Before you append \`blocked:\` about the pipeline, run \`no-mistakes daemon status\` and
\`no-mistakes axi status\`. If the daemon socket refuses connections or is missing, append
\`blocked: {the daemon error}\` and stop even when the local run record still says running or
fixing, because that record can be stale after the daemon exits. A run record failed with a
daemon error is also a real block.
Only after ruling out socket refusal, if the run is still running or fixing, reattach and keep
going. A drive-call error, timeout, slow read, or generic unreachability is NOT a daemon error:
the daemon accepts \`respond\` immediately and runs the round in the background, so a killed or
timed-out call was only waiting for a read while the run kept working.

$INBOX_SECTION

Expand Down
83 changes: 81 additions & 2 deletions bin/fm-crew-state.sh
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,13 @@
# 3. Reconcile the status log: if its last line says needs-decision/blocked but
# the run-step shows the run moved on, the log is deterministically stale and
# is flagged superseded. A genuinely parked run plus a needs-decision log
# agree, and are reported as parked.
# agree, and are reported as parked. A `blocked:` line that reports a
# refused or missing daemon socket remains blocked even if an attributed
# run record is stale or terminal. Other daemon, timeout, or unreachability
# claims are superseded BECAUSE THE RUN IS ALIVE when the run is
# running/fixing with recent reported activity: a killed or timed-out drive
# call is not daemon death, so that claim is answered by steering the crew
# to reattach, not by escalating.
# 4. No run for this crew (pre-validation, or kind=scout): fall back to the
# recorded backend's pane busy state, then the status log's last line only
# when its verb maps to a recognized run-state. Decision-only events such as
Expand Down Expand Up @@ -315,6 +321,61 @@ log_reports_ci_ready() {
esac
}

# 0 when a status-log line reports positive daemon socket failure rather than a
# client-side timeout or generic unreachability.
log_reports_daemon_socket_down() { # <line>
local line
line=$(printf '%s' "$1" | tr '[:upper:]' '[:lower:]')
case "$line" in
*daemon*|*no-mistakes*) ;;
*) return 1 ;;
esac
case "$line" in
*"connection refused"*|*"connections refused"*|*"socket refused connection"*|*"socket refuses connection"*|*"socket refusing connection"*|*"socket missing"*|*"socket is missing"*|*"missing socket"*) return 0 ;;
esac
return 1
}

# 0 when a status-log line blames the pipeline's transport rather than the work.
# None of these claims alone is evidence the daemon died: a drive call is only
# waiting for a read while the fix round runs in the background.
log_claims_pipeline_unreachable() { # <line>
case "$(printf '%s' "$1" | tr '[:upper:]' '[:lower:]')" in
*daemon*|*timeout*|*"timed out"*|*unreachab*) return 0 ;;
esac
return 1
}

# Rows of the `active_steps[N]{...}:` table in the captured run output
# ($RUN_OUT), which the pipeline emits only while a step is actually running or
# fixing. Column order is deliberately not assumed: the header's own indentation
# bounds the block, and callers below read the table as text.
nm_active_steps_rows() {
printf '%s\n' "$RUN_OUT" | awk '
/^[[:space:]]*active_steps\[[0-9]+\]\{/ { hdr = index($0, "active_steps"); inblock = 1; next }
inblock {
if ($0 ~ /^[[:space:]]*$/) { inblock = 0; next }
match($0, /[^ \t]/)
if (RSTART <= hdr) { inblock = 0; next }
print
}
'
}

# 0 when the pipeline itself reports RECENT activity on an actively running or
# fixing step. The client prefixes a step's `last_activity` with `quiet` once no
# step log or native-agent lifecycle event has arrived for longer than its
# configured quiet warning, so its own recency verdict is the signal here rather
# than a second threshold invented in firstmate. Positive evidence is required:
# an absent table is not recency, so a run record that merely still says
# `running` while nothing executes it never reads as alive.
nm_run_activity_is_recent() {
local rows
rows=$(nm_active_steps_rows)
[ -n "$rows" ] || return 1
! printf '%s\n' "$rows" | grep -q 'quiet'
}

nm_ci_step_status() {
local row rest
row=$(printf '%s\n' "$RUN_OUT" | grep -E '^[[:space:]]*ci,[[:space:]]*"?(running|fixing)"?[[:space:]]*,' | head -1)
Expand Down Expand Up @@ -548,11 +609,29 @@ if [ "$HAVE_RUN" = 1 ]; then
# Reconcile the status log. A needs-decision/blocked log line that the run-step
# has moved past (anything but a genuinely parked run) is deterministically
# stale: the gate resolved and the run resumed or finished.
#
# A refused or missing daemon socket is positive daemon-down evidence and
# outranks any attributed run record, including a terminal one left behind
# after the daemon stopped. Other blocked claims caused by a timed-out drive
# call are contradicted only when the run reports recent
# activity; the answer is then to steer the crew to reattach without touching
# the shared daemon.
case "$LOG_VERB" in
needs-decision|blocked)
if [ "$LOG_VERB" = blocked ] \
&& log_reports_daemon_socket_down "$LOG_LINE"; then
emit blocked status-log "$(status_line_note "$LOG_LINE")${SEP}daemon socket down despite attributed run record"
fi
if [ "$RUN_STATE" != parked ]; then
if [ "$RUN_STATE" = working ]; then
RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded by active run"
if [ "$LOG_VERB" = blocked ] \
&& log_claims_pipeline_unreachable "$LOG_LINE" \
&& { [ "$RUN_STATUS" = running ] || [ "$RUN_STATUS" = fixing ]; } \
&& nm_run_activity_is_recent; then
RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded: run alive, not a daemon failure (steer reattach)"
else
RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded by active run"
fi
else
RUN_DETAIL="$RUN_DETAIL${SEP}status-log superseded (run $RUN_STATE)"
fi
Expand Down
6 changes: 6 additions & 0 deletions bin/fm-dod-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -216,6 +216,12 @@ When the captain's intent refers to a report, decision, or PR ("do items 1, 2, 3
This replaces the no-mistakes skill's advice to enrich \`--intent\` with decisions and tradeoffs; that advice does not apply to Firstmate-dispatched work.
Do not hand-edit, commit, or fix findings yourself while a run is active - the pipeline applies every fix.

One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain.
So background the drive call and poll \`no-mistakes axi status\` from a separate call instead of sitting in one blocking hold your harness will kill.
Where a harness's own command limit is not established, assume it bounds commands and use that same background-and-poll shape.
A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working.
Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real.

Two firstmate-specific rules layer on top of that guidance:
- ask-user findings are never yours to answer: escalate to firstmate using rule 6's ask-user format and stop.
Firstmate applies \`ask-user-authority\` and obtains any required captain decision.
Expand Down
3 changes: 2 additions & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,8 @@ The explicit resolution is written by the actor that answers, not the busy worke
This home's answerer close, pending-reply escalation close, and captain-held transfer use the provenance-guarded append owned by `bin/fm-wake-lib.sh`, so they advance the watcher marker only across their own bytes when all earlier bytes were already announced; pending or interleaved foreign bytes fail toward an ordinary wake.
A turn-ended-only queue row omits its historical status annotation when that status file exactly matches the same seen marker.
Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line.
`bin/fm-crew-state.sh <id>` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed.
`bin/fm-crew-state.sh <id>` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed, except that a `blocked:` event reporting a refused or missing daemon socket outranks a potentially stale active run record.
For other daemon, timeout, or unreachability claims, a running or fixing run with recent pipeline-reported activity supersedes the event and names reattachment as the recovery instead of surfacing a false block.
[`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules.
A run head the task copy cannot resolve locally is attributed only when the pipeline's own runs ledger proves it is an active continuation of the submitted head, so a pipeline fix round never reads as an older failed run.
During no-mistakes' `ci` monitor phase, it also reads the ci step log tail because `axi status` reports both "still waiting on checks" and "checks green, waiting on merge" as `ci,running`.
Expand Down
Loading
Loading