Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .agents/skills/operational-home-layout/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@ config/trace-context optional presence flag enabling default-off native W3C tra
config/lavish-axi-host optional one-line per-machine Lavish server address; LOCAL, gitignored, inherited by secondmate homes, and exported into every worker launch; see docs/configuration.md "Lavish server address" for opening versus polling
config/brief-include.md optional standing worker instructions appended verbatim as the last section of every ship and scout scaffold; LOCAL, gitignored, and not inherited; keep its text out of `## Firstmate spec`; see docs/configuration.md "Home brief include"
config/fleet-ledger optional presence flag opting this home in to the default-off fleet activity ledger state/fleet-ledger.jsonl that outside tools can follow; LOCAL, gitignored, and not inherited; see docs/fleet-ledger.md
config/wait-no-turns optional presence flag opting this home into default-off waiting-worker behavior (brief waiting section, foreground pipeline drive, pending-reply hold, one fire-and-forget retry ring); LOCAL, gitignored, and not inherited; see docs/configuration.md "Waiting worker spends no turns"
config/turnend-churn-absorb optional presence flag opting this home into the default-off absorb of bare turn-end wakes on pane churn; LOCAL, gitignored, and not inherited; see docs/configuration.md "Turn-end pane-churn absorb"
config/wedge-defer-parked-gate optional presence flag opting this home into the default-off deferral of a wedge escalation for a lane parked at a validation gate awaiting the supervisor's own still-open decision; LOCAL, gitignored, and not inherited; see docs/configuration.md "Parked-gate wait deferral"
config/cmux-socket-password optional cmux control-socket password; LOCAL, gitignored; read fresh on every cmux CLI call and passed through without ever overriding an operator's own ambient CMUX_SOCKET_PASSWORD when absent (docs/cmux-backend.md "Setup")
Expand Down
41 changes: 36 additions & 5 deletions bin/fm-brief.sh
Original file line number Diff line number Diff line change
Expand Up @@ -348,17 +348,48 @@ INBOX_DIR=$(shell_quote "$STATE/$ID.inbox")

# The receive-and-ack half of the steering-inbox contract, included in every
# scaffold kind. The record format, doorbell line, and re-ring ladder are
# owned by bin/fm-task-inbox-lib.sh; the doorbell itself is self-describing,
# so this section is reinforcement for the natural-checkpoint habit, not the
# only carrier of the instruction.
# owned by bin/fm-task-inbox-lib.sh. The doorbell names the inbox as
# "$FM_TASK_INBOX", which bin/fm-spawn.sh exports into every launch; the full
# path here remains the fallback for a worker launched without that export.
# The doorbell itself is self-describing, so this section is reinforcement
# for the natural-checkpoint habit, not the only carrier of the instruction.
# config/wait-no-turns (docs/configuration.md) adds the line that a waiting
# worker does not poll the inbox: checkpoint checks happen during active work,
# so waiting still spends no turns.
IFS= read -r -d '' INBOX_SECTION <<EOF || true
# Firstmate instruction inbox
Firstmate steers you through durable message files in $INBOX_DIR.
When a terminal message says an instruction is waiting there - and at any natural checkpoint when you are unsure - list $INBOX_DIR/*.msg, read and act on each message in numeric order, then acknowledge each handled message by moving it: \`mv $INBOX_DIR/NNN.msg $INBOX_DIR/handled/\`.
The move IS the acknowledgement: without it firstmate rings again and eventually treats you as stuck. An empty or absent inbox needs no action.
EOF
if [ -e "$CONFIG/wait-no-turns" ]; then
INBOX_SECTION+="Do not poll or list the inbox while waiting; a waiting instruction rings."$'\n'
fi
INBOX_SECTION=${INBOX_SECTION%$'\n'}

# How a crewmate or scout waits. Every model turn resends the whole context, so
# a wait must cost no turns: a decision wait ends the turn, and an external
# wait sleeps in one bounded blocking shell command sized to the harness.
# Emitted only when config/wait-no-turns is present.
IFS= read -r -d '' WAIT_SECTION <<'EOF' || true
# Waiting
Every turn you take resends your whole context, so a wait must cost no turns.
After you append `needs-decision:` or `blocked:`, end your turn at once: do not check the inbox, the status file, or anything else, because the answer arrives as a terminal message that starts your next turn.
Wait on anything external - a pipeline gate, PR checks, a heavy-test slot - with ONE blocking shell command that returns when the state changes: `no-mistakes axi run` or `respond` with `--wait`, `gh pr checks <pr> --watch`, or `until <condition>; do sleep 30; done` for anything else.
Never spend turns on `sleep` followed by a status check, and never background a command in order to poll it.
In Claude Code that `until` loop in a single Bash call is the sanctioned foreground wait: when the harness refuses a sleep-then-check command and points you at backgrounding instead, reissue the wait as the loop rather than accepting the background.
Bound that command by what your harness lets one command run: in Pi pass the bash tool a `timeout` of at most 2700 seconds, because Pi sets none by default; in Claude Code pass the Bash tool its maximum `timeout` of 600000 ms, because its default is 2 minutes; in Codex keep waiting on a still-running command with empty `write_stdin` polls of up to 300000 ms; elsewhere pass your shell tool its largest timeout and assume at most 10 minutes.
Give any `--wait` a duration a little under that bound.
When the bound passes with nothing changed, run the same blocking command again, with no status check in between.
The one exception is `respond`: it sent its answer before it began waiting, so reattach with `no-mistakes axi run --wait` instead, and never send the same `respond` again, because it would answer whichever gate parks next without you reading it.
A wait your shell can watch this way needs no `paused:` line, except your own pipeline run, a long foreground command, or your own validation round, which you declare once just before its blocking hold: append `paused:` once just before its first blocking command, then stay in the command, and never append it again as you reissue that command.
EOF
WAIT_SECTION=${WAIT_SECTION%$'\n'}
WAIT_BLOCK=
if [ -e "$CONFIG/wait-no-turns" ]; then
WAIT_BLOCK="$WAIT_SECTION"$'\n\n'
fi

if [ "$KIND" = secondmate ]; then
SECONDMATE_PROJECTS=""
idx=1
Expand Down Expand Up @@ -576,7 +607,7 @@ $CREWMATE_PAUSE_INSTRUCTIONS
Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved [at=<epoch>]: {how it cleared}\` yourself (same \`[key=<slug>]\` if you opened it with one) as you resume.
$SHARED_INFRA_RULE

$INBOX_SECTION
$WAIT_BLOCK$INBOX_SECTION

# Definition of done
Write your findings to \`$DATA/$ID/report.md\`.
Expand Down Expand Up @@ -654,7 +685,7 @@ $ASK_USER_BLOCK
Firstmate's reply normally writes that closing line at answer time; when a blocker or wait clears WITHOUT a firstmate reply, append \`resolved [at=<epoch>]: {how it cleared}\` yourself (same \`[key=<slug>]\` if you opened it with one) as you resume.
$SHARED_INFRA_RULE

$INBOX_SECTION
$WAIT_BLOCK$INBOX_SECTION

# Project memory
A project's \`AGENTS.md\` or \`CLAUDE.md\` is loaded into every agent session in that project, so edit it only to correct information that is factually wrong - including information your own change made wrong - and never to add knowledge because it is missing.
Expand Down
19 changes: 19 additions & 0 deletions bin/fm-classify-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -922,6 +922,25 @@ EOF
printf '%s\n' "$current"
}

# The subset of status_open_decisions the task raised about its own work: a
# reserved-namespace key is raised by a supervisor library about the task (a
# pending-reply escalation), a `remote-reply-continuity-` key is the parent's
# own blocker about a broken remote reply mirror
# (bin/fm-procevent-remote-reply.sh), and a `captain-hold-` key relays a child
# decision a secondmate escalated to the captain (bin/fm-captain-hold.sh) while
# it keeps working, so the task is not waiting on any of them. Pending-reply
# recovery and a fire-and-forget retry ring consult this set and leave a task
# alone while it is non-empty.
status_own_open_decisions() { # <status-file>
local line prefix
status_open_decisions "$1" | while IFS= read -r line || [ -n "$line" ]; do
for prefix in ${FM_CLASSIFY_RESERVED_KEY_PREFIXES:-$FM_CLASSIFY_RESERVED_KEY_PREFIXES_DEFAULT} remote-reply-continuity- captain-hold-; do
case "$line" in "$prefix"*) continue 2 ;; esac
done
printf '%s\n' "$line"
done
}

# 0 when the fold above still holds at least one decision OPENED by
# `needs-decision` - the status side's own record that a human was asked
# something and has not answered. A `blocked` record is deliberately not this: a
Expand Down
24 changes: 18 additions & 6 deletions bin/fm-dod-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -280,12 +280,28 @@ EOF
# Written once; only the two sentences about a green PR depend on the forge,
# because on gerrit the ci step is skipped and there is no PR to report.
fm_nm_driving_block() { # <forge>
local pr_return_line='' pr_reattach_clause=';'
local pr_return_line='' pr_reattach_clause=';' drive_block wait_cfg
if [ "$1" != gerrit ]; then
pr_return_line="Only a drive call's return reports the green PR: \`no-mistakes axi status\` shows progress but never reports \`checks-passed\` while the ci step is still monitoring the PR for merge, so never wait on a status poll for the next gate or outcome.
"
pr_reattach_clause="; once checks are green it returns \`checks-passed\` immediately, and"
fi
# config/wait-no-turns selects the foreground drive. Absent, the text matches
# the backgrounded drive a home had before that flag.
wait_cfg=${CONFIG:-${FM_CONFIG_OVERRIDE:-${FM_HOME:-}/config}}
if [ -e "$wait_cfg/wait-no-turns" ]; then
drive_block="Drive the run with ONE foreground \`no-mistakes axi run\` and let it block.
It bounds its own hold for you: \`--wait\` (default 8m) exists precisely so a harness with a ten-minute command cap gets a structured return instead of being killed mid-hold.
Declare that wait using the brief's status-reporting rule before the foreground drive call.
Never background a wait, and never arm a timer to stand in for one: a backgrounded call returns in milliseconds, so it does not wait at all, and every timer left behind fires later as a paid wake for nothing.
${pr_return_line}Whenever a drive call returns without a gate or an outcome - its own wait elapsed, or it was killed or timed out - that is not a failure: reattach at once by re-running \`no-mistakes axi run\` without flags, and issue the same foreground call again, one at a time, until a gate or outcome comes back${pr_reattach_clause} if it refuses because no run is active, read the finished outcome from \`no-mistakes axi status\`."
else
drive_block="One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain.
So background the drive call instead of sitting in one blocking hold your harness will kill, and read its return when it finishes.
Declare that wait using the brief's status-reporting rule before waiting on the backgrounded drive call.
Where a harness's own command limit is not established, assume it bounds commands and use that same backgrounded shape.
${pr_return_line}Whenever a drive call returns without a gate or an outcome - its own wait elapsed, or it was killed or timed out - reattach at once by re-running \`no-mistakes axi run\` without flags, backgrounded the same way${pr_reattach_clause} if it refuses because no run is active, read the finished outcome from \`no-mistakes axi status\`."
fi
cat <<EOF
You drive no-mistakes by responding to its gates, not by implementing fixes.
Follow the guidance no-mistakes itself provides for the mechanics: it loads when you invoke /no-mistakes, and \`no-mistakes axi run --help\` plus the \`help\` lines in each \`axi\` response are authoritative and version-matched to the installed binary.
Expand All @@ -299,11 +315,7 @@ When the captain's intent refers to a report, decision, or PR ("do items 1, 2, 3
This replaces the no-mistakes skill's advice to enrich \`--intent\` with decisions and tradeoffs; that advice does not apply to Firstmate-dispatched work.
Do not hand-edit, commit, or fix findings yourself while a run is active - the pipeline applies every fix.
One drive call blocks until the next gate or outcome, which routinely outlives what your harness lets a single command run: Claude Code kills a command at ten minutes maximum, while one fix round is capped around thirty minutes and up to three rounds chain.
So background the drive call instead of sitting in one blocking hold your harness will kill, and read its return when it finishes.
Declare that wait using the brief's status-reporting rule before waiting on the backgrounded drive call.
Where a harness's own command limit is not established, assume it bounds commands and use that same backgrounded shape.
${pr_return_line}Whenever a drive call returns without a gate or an outcome - its own wait elapsed, or it was killed or timed out - reattach at once by re-running \`no-mistakes axi run\` without flags, backgrounded the same way${pr_reattach_clause} if it refuses because no run is active, read the finished outcome from \`no-mistakes axi status\`.
$drive_block
A killed or timed-out call is never evidence the daemon died: the daemon accepts your response immediately and runs the round in the background, so the call was only ever waiting for a read while the run kept working.
Reattach and keep going rather than reporting the pipeline blocked; rule 7 owns the checks that decide when a pipeline block is real.
Expand Down
16 changes: 12 additions & 4 deletions bin/fm-git-strip-ai-trailers.sh
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,9 @@
# because git -c core.hooksPath=<this dir> (or a child process that
# inherits it) carries the override there, and a lookup that honored it
# would find this directory again and never run the repository's own
# hook - a skipped pre-push guard. A lookup that fails exits nonzero
# rather than skipping the repository's hook. Does not touch the
# hook - a skipped pre-push guard. An empty core.hooksPath means no
# repository hook, as in plain git; any other failed lookup exits
# nonzero rather than skipping the repository's hook. Does not touch the
# project's git config; the caller prefixes the pane with
# GIT_CONFIG_COUNT / GIT_CONFIG_KEY_0 / GIT_CONFIG_VALUE_0.
#
Expand Down Expand Up @@ -153,14 +154,21 @@ write_executable() {
# is the other environment channel that can carry this directory as
# core.hooksPath; only the repository's config files name its own hooks. Skip
# when the lookup still names this launch's own hooks dir, meaning those files
# point here, so the wrapper cannot recurse into itself.
# point here, so the wrapper cannot recurse into itself. An empty
# core.hooksPath makes that lookup fail, but plain git reads it as "no hooks",
# so the wrapper runs none; any other failure reruns the lookup to show git's
# error and refuses.
runtime_chain_body() {
local ours=$1
cat <<EOF
unset GIT_CONFIG_COUNT GIT_CONFIG_KEY_0 GIT_CONFIG_VALUE_0
ours=$(quote_for_hook "$ours")
name=\${0##*/}
orig=\$(unset GIT_CONFIG_PARAMETERS; git rev-parse --path-format=absolute --git-path hooks) || {
orig=\$(unset GIT_CONFIG_PARAMETERS; git rev-parse --path-format=absolute --git-path hooks 2>/dev/null) || {
if hooks_path=\$(unset GIT_CONFIG_PARAMETERS; git config --get --type=path core.hooksPath 2>/dev/null) && [ -z "\$hooks_path" ]; then
exit 0
fi
(unset GIT_CONFIG_PARAMETERS; git rev-parse --path-format=absolute --git-path hooks >/dev/null)
echo "fm-git-strip-ai-trailers: cannot resolve this repository's hooks directory; refusing to skip its \$name hook" >&2
exit 1
}
Expand Down
11 changes: 9 additions & 2 deletions bin/fm-pending-reply-lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,10 @@
# Safety property (captain direction 2026-07-22): a secondmate agent may ignore
# the marker and answer only in its visible conversation. The parent must notice
# the missing correlated report without scraping that conversation, send exactly
# one automatic recovery request asking for a repost through the parent channel,
# and escalate once if the recovery turn also completes without a correlated
# one automatic recovery request asking for a repost through the parent channel
# (held back, when config/wait-no-turns is present, while the mate waits on its
# own open decision or blocker), and
# escalate once if the recovery turn also completes without a correlated
# report. Never loop, never repeatedly inject, never silently expire unresolved
# records, and never treat wrong-home or structured-home heuristics as
# acknowledgement. A same-basename restatement-copy of the mate home's
Expand Down Expand Up @@ -963,6 +965,11 @@ fm_pending_reply_send_recovery() { # <state-dir> <corr_id>
task_id=$(fm_pending_reply_get "$rec" task_id)
# A remote mate's report may exist and simply not have been mirrored yet.
fm_pending_reply_missing_report_is_evidence "$state" "$task_id" "$completed" || return 1
# config/wait-no-turns: a mate waiting on its own open decision or blocker
# is never poked. The recovery stays unattempted until the answer lands.
if [ -e "${FM_CONFIG_OVERRIDE:-${FM_HOME:-}/config}/wait-no-turns" ]; then
[ -z "$(status_own_open_decisions "$state/$task_id.status")" ] || return 1
fi
status_file=$(fm_pending_reply_get "$rec" parent_status)
parent_home=$(fm_pending_reply_get "$rec" parent_home)
msg=$(fm_pending_reply_recovery_message "$rec")
Expand Down
21 changes: 18 additions & 3 deletions bin/fm-send.sh
Original file line number Diff line number Diff line change
Expand Up @@ -51,7 +51,9 @@
# watcher re-rings an unacknowledged message while its endpoint remains
# available, escalates after the bounded ladder, and instead routes a positively
# dead or missing endpoint directly to recovery without typing. An explicit
# fire-and-forget record is excluded from that ladder.
# fire-and-forget record is excluded from that ladder; when config/wait-no-turns
# is present and its ring here was skipped or failed, the watcher rings it
# exactly once more.
# bin/fm-task-inbox-lib.sh owns the record format, the doorbell line, and the
# re-ring ladder. The composer pre-check before the ring is ADVISORY only: when
# the composer visibly holds pending text the ring is skipped with a notice and
Expand Down Expand Up @@ -1085,9 +1087,22 @@ else
# bounded re-ring ladder or direct unavailable-endpoint recovery.
ring_rc=0
fm_task_inbox_ring "$TARGET_BACKEND" "$T" "$INBOX_RECORD" "$EXPECTED_LABEL" || ring_rc=$?
ring_retry="the watcher will re-ring"
if [ -n "$FIRE_AND_FORGET_ID" ] \
&& [ -e "${FM_CONFIG_OVERRIDE:-$FM_HOME/config}/wait-no-turns" ]; then
case "$ring_rc" in
1|2)
if fm_task_inbox_mark_retry "$STATE" "$INBOX_TASK_ID" "$INBOX_RECORD"; then
ring_retry="the watcher will ring it once more"
else
ring_retry="its one retry ring could not be recorded, so nothing will ring it again"
fi
;;
esac
fi
case "$ring_rc" in
1) echo "fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;;
2) echo "fm-send: doorbell did not reach $T; the steer is durably recorded at $INBOX_RECORD and the watcher will re-ring" >&2 ;;
1) echo "fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at $INBOX_RECORD and $ring_retry" >&2 ;;
2) echo "fm-send: doorbell did not reach $T; the steer is durably recorded at $INBOX_RECORD and $ring_retry" >&2 ;;
3) echo "fm-send: doorbell not typed because the agent in $T has exited; the steer is durably recorded at $INBOX_RECORD for recovery (stuck-crewmate-recovery), and the watcher will not re-ring a dead pane" >&2 ;;
esac
exit 0
Expand Down
Loading
Loading