Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .agents/skills/bearings/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ It never tears down a task, merges a PR, dispatches new work, or mutates any tas
When the captain asks to include PRs, use the command's live-PR opt-in; otherwise keep the default local-only read.
If the command is unavailable, fall back to `bin/fm-fleet-snapshot.sh --json` and `bin/fm-crew-state.sh <id>`; never infer current state from a raw `tail` of `state/<id>.status`, which is append-only wake-event history whose last line goes stale.
For registered secondmates, use the snapshot's structured-home classification and provenance; a parent event or bounded terminal contradiction is fallback evidence, never authority over readable structured home state.
Structured captain-held decisions from `decision-hold-lifecycle` (#593) are not yet ported on this fork; `decisions_open` surfaces fork-native parked `needs-decision` / `blocked` status hints instead, and does not scrape reports or visual-review artifacts to invent decisions.
`decisions_open` surfaces structured captain holds from `decision-hold-lifecycle` plus parked `needs-decision` / `blocked` status hints; it does not scrape reports or visual-review artifacts to invent decisions.
A queued item under `gates` only becomes "next work" when its blocker is gone and its time/date gate has arrived; until then it stays queued with the reason.
The `(main-inventory)` gate is an action-free integrity warning rather than queued work: render it under Charted Next with the related `omitted` disclosure, never invent an Underway row from backlog-only state, and never move it into Captain's Call.

Expand Down
40 changes: 40 additions & 0 deletions .agents/skills/decision-hold-lifecycle/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
---
name: decision-hold-lifecycle
description: >-
Agent-only policy for completing investigations and visual reviews without losing unresolved captain decisions.
Load before treating an investigation, scout report, structured review, or Lavish review as complete, before ending a visual review that exposed a decision, and when recording or routing the captain's answer.
user-invocable: false
metadata:
internal: true
---

# Durable unresolved-decision lifecycle

This skill is the single policy owner for unresolved captain decisions discovered by an investigation or visual review.

## Policy

Every unresolved decision that belongs to the captain and is discovered while producing, reading, presenting, or ending an investigation or visual review must become a structured captain-held work item in the authoritative backlog of the home that owns the originating work before that work or review may be treated as complete.
The agent performs the semantic inventory because scripts must not infer decisions from report prose, visual-review artifacts, terminal output, or chat.
Give each distinct unresolved decision a stable privacy-safe key, register it through `bin/fm-decision-hold.sh hold`, and use the same key on retry so registration is idempotent while different decisions retain different durable identities.
After inventorying the whole report and review surface, run `bin/fm-decision-hold.sh complete` with every unresolved key, or with `--none` only when the reviewed surface contains no unresolved captain decision.
A completed investigation and an ended visual review use this same owner and completion command; a visual tool, including Lavish, never owns a parallel completion policy.
Run the command in the originating work's authoritative `FM_HOME`; main-home work creates main-home holds, and secondmate-owned work creates holds in that secondmate home's backlog rather than copying them into the main backlog.
Do not close a hold merely because the originating investigation completed, its report was archived, its visual review ended, or its task was torn down.
The hold remains the authoritative Captain's Call item until the captain's answer is durably recorded, dependent work is created in the same backlog and blocked by that hold, and `bin/fm-decision-hold.sh resolve` routes the answer by clearing those dependency edges before closing the hold.
Resolved findings, recommendations that need no captain choice, and prose that merely sounds decision-like do not create holds.
Bearings reads the resulting structured state and must never compensate by scraping historical reports, visual-review artifacts, terminal output, chat, or other prose.

## Operating sequence

1. Read the complete investigation result and complete the visual review before declaring either complete.
2. Inventory only genuine unresolved choices that require the captain.
3. For each choice, choose a stable key and use the script's `hold` command with a concise title, reason, and repository.
4. Run the script's `complete` command with the full unresolved-key inventory for that review pass.
5. Relay the choices to the captain as decisions from Bearings' Captain's Call section under `AGENTS.md` section 9; do not use the word hold in captain chat.
6. After the captain decides, record dependent work with normal tasks-axi commands and block it by the hold identity.
7. Put the captain's exact durable decision in a file and use the script's `resolve` command with every routed task.
8. Confirm Bearings no longer shows the closed hold and that routed work remains in structured backlog state.

`bin/fm-decision-hold.sh --help` owns command syntax, identity construction, completion attestation, retry behavior, and close ordering.
`docs/decision-hold-lifecycle.md` records the mechanism and regression evidence without restating this policy.
2 changes: 1 addition & 1 deletion .agents/skills/firstmate-orca/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ For a messy Orca-backed task:
6. Stop and inspect if the recorded worktree path, Orca worktree id, or project checkout no longer matches expectations.

Teardown remains governed by the normal firstmate landing rules.
Scout work can be torn down after the report exists.
Scout work can be torn down after the report exists and the `decision-hold-lifecycle` completion gate passes.
Ship work can be torn down only after the work is landed by its project mode.

## Smoke Test
Expand Down
14 changes: 13 additions & 1 deletion .agents/skills/harness-adapters/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,19 @@ Every verified primary harness also has a wired PreToolUse-equivalent hook that
`claude` and `codex` block directly through PreToolUse hooks; `grok` blocks the same way but requires every `$VAR` reference in its hook `command` string to carry an inline `:-default` or it fails to launch the hook entirely.
`opencode` and `pi` block by throwing from `tool.execute.before` / returning `{block: true}` from `tool_call`.
The exact hook files, commands, output-shaping quirks (Claude Code only honors the deny when stdout is empty), and validation transcripts are owned by `docs/arm-pretool-check.md`.
When changing any primary PreToolUse hook, validate the real harness behavior in a scratch project before trusting it, then update that doc.
When changing any watcher-arm PreToolUse hook, validate the real harness behavior in a scratch project before trusting it, then update that doc.

## Primary delegation-shape guard

Claude exposes built-in delegation, scheduling, and worktree tools that a primary session can use to create work with no `state/<id>.meta`, which makes the whole guard stack inert because every guard counts that metadata.
The shipped mechanism is `bin/fm-subagent-pretool-check.sh`, a primary-home PreToolUse guard that denies a delegation-SHAPED tool name.
Claude primaries should also use an untracked per-home local `permissions.deny` list as hardening for known Claude delegation tools, because it removes them from the model's schema so they are never offered.
That deny list must not ship in tracked `.claude/settings.json` because it is Claude-only rather than harness-agnostic, and because tracked project settings propagate into linked worktrees where they disarm legitimate crewmates.
`docs/subagent-guard.md` owns the full contract, the local deny-list recommendation, the `FM_ALLOW_SUBAGENT=1` escape hatch, and the per-harness applicability review.

Two verified facts worth pinning here.
The subagent tool presents to the model as `Agent`, and on Claude Code 2.1.217 both `Agent` and `Task` work as `permissions.deny` keys, verified by an A/B with a nonsense-name control.
`permissions.allow` is a pre-approval list rather than an availability list, so there is no fail-closed positive allowlist.

## Crew process-signal guard

Expand Down
13 changes: 13 additions & 0 deletions .claude/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,19 @@
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/bin/fm-arm-pretool-check.sh --claude"
},
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/bin/fm-cd-pretool-check.sh --claude"
}
]
},
{
"matcher": ".*",
"hooks": [
{
"type": "command",
"command": "\"$CLAUDE_PROJECT_DIR\"/bin/fm-subagent-pretool-check.sh --claude"
}
]
}
Expand Down
5 changes: 5 additions & 0 deletions .codex/hooks.json
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,11 @@
"type": "command",
"command": "bash -lc 'payload=$(cat 2>/dev/null || true); [ -n \"$payload\" ] || exit 0; command -v jq >/dev/null 2>&1 || exit 0; root=$(pwd -P) || exit 0; [ -x \"$root/bin/fm-arm-pretool-check.sh\" ] || exit 0; [ -f \"$root/AGENTS.md\" ] || exit 0; [ -f \"$root/.codex/hooks.json\" ] || exit 0; jq -e \"any(.hooks.PreToolUse[]?.hooks[]?.command?; type == \\\"string\\\" and contains(\\\"fm-arm-pretool-check.sh\\\"))\" \"$root/.codex/hooks.json\" >/dev/null 2>&1 || exit 0; printf \"%s\" \"$payload\" | \"$root/bin/fm-arm-pretool-check.sh\"'",
"timeout": 10
},
{
"type": "command",
"timeout": 10,
"command": "bash -lc 'payload=$(cat 2>/dev/null || true); [ -n \"$payload\" ] || exit 0; command -v jq >/dev/null 2>&1 || exit 0; root=$(pwd -P) || exit 0; [ -x \"$root/bin/fm-cd-pretool-check.sh\" ] || exit 0; [ -f \"$root/AGENTS.md\" ] || exit 0; [ -f \"$root/.codex/hooks.json\" ] || exit 0; jq -e \"any(.hooks.PreToolUse[]?.hooks[]?.command?; type == \\\"string\\\" and contains(\\\"fm-cd-pretool-check.sh\\\"))\" \"$root/.codex/hooks.json\" >/dev/null 2>&1 || exit 0; printf \"%s\" \"$payload\" | \"$root/bin/fm-cd-pretool-check.sh\"'"
}
]
}
Expand Down
16 changes: 16 additions & 0 deletions .grok/hooks/fm-primary-cd-check.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "bash -lc '[ -n \"${GROK_WORKSPACE_ROOT:-}\" ] || exit 0; exec \"${GROK_WORKSPACE_ROOT:-}/bin/fm-cd-pretool-check.sh\"'",
"timeout": 10
}
]
}
]
}
}
9 changes: 9 additions & 0 deletions .no-mistakes.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,14 @@
# Per-repo no-mistakes overrides.

# firstmate is an agent-orchestration repo: its AGENTS.md installs a fleet-captain
# identity. Disable project-level agent settings/instructions for gate agents so a
# no-mistakes review/test/document/lint/pr/rebase/ci agent never adopts that
# identity or drives the fleet. Trusted-only: a pushed branch cannot turn this off,
# so it is honored only from the default-branch copy of this file. Layered above
# the NO_MISTAKES_GATE lifecycle refusal (bin/fm-gate-refuse-lib.sh) and the
# HEAD-continuity guard; see docs/architecture.md "No-mistakes gate authority boundary."
disable_project_settings: true

# Run the firstmate bash behavior suite deterministically as the test-step
# baseline, instead of delegating to an agent (an agent-driven test step has
# crashed the daemon). Mirrors .github/workflows/ci.yml: iterate every
Expand Down
64 changes: 64 additions & 0 deletions .opencode/plugins/fm-primary-cd-check.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
import { realpathSync } from "node:fs";
import { resolve } from "node:path";
import { spawn } from "node:child_process";

// PreToolUse seatbelt for OpenCode: block a stray persistent top-level `cd` in
// the primary firstmate checkout before the agent's bash tool relocates the
// shell out of the home (see bin/fm-cd-pretool-check.sh and docs/cd-guard.md).
// This mirrors fm-primary-pretool-check.js, calling the cd-guard owner instead
// of the watcher-arm one. tool.execute.before can block by throwing (verified
// 2026-07-09 against OpenCode 1.17.15 for the watcher-arm plugin; the same
// mechanism carries this guard). The owner script is itself inert outside the
// real primary checkout, so a crewmate/scout worktree is never affected.

function runProcess(command, args) {
return new Promise((resolvePromise) => {
const child = spawn(command, args, { stdio: ["ignore", "pipe", "pipe"] });
let stdout = "";
let stderr = "";
child.stdout.on("data", (chunk) => {
stdout += chunk.toString();
});
child.stderr.on("data", (chunk) => {
stderr += chunk.toString();
});
child.on("error", () => resolvePromise({ code: 0, stdout: "", stderr: "" }));
child.on("close", (code) => resolvePromise({ code: code ?? 0, stdout, stderr }));
});
}

async function resolveRoot(anchor) {
if (!anchor) return "";
const result = await runProcess("git", ["-C", anchor, "rev-parse", "--show-toplevel"]);
const root = result.stdout.trim();
if (result.code === 0 && root) return root;
try {
return realpathSync(anchor);
} catch {
return resolve(anchor);
}
}

export const FmPrimaryCdCheck = async ({ directory, worktree }) => {
const root = worktree ? (() => {
try {
return realpathSync(worktree);
} catch {
return resolve(worktree);
}
})() : await resolveRoot(directory);

return {
"tool.execute.before": async (input, output) => {
if (!root || input?.tool !== "bash") return;
const command = output?.args?.command;
if (!command || typeof command !== "string") return;

const result = await runProcess(`${root}/bin/fm-cd-pretool-check.sh`, ["--command", command]);
if (result.code !== 2) return;

const reason = result.stderr.trim() || "denied by the cd-guard PreToolUse seatbelt";
throw new Error(reason);
},
};
};
1 change: 1 addition & 0 deletions .opencode/plugins/fm-primary-turnend-guard.js
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ function runProcess(command, args, input = "") {
child.stderr.on("data", (chunk) => {
stderr += chunk.toString();
});
child.stdin.on("error", () => {});
child.on("error", () => resolve({ code: 0, stdout: "", stderr: "" }));
child.on("close", (code) => resolve({ code: code ?? 0, stdout, stderr }));
child.stdin.end(input);
Expand Down
29 changes: 21 additions & 8 deletions .pi/extensions/fm-primary-turnend-guard.ts
Original file line number Diff line number Diff line change
Expand Up @@ -70,15 +70,16 @@ function runGuard(): Promise<{ code: number; stderr: string }> {
});
}

// PreToolUse seatbelt (bin/fm-arm-pretool-check.sh; docs/arm-pretool-check.md).
// Piggybacks on this same extension file rather than a separate one so no
// second Pi -e flag is needed at launch - the primary already loads this
// file for the turn-end guard, and pi.on("tool_call", ...) can block
// (verified 2026-07-09 against pi 0.80.5: returning {block: true} prevents
// the bash command from running).
function runPretoolCheck(command: string): Promise<{ code: number; stderr: string }> {
// PreToolUse seatbelts (bin/fm-arm-pretool-check.sh, docs/arm-pretool-check.md;
// bin/fm-cd-pretool-check.sh, docs/cd-guard.md). Both piggyback on this same
// extension file rather than separate ones so no extra Pi -e flag is needed at
// launch - the primary already loads this file for the turn-end guard, and
// pi.on("tool_call", ...) can block (verified 2026-07-09 against pi 0.80.5:
// returning {block: true} prevents the bash command from running). Each owner
// script owns its own decision and is inert outside the real primary checkout.
function runChecker(script: string, command: string): Promise<{ code: number; stderr: string }> {
return new Promise((resolveResult) => {
const child = spawn(`${root}/bin/fm-arm-pretool-check.sh`, ["--command", command], {
const child = spawn(`${root}/bin/${script}`, ["--command", command], {
stdio: ["ignore", "ignore", "pipe"],
});
let stderr = "";
Expand All @@ -90,6 +91,14 @@ function runPretoolCheck(command: string): Promise<{ code: number; stderr: strin
});
}

function runPretoolCheck(command: string): Promise<{ code: number; stderr: string }> {
return runChecker("fm-arm-pretool-check.sh", command);
}

function runCdCheck(command: string): Promise<{ code: number; stderr: string }> {
return runChecker("fm-cd-pretool-check.sh", command);
}

export default function (pi: ExtensionAPI) {
pi.on?.("session_start", () => {
markLoaded();
Expand All @@ -99,6 +108,10 @@ export default function (pi: ExtensionAPI) {
if (event.type !== "tool_call" || event.toolName !== "bash") return {};
const command = String((event.input as { command?: unknown })?.command ?? "");
if (!command) return {};
const cdResult = await runCdCheck(command);
if (cdResult.code === 2) {
return { block: true, reason: cdResult.stderr.trim() || "denied by the cd-guard PreToolUse seatbelt" };
}
const result = await runPretoolCheck(command);
if (result.code !== 2) return {};
return { block: true, reason: result.stderr.trim() || "denied by the watcher-arm PreToolUse seatbelt" };
Expand Down
Loading