Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .agents/skills/harness-adapters/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ This preserves launch success instead of passing a known-bad value.
| Busy-pane signature | `esc to interrupt` |
| Exit command | `/exit` |
| Interrupt | single Escape |
| Skill invocation | `/<skill>` (e.g. `/code-review`) |
| Skill invocation | `/<skill>` (e.g. `/security-review`). NOT `/code-review` or `/verify`: those two Claude Code built-ins are `disable-model-invocation`, so a crewmate can never self-invoke them (verified by execution, 2.1.220). Whether typing one into a crew's composer with `fm-send` counts as the human-typed turn the gate wants is UNVERIFIED - do not rely on it. This is why `bin/fm-brief.sh` names no review command. |

First launch in a fresh worktree, or first ever on a machine, may show a trust or bypass-permissions confirmation.
After every spawn, peek the pane within about 20 seconds.
Expand Down Expand Up @@ -204,7 +204,7 @@ Launch with a positional prompt: `grok --always-approve "$(cat <brief>)"`.
| Busy-pane signature | `Ctrl+c:cancel` (the mid-turn cancel hint in grok's keybind bar, shown iff a turn is running; the spinner line is a braille glyph + `<status>… N.Ns` + `[stop]`, e.g. `⠹ Thinking… 1.1s … [stop]`). Idle keybind bar shows only `Shift+Tab:mode │ Ctrl+.:shortcuts`. The ASCII `Ctrl+c:cancel` is the busy regex (avoids locale fragility of matching braille). |
| Exit command | `Ctrl+Q` double-press within 1000ms (it is a confirmed destructive action). Prints `Resume this session with: grok --resume <session-id>`. `Ctrl+D` is the quit key in VS Code family terminals. NOT `/exit` and NOT `Ctrl+C`. |
| Interrupt | single `Ctrl+C` (cancels the current turn; the footer shows `Ctrl+c:cancel` mid-turn). `Esc` only moves focus to the scrollback, it does NOT interrupt. |
| Skill invocation | `/<skill>` (e.g. `/code-review`), same as claude; verified end to end that grok discovers a user-level skill and invokes it. Opens a slash-autocomplete popup, so a too-fast Enter selects the popup entry instead of sending. For an argument-taking command that first Enter does not submit at all - it expands the selection into an argument-hint placeholder in the composer (e.g. `/compact` -> `/compact compaction instructions`, live-verified), leaving real text still sitting there unsubmitted; a genuine second Enter is required. `fm-send`'s retried Enter lands it, because `fm_tmux_composer_state` correctly recognizes that placeholder-filled text as still-pending. |
| Skill invocation | `/<skill>` (e.g. `/security-review`), same as claude, and subject to the same caveat about Claude Code's gated built-ins; verified end to end that grok discovers a user-level skill and invokes it. Opens a slash-autocomplete popup, so a too-fast Enter selects the popup entry instead of sending. For an argument-taking command that first Enter does not submit at all - it expands the selection into an argument-hint placeholder in the composer (e.g. `/compact` -> `/compact compaction instructions`, live-verified), leaving real text still sitting there unsubmitted; a genuine second Enter is required. `fm-send`'s retried Enter lands it, because `fm_tmux_composer_state` correctly recognizes that placeholder-filled text as still-pending. |
| Autonomy | `--always-approve` (footer shows `· always-approve`); auto-approves every tool execution, verified to run fully unattended. `--permission-mode bypassPermissions` is the stronger equivalent. |
| Env marker | `GROK_AGENT=1`, set for child/tool processes. grok does NOT set `CLAUDECODE` despite Claude compatibility, so the marker is unambiguous. |
| Resume | `grok --resume <session-id>` (id printed on exit) or `grok -c` / `--continue` (most recent for the cwd); `--fork-session` branches a new session id. |
Expand Down
9 changes: 6 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,8 +145,9 @@ After any merge you perform without asking, post a one-line "merged <full PR URL
Crewmates run it inside their worktree and commit the result; firstmate never installs hooks into a clone itself.
A blocked commit or push means the floor did its job - never steer a crewmate around a hook.

**The judgment layer** is review: the crewmate's own `/code-review` and `/verify`, then firstmate's independent, direction-aware review of the pushed branch (section 6).
**The judgment layer** is review: the crewmate's own independent review of its diff and its end-to-end exercise of the change, then firstmate's independent, direction-aware review of the pushed branch (section 6).
The crew does not mark its own homework.
The scaffold names no command for either step and makes the crew name the mechanism it used instead; `bin/fm-brief.sh`'s header owns why.

**Tending** is keeping quality from decaying over time, a bounded slice at a time.
When the user invokes `/code-shape`, load that skill: it asks the review ledger (`bin/fm-review-ledger.sh`) for the next codebase slice - preferring recently-changed code, else never-reviewed code, never re-reviewing an unchanged slice - and ships one direct-fix crew scoped to it for duplication, dead code, testability, and comment hygiene.
Expand Down Expand Up @@ -237,8 +238,10 @@ Add ship and scout tasks to `data/backlog.md` under In flight; a secondmate spaw

### Review and ship

A ship crewmate pushes its branch and reports `review-ready:` (mode `PR`) or `done: ready in branch fm/<id>` (`local-only`).
**Your review is independent of the crew's own `/code-review`, `/verify`, and hook run, not a rubber stamp for them.**
A ship crewmate pushes its branch and reports `review-ready:` (mode `PR`) or `done: ready in branch fm/<id>` (`local-only`), each carrying `reviewed by: <mechanism> - <what it found>`.
**Your review is independent of the crew's own review, its end-to-end exercise, and the hook run, not a rubber stamp for them.**
A report missing that clause means the crew never got a review: send it back rather than reviewing on top of a gate that did not run.
Judge the outcome half against the diff you are about to read anyway - "no findings" on a diff with obvious defects is how an un-run review gives itself away.

1. Read the diff with `bin/fm-review-diff.sh <id>` (summary first, then `--full` or `--files <path>`) - never a raw `git diff`, which can be stale.
2. Review it against the project's direction (section 5, gate 4). Mechanical quality is the hooks' and CI's job; you are looking for drift, wrong-shaped solutions, and scope creep.
Expand Down
3 changes: 2 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@
Thanks for wanting to contribute.

This fork ships changes through an ordinary pull request, gated by CI and by review.
There is no separate validation pipeline to install: the quality gate is Claude Code hooks, the built-in `/code-review` skill, and the CI workflow in `.github/workflows/ci.yml`.
There is no separate validation pipeline to install: the quality gate is Claude Code hooks, an independent review of the diff, and the CI workflow in `.github/workflows/ci.yml`.
An agent contributor cannot invoke Claude Code's built-in `/code-review` or `/verify` and must obtain that review another way, such as a review subagent over the diff; `bin/fm-brief.sh`'s header owns why.

## Workflow

Expand Down
62 changes: 48 additions & 14 deletions bin/fm-brief.sh
Original file line number Diff line number Diff line change
Expand Up @@ -27,16 +27,33 @@
# with needs-decision rather than quietly working against it.
# For ship tasks, the definition of done is shaped by the project's delivery mode
# (data/projects.md via fm-project-mode.sh; see AGENTS.md task lifecycle):
# PR implement -> /code-review + /verify -> push the branch, open NO PR ->
# report `review-ready: branch fm/<id> pushed, no PR` and STOP (default).
# PR implement -> independent review + end-to-end exercise -> push the branch,
# open NO PR -> report
# `review-ready: branch fm/<id> pushed, no PR - reviewed by: <mechanism> -
# <what it found>` and STOP (default).
# Firstmate reviews the pushed branch against the direction; findings are
# fixed in place and re-signalled `review-ready:`, and only an approval opens
# the PR (plain gh) with `done: PR <url>`. The review gate sits BEFORE the PR so
# a finding costs one fix, not a re-review plus a re-run of the crew's verify and
# full suite and the PR's CI - the largest source of rework measured in the fleet.
# a finding costs one fix, not a re-review plus a re-run of the crew's end-to-end
# exercise and full suite and the PR's CI - the largest source of rework measured
# in the fleet.
# The push still happens, so the work is durable against a box reboot.
# local-only implement on branch, stop and report "ready in branch" (no push/PR);
# local-only implement on branch, stop and report
# `done: ready in branch fm/<id> - reviewed by: <mechanism> - <what it found>`
# (no push/PR);
# firstmate reviews, user approves, firstmate merges to local main
# Both modes require an INDEPENDENT review of the diff and name no command to get one.
# Naming one is unsafe in two directions: a harness may refuse to let an agent invoke its
# review command (Claude Code's built-in `/code-review` and `/verify` are both
# disable-model-invocation, i.e. runnable only in a turn a human typed them in, which a
# crewmate never has), and a same-named command from another source may review a different
# artifact (an already-open PR rather than the working diff). The scaffold therefore demands
# the OUTCOME - a fresh reader over the diff - and makes the crew report BOTH the mechanism and
# what the review found, on the status line it already writes. The outcome half is what makes
# the claim checkable: firstmate reads the same diff independently, so "no findings" on a diff
# with obvious defects is the tell that no review ran. A mechanism alone is unfalsifiable.
# The gate is built once by review_gate() and interpolated into both DODs, because two copies
# of one contract drift the moment only one is edited.
# Ship briefs begin with a worktree-isolation assertion before the branch step.
# Scout tasks ignore mode - their deliverable is a report, not a merge - but they still
# carry the direction, because a recommendation that ignores it is worthless.
Expand Down Expand Up @@ -292,6 +309,25 @@ read -r MODE _ <<EOF
$("$FM_ROOT/bin/fm-project-mode.sh" "$REPO")
EOF

# The review gate is identical in both delivery modes, so it is built once here and
# interpolated into each DOD rather than written twice: two copies drift the moment
# only one is edited. Only the line the crew reports it on differs, so that is the
# single parameter. $REPORT_LINE names it (a `review-ready:` or a `done:` line).
review_gate() {
local report_line=$1
cat <<EOF
2. Obtain an INDEPENDENT code review of your diff and fix what it finds.
Do not assume any particular review command exists. This brief deliberately names none: firstmate cannot know which commands your harness exposes, some harnesses refuse to let an agent invoke its review command at all, and a command that shares a name across harnesses can review a different artifact than your working diff.
Use whatever your harness actually provides. A review subagent handed your diff works on every harness and is always an acceptable choice.
It must be a fresh reader over the diff, not you re-reading your own work.
**Report the review on your $report_line, as \`reviewed by: <mechanism> - <what it found>\`.** Name the mechanism, then say what came back: the findings you acted on, or \`no findings\` if it genuinely returned none.
Naming a mechanism but no outcome does not count: firstmate reads the same diff independently, and a review that found nothing worth reporting on a diff with obvious defects is how it detects a review that never ran. Reporting nothing at all means the review did not happen and firstmate will send it back.
Fix the real findings yourself. A finding that turns on a human judgment call - a product choice, a destructive or irreversible action, a security trade-off - is NOT yours to decide: escalate it under rule 6 and stop.
3. Exercise the change end-to-end - drive the affected flow in the real app, not just the tests.
Do not assume a named command for this either; use whatever your harness provides.
EOF
}

case "$MODE" in
local-only)
RULE1="1. Never push to any remote and never open a PR. Work only on your \`fm/$ID\` branch; firstmate handles the merge into local \`main\`."
Expand All @@ -301,11 +337,10 @@ This project ships **local-only**: no remote, no PR.

1. Implement the change and commit it on your branch \`fm/$ID\`. Do NOT push, do NOT open a PR, do NOT merge.
Run the tests your change AFFECTS as you go. Run the project's FULL suite EXACTLY ONCE, at the end, before step 6 - not after every edit.
2. Run \`/code-review\` and address what it finds. Fix the real findings; a finding that is a human judgment call is not yours to decide - escalate it under rule 6.
3. Run \`/verify\` to exercise the change end-to-end - drive the affected flow in the real app, not just the tests.
$(review_gate "\`done:\` line at step 6")
4. Direction check: in one line, state how this change honors the Direction above. If it moves against the direction, stop and escalate under rule 6 instead of shipping it.
5. Keep your branch a clean fast-forward onto the current default branch - if \`main\` has advanced, rebase onto it so the eventual merge stays a fast-forward.
6. Append \`done: ready in branch fm/$ID\` to the status file and stop.
6. Append \`done: ready in branch fm/$ID - reviewed by: {mechanism} - {what it found}\` to the status file and stop, filling in both halves from step 2.

Firstmate then reviews your branch diff against the project's direction, the user approves, and firstmate merges it into local \`main\`.
EOF
Expand All @@ -320,17 +355,16 @@ Firstmate reviews your pushed branch BEFORE any PR exists, so its findings cost

1. Implement the change and commit it on your branch.
Run the tests your change AFFECTS as you go. Run the project's FULL suite EXACTLY ONCE, at the end, before step 6 - not after every edit. Re-running the whole suite per edit was the single biggest time sink measured in this fleet.
2. Run \`/code-review\` and address what it finds.
Fix the real findings yourself. A finding that turns on a human judgment call - a product choice, a destructive or irreversible action, a security trade-off - is NOT yours to decide: escalate it under rule 6 and stop.
3. Run \`/verify\` to exercise the change end-to-end - drive the affected flow in the real app, not just the tests.
$(review_gate "\`review-ready:\` line at step 6, and again in the PR body at step 7")
**State in the PR body how you exercised it.**
4. Satisfy the project's quality hooks. They run automatically on commit and push (secret scan, lint, typecheck, tests). A blocked commit or push means the gate caught something real; fix the cause, never work around the gate.
5. Direction check: in one line, state how this change honors the Direction above.
If the task as specified would move AGAINST the direction, do not quietly implement it - escalate under rule 6.
6. **Push your branch. Open NO PR.** The push makes your work durable; the PR would only make firstmate's review expensive to act on.
Append \`review-ready: branch fm/$ID pushed, no PR\` to the status file and STOP. Firstmate now reviews your diff against the direction.
Append \`review-ready: branch fm/$ID pushed, no PR - reviewed by: {mechanism} - {what it found}\` to the status file and STOP, filling in both halves from step 2. Firstmate now reviews your diff against the direction.
7. Firstmate replies with one of two things:
- **Findings.** Fix them IN PLACE on the same branch, push again, and append \`review-ready:\` again. Repeat until firstmate approves. No PR exists yet, so there is nothing to churn.
- **Approval.** Open the PR with plain \`gh\` (\`gh pr create\`), append \`done: PR {url}\` to the status file, and stop.
- **Findings.** Fix them IN PLACE on the same branch, push again, and append \`review-ready:\` again, reporting the review of what you changed this round. Repeat until firstmate approves. No PR exists yet, so there is nothing to churn.
- **Approval.** Open the PR with plain \`gh\` (\`gh pr create\`), stating in the body how you reviewed and exercised the change, append \`done: PR {url}\` to the status file, and stop.

Do NOT merge the PR, and do not wait for CI yourself. Firstmate watches CI; the user merges.
Once the PR is open your work is on the remote, so firstmate releases your worktree at that point - finish step 7 and stop cleanly.
Expand Down
5 changes: 3 additions & 2 deletions bin/fm-hooks-install.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,9 @@
# Ensure a project worktree has a mechanical quality floor: Claude Code hooks that
# enforce secret-scanning, lint, typecheck, and tests without an agent's cooperation.
# Hooks are the floor that cannot be talked out of it. The judgment layer on top of
# them is the crewmate's own /code-review pass and firstmate's independent,
# direction-aware review of the diff before it reaches the user.
# them is the crewmate's own independent review of its diff and firstmate's independent,
# direction-aware review of the diff before it reaches the user. The crewmate's review
# names no command by design; bin/fm-brief.sh's header owns why.
#
# This is a worktree utility for crewmates, not a supervision script, so it does not
# call fm-guard.sh, and firstmate never runs it against a project clone itself:
Expand Down
7 changes: 6 additions & 1 deletion bin/fm-promote.sh
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,11 @@
# (inventory scratch state, reset to a clean default-branch base, carry over only
# intended fix changes, create branch fm/<task-id>, implement, then report done
# according to the project's delivery mode).
# The scout brief carries no ship review gate and is never regenerated on promotion,
# so those instructions must carry it themselves: an independent review of the diff
# and a `reviewed by: <mechanism> - <what it found>` report. Without it a promoted crew
# cannot satisfy the rule in AGENTS.md's "Review and ship" and would be bounced forever.
# bin/fm-brief.sh's header owns why the gate names no command.
# Usage: fm-promote.sh <task-id>
set -eu

Expand All @@ -26,4 +31,4 @@ mv "$TMP" "$META"

HOME_Q=$(printf '%q' "$FM_HOME")
echo "promoted $ID to ship (teardown protection restored)"
echo "next: FM_HOME=$HOME_Q bin/fm-send.sh fm-$ID '<ship instructions: review scratch state with git status and git log; reset to a clean default-branch base; carry over only intended fix changes; create branch fm/$ID; implement; report done>'"
echo "next: FM_HOME=$HOME_Q bin/fm-send.sh fm-$ID '<ship instructions: review scratch state with git status and git log; reset to a clean default-branch base; carry over only intended fix changes; create branch fm/$ID; implement; get an INDEPENDENT review of your diff by whatever mechanism your harness provides (a review subagent over the diff always works) and fix what it finds; report done with \"reviewed by: <mechanism> - <what it found>\">'"
2 changes: 1 addition & 1 deletion docs/proposals/context-management.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Harness posture: full stack on Claude/Agent-SDK, a bounded-tools-plus-reset floo
## 1. What this is

A long-running firstmate session accumulates context the same way any agent does: every wake, pane peek, crew-state read, review diff, and fleet snapshot adds tokens, and a supervising session lives for hours across many tasks.
Crews carry even more risk per session, because they do the heavy reads, run `/code-review`, `/verify`, and full test suites, and today get zero context guidance in their brief.
Crews carry even more risk per session, because they do the heavy reads, run their own diff review, end-to-end exercise, and full test suites, and today get zero context guidance in their brief.
Left alone, both drift toward the window limit, and the harness falls back to lossy auto-compaction that keeps what it guesses is important rather than what firstmate chose.

This proposal makes context reset a routine, lossless, first-class operation rather than a crash-only event.
Expand Down
Loading