Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
47 commits
Select commit Hold shift + click to select a range
2c93585
refactor(bin): give launch knowledge one owner in fm-launch-lib.sh
sbracewell64 Jul 26, 2026
6e65a07
no-mistakes(review): map fm-launch-lib.sh into test selection, fixtur…
sbracewell64 Jul 26, 2026
7fd8e55
no-mistakes(document): point launch-command docs at new fm-launch-lib…
sbracewell64 Jul 26, 2026
e067656
no-mistakes(review): use verified opencode --auto primary shape and t…
sbracewell64 Jul 26, 2026
dbfd400
no-mistakes(document): point grok adapter launch line at fm-launch-li…
sbracewell64 Jul 26, 2026
abe3c8d
no-mistakes(review): add grok --trust, refuse kimi primary, pin every…
sbracewell64 Jul 26, 2026
bf87246
no-mistakes(review): record primary autonomy evidence and consumer ob…
sbracewell64 Jul 26, 2026
1154677
no-mistakes(review): bind consumer obligation to any bypass flag, add…
sbracewell64 Jul 26, 2026
0f42855
no-mistakes(review): bind consumer obligation to every primary, fix p…
sbracewell64 Jul 26, 2026
a083886
no-mistakes(document): point harness-adapters launch facts at fm-laun…
sbracewell64 Jul 26, 2026
cca9d3a
feat(bin): add the fleet launcher menu, Herdr gate, and reattach guard
sbracewell64 Jul 29, 2026
74be2a6
no-mistakes(review): fix glob split, enforce backend, guard reattach,…
sbracewell64 Jul 29, 2026
fa1dcc7
test: replace launch-lib one-owner source assertions with behavioral …
sbracewell64 Jul 30, 2026
06a6fbd
no-mistakes(review): add launch lock, reject leading-zero input, refr…
sbracewell64 Jul 30, 2026
9dbaa00
Merge fm/fm-launch-menu: fleet launcher menu (upstream PR #1288)
sbracewell64 Jul 31, 2026
d965699
feat(bin): mark crewmate and scout steers as from-firstmate and teach…
sbracewell64 Jul 31, 2026
e960b84
feat(bin): enforce the zero-budget model rule with a model registry, …
sbracewell64 Jul 31, 2026
cf8f2c6
feat(bin): trigger the 70% compaction doctrine from Claude's host-com…
sbracewell64 Jul 31, 2026
3edbc23
CI mirror: feat(bin): add fleet admission control stages 0 and 1 (#8)
sbracewell64 Jul 31, 2026
8e1eb2d
feat(bin): wake firstmate when a monitored PR goes conflicting (#21)
sbracewell64 Aug 2, 2026
0c83af2
test: reap leaked background processes on every suite ending (#16)
sbracewell64 Aug 2, 2026
988dd4b
fix(composer): read an NBSP-padded empty composer as empty and fail u…
sbracewell64 Aug 2, 2026
1ba6ffe
fix(bin): allowlist the opencode turn-end plugin in teardown's dirty …
sbracewell64 Aug 2, 2026
61a5ce8
feat(bin): carry standing worker rules into every generated brief (#23)
sbracewell64 Aug 2, 2026
abc2e7b
fix(bin): prevent treehouse allocation from detaching live work (#25)
sbracewell64 Aug 2, 2026
ecbbe30
fork-landing: Windows-to-WSL launcher bridge reconciled onto trunk (#31)
sbracewell64 Aug 3, 2026
a004e4c
Merge upstream 33a4287 into the fork trunk queue
sbracewell64 Aug 3, 2026
490a371
reconcile: fork trunk with upstream main (0 behind, queue intact, 7 c…
sbracewell64 Aug 3, 2026
1149178
Merge pull request #33 from sbracewell64/fm/fork-upstream-resync-conf…
sbracewell64 Aug 3, 2026
5e6e90f
fix(bin): let a released task land through the sanctioned merge path …
sbracewell64 Aug 3, 2026
ef74752
fix(bin): recognize kimi variants and pi launcher basenames in harnes…
sbracewell64 Aug 3, 2026
817ee8e
CI mirror: refuse local merges only on genuine working-tree collision…
sbracewell64 Aug 3, 2026
631c248
feat(bin): wake-outcome ledger for measured coordinator attention cos…
sbracewell64 Aug 3, 2026
50adfce
fix(bin): measure pool-slot safety against the landing target (#36)
sbracewell64 Aug 4, 2026
cc242ae
feat: add research-approved-work skill and deterministic corpus scann…
sbracewell64 Aug 5, 2026
9f914d1
feat: add the canonical LoopSpec representation, schema and register …
sbracewell64 Aug 5, 2026
ad53e97
reconcile: one landable base from the diverged fork and upstream trun…
sbracewell64 Aug 5, 2026
27c7a2d
fix(bin): stop wedge alarms on settled terminal tasks (#40)
sbracewell64 Aug 6, 2026
4bee9f1
fix(bin): escalate blocked-on-human transitions past status-line dedu…
sbracewell64 Aug 6, 2026
dda7f30
fix: stop rechecking captain-gated pauses in away mode (land of upstr…
sbracewell64 Aug 6, 2026
3611e49
fix(bin): separate task read and contribution bases (land of upstream…
sbracewell64 Aug 6, 2026
f90ed1d
fix(bin): detect child-process work during supervision (land of upstr…
sbracewell64 Aug 6, 2026
aef7a6d
fix(bin): resolve the wake sequence instead of trusting a supplied on…
sbracewell64 Aug 6, 2026
561f0bb
fix(bin): keep the fleet view working past the per-argument cap (#46)
sbracewell64 Aug 6, 2026
744f982
fix(bin): resolve the fork trunk as a landing target (#47)
sbracewell64 Aug 6, 2026
ed376cf
fix(bin): refuse merges without verified green checks (land of upstre…
sbracewell64 Aug 6, 2026
2494ccc
docs(verification): record what CI does not prove about harness adapters
sbracewell64 Aug 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .agents/skills/afk/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,8 @@ Classify each wake this way:
Other signals with no captain-relevant status -> self-handle.
- `signal` or `stale` for a declared `paused:` external wait -> self-handle and track the pause rather than a wedge.
If it remains declared and idle past `FM_PAUSE_RESURFACE_SECS` (default 3600s), housekeeping sends one awaiting-external recheck and resets the pause window.
The recheck exists to re-ask a wait that can change without the captain, so it is skipped for a wait the backlog records as captain-gated (`hold_kind: captain`): that one clears only when the captain acts, and the captain acting is already the away-mode exit, which runs the full return catch-up.
Suppression is a cadence decision only - the wait is still tracked, still reset each window, and still as visible as before in the backlog digest, the fleet view, and the return catch-up - and any kind that cannot be established is rechecked normally rather than dropped.
- `check` -> always escalate. Check scripts print only when firstmate should wake.
- `stale` with a terminal status or bare legacy captain-relevant line -> escalate.
Nonterminal progress remains transient even when its prose contains a legacy free-text token or its seen-status marker already matches, so record a marker and self-handle.
Expand Down
36 changes: 30 additions & 6 deletions .agents/skills/bootstrap-diagnostics/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
name: bootstrap-diagnostics
description: >-
Agent-only handling playbook for session-start bootstrap diagnostics.
Use whenever the session-start digest's bootstrap section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, FLEET_SYNC, PR_CHECK_MIGRATION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, NUDGE_SECONDMATES, or FMX - or when a standalone bin/fm-bootstrap.sh run prints one of those lines.
Use whenever the session-start digest's bootstrap section prints an actionable diagnostic line - MISSING, MISSING_MANUAL, BACKEND_INVALID, NEEDS_GH_AUTH, TANGLE, STARTUP_MEMORY_BUDGET, CREW_DISPATCH invalid, MODEL_REGISTRY, MODEL_PRICE, MODEL_VERIFY, ADMISSION_CONTROL, WAKE_LEDGER, FLEET_SYNC, PR_CHECK_MIGRATION, SECONDMATE_SYNC, SECONDMATE_LIVENESS, SECONDMATE_HANDOFF, NUDGE_SECONDMATES, or FMX - or when a standalone bin/fm-bootstrap.sh run prints one of those lines.
A silent bootstrap section, or a BOOTSTRAP_INFO fact, means no skill load.
user-invocable: false
metadata:
Expand All @@ -19,16 +19,37 @@ When any diagnostic needs captain attention, report the plain consequence and re
- `MISSING: <tool> (install: <command>)` - list the missing tools to the captain with a one-line purpose each plus the printed install commands, wait for consent (one approval may cover the list), then run `bin/fm-bootstrap.sh install <approved tools...>`.
For `treehouse`, this also covers an installed version whose `treehouse get` lacks `--lease`; treat it as an upgrade request.
For `no-mistakes`, this also covers an installed version older than 1.31.2, because crewmate validation briefs delegate gate mechanics to no-mistakes' version-matched guidance.
For `tasks-axi`, this also covers an installed build that fails the compatibility probe (`docs/configuration.md` "Backlog backend" owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report.
For `gh-axi`, this also covers an installed version below the bootstrap-owned floor; treat it as an upgrade request so non-interactive PR merges keep a working bare `--squash` shorthand.
For `tasks-axi`, this also covers an installed build that fails the compatibility probe (`bin/fm-tasks-axi-lib.sh` owns the definition); `config/backlog-backend=manual` only suppresses the verbose `BOOTSTRAP_INFO: tasks-axi available` fact, not this missing-tool report.
For `quota-axi`, bootstrap requires it because firstmate reads its current output directly before resolving every crew-dispatch profile array; without it, report the missing requirement and do not choose around an unexamined candidate.
- `MISSING_MANUAL: <tool> (instructions: <url>)` - tell the captain why the tool is required and give them the printed instructions URL, but do not pass the tool to `bin/fm-bootstrap.sh install`; wait for the captain to complete the manual installation, then rerun session start to confirm the dependency is present.
- `BACKEND_INVALID: <name> (known: <names>)` - the resolved runtime backend has no verified dependency or lifecycle contract, so do not dispatch work until the invalid `FM_BACKEND` or `config/backend` value is corrected to one of the listed backends.
- `NEEDS_GH_AUTH` - ask the captain to run `! gh auth login` (interactive; you cannot run it for them).
- `TANGLE: <remediation>` - the primary checkout is stranded on a feature branch instead of its default branch; `AGENTS.md` section 8 explains why this guard exists and what it protects.
The work is safe on that branch ref; restore the primary to its default branch with the printed `git -C <root> checkout <default>`, then re-validate that branch in a proper worktree.
This is the only sanctioned firstmate-initiated git write to the primary, and it is a non-destructive branch switch that strands nothing.
- `STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - <reason>` - the visible startup-memory budget is not a safe one-line positive decimal file; do not infer the default or propagate it. Correct the local primary file, then rerun session start so the normal convergence path can deliver the validated value to secondmate homes.
- `STARTUP_MEMORY_BUDGET: invalid config/startup-memory-budget - <reason>` - the visible startup-memory budget is not a safe one-line positive decimal file; do not infer the default or propagate it.
Correct the local primary file, then rerun session start so the normal convergence path can deliver the validated value to secondmate homes.
- `CREW_DISPATCH: invalid config/crew-dispatch.json - <reason>` - the optional dispatch profile file exists but failed low-cost bootstrap validation; stop profile-based dispatch, report the actionable error, and require correction of the malformed schema, unverified harness name, or invalid harness/effort pair rather than falling back around it or selecting a bad profile.
- `MODEL_REGISTRY: invalid config/models.json - <reason>` - the model registry exists but failed schema validation, so every provider-prefixed model is now refused at spawn until it is corrected.
Fix the registry; never delete it to clear the error, because deleting it silently disables zero-budget enforcement rather than restoring it.
- `MODEL_REGISTRY: <model> <problem>` - the dispatch config and the registry disagree: the model is unregistered, carries a non-approved status, or has no current live-probe record.
This is the check that catches a bad model before any worker is launched against it, so correct the dispatch rule or complete the model's admission; do not weaken the check to make the line go away.
- `MODEL_REGISTRY: no config/models.json, ...` - routed provider models exist but nothing enforces the zero-budget rule for them.
This is a standing gap rather than a failure, and it stays inert by design; raise it with the captain rather than treating it as a startup blocker.
- `MODEL_PRICE: <model> ... no longer zero` - an allowlisted model is no longer free at its provider, which is the exposure a name-only allowlist cannot see.
Suspend that route immediately, then re-verify its cost class before it is routed to again.
- `MODEL_PRICE: <model> price drifted ...` or `... catalogue source is unreadable` - re-verify the cost class and update the recorded price, or repair the declared catalogue path so the check stops being blind.
- `MODEL_VERIFY: <model> ...` - a live probe failed. Load `model-onboarding` and read the shape before reacting: a provider refusal means this account can never use that model and it must leave routing, while a local failure is a configuration error on this machine and not a provider fact.
- `ADMISSION_CONTROL: invalid config/crew-dispatch.json _scheduling.admission_control - <reason>` - the optional fleet-admission policy exists but failed schema validation, so the fleet's admission layer cannot resolve a band.
Do not dispatch new work in this home until the named field is corrected; an unknown field is refused rather than ignored precisely so a typo cannot silently disable a safety condition.
`docs/configuration.md` "Fleet admission control" owns the schema, and `fleet-admission` owns what firstmate does with a resolved band.
- `WAKE_LEDGER: <n> outcome record(s) join no wake record` - that many recorded supervision costs point at a wake this home never drained, so every figure drawn from the ledger overcounts by up to that number.
Treat it as a measurement defect, never as supervision work: report the count rather than any rate or total computed from the file, until the captain decides what to purge.
The records are append-only evidence, so never rewrite or migrate the file to clear the count; `bin/fm-wake-ledger.sh reconcile` restates it on demand and that script owns the join.
A count that grows during a session means something is still recording against unresolvable sequences, which is a bug to escalate rather than a backlog of old damage.
- `WAKE_LEDGER: the wake ledger could not be read ...` - the file exists but could not be opened, so the count above is unavailable rather than zero.
Repair its permissions or path before quoting any supervision-cost figure; an unreadable ledger is reported precisely so it cannot pass as a clean one.
- `FLEET_SYNC: <repo>: skipped: <reason>` - a benign one-off skip (offline, no origin, local-only); bootstrap continued, investigate only if it blocks work.
A skip can also report the bounded fleet-refresh timeout (`FM_FLEET_SYNC_BOOTSTRAP_TIMEOUT`, or a fleet-size-aware default with a 20 second floor); a timeout never blocks startup.
- `FLEET_SYNC: <repo>: recovered: <detail>` - the clone had drifted onto a clean detached HEAD holding no unique commits and the sync self-healed it (re-attached the default branch and fast-forwarded); no action needed, it is reported only so the self-heal is visible.
Expand All @@ -44,10 +65,13 @@ When any diagnostic needs captain attention, report the plain consequence and re
Resume the emitted supervision protocol after finishing the session-start wake handling.
- Any other `PR_CHECK_MIGRATION:` refusal means migration did not complete safely, whether because watcher exclusion, a private path, a diagnostic, quarantine validation, or marker publication could not be proved.
Keep each affected poll unavailable, inspect the named private state path, and do not bypass the migration or execute a quarantined artifact; a completed safe-scan marker allows unrelated authenticated polls to continue while private repair remains pending.
- `SECONDMATE_SYNC: secondmate <id>: skipped: <reason>` - the local-HEAD secondmate sync left a live secondmate home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing the primary target commit, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update.
- `SECONDMATE_SYNC: secondmate <id>: skipped: <reason>` - secondmate convergence left a live home on its existing checkout because the home was dirty, diverged, unsafe, on the wrong branch, missing its placement-specific target commit, unreachable, or otherwise not fast-forwardable, or because inherited local-material propagation failed; bootstrap continued, but inspect the reason because the secondmate's tracked instructions, inherited settings, or shared captain preferences may be stale after a primary update.
- `SECONDMATE_LIVENESS: secondmate <id>: skipped: <reason>|respawn failed after <cause>: <reason>` - the session-start liveness sweep could not guarantee that the registered secondmate is running a real agent process.
Investigate the reason because that secondmate is not guaranteed live.
- `NUDGE_SECONDMATES: secondmate <id>: send failed: <reason>` - the secondmate sweep fast-forwarded a running secondmate home and its loaded instruction surface (`AGENTS.md`, `bin/`, or `.agents/skills/`) changed, but the deterministic `fm-send.sh fm-<id>` re-read nudge failed.
Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send.
- `SECONDMATE_HANDOFF: secondmate <id>: pending delivery: <n> item(s)` - queued work has already left the main dispatchable backlog and remains safe in the named remote route's backlog-format outbox.
Preserve that outbox and rerun `bin/fm-backlog-handoff.sh --resume-pending` after same-host connectivity returns; never re-add or dispatch the items from the main backlog.
An unsafe-outbox variant requires path and file-type inspection before any retry.
- `NUDGE_SECONDMATES: secondmate <id>: send failed: <reason>` - secondmate convergence changed a running home's loaded instructions or inherited config, but the deterministic `fm-send.sh fm-<id>` re-read nudge failed.
Inspect the reason, keep the pending marker under `state/.secondmate-nudge-pending/` intact, and rerun session start after the endpoint or metadata issue is fixed so bootstrap can retry the exact same marked send on the same local or remote route.
- `FMX: X mode on ...` / `FMX: X mode off ...` - bootstrap confirmed or removed the local X-mode poll artifacts (`docs/configuration.md` "X mode (.env)").
Only when a running watcher needs the cadence transition applied immediately, restart the home-scoped watcher through the emitted harness supervision protocol; bootstrap deliberately never restarts the watcher itself.
29 changes: 29 additions & 0 deletions .agents/skills/firstmate-coding-guidelines/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,11 @@ Briefs for tasks that touch firstmate's own tracked material should tell the cre
Firstmate adds this skill's load instruction to firstmate-repo briefs by hand instead.
`CONTRIBUTING.md`'s "Development" section carries the same instruction as a durable reminder.

## Generated worker discipline

For firstmate-repo work, the generated `# Branch conflict resolution` and `# Verification discipline` sections in [`bin/fm-brief.sh`](../../../bin/fm-brief.sh) own the full worker rules.
Follow both for branch shipping and verification even if the current task was scaffolded before those sections existed.

## Compatibility and enforcement

Before changing shared tracked behavior, review every affected supported primary harness and runtime backend rather than checking only the adapters active in the current fleet.
Expand All @@ -81,6 +86,29 @@ Mark an axis not applicable only after inspecting its integration surface, and u
For critical safety, routing, startup, and supervision infrastructure, prefer deterministic and idempotent enforcement over relying on agent memory alone.
Keep instructions as the authority and discovery layer, but make repeated execution converge safely and make invalid or unsafe states fail closed wherever the runtime can enforce them.

### Harness-dependent checks

This section is the single owner of the rule and of how to satisfy it.

A check is harness-dependent when its verdict comes from something the vendor emits: a process name, rendered output, a spinner or keybind glyph, a banner, or a key the harness binds.
Anything in that class must be proven end to end against the real harness, because a stub or fake agent can only confirm the assumption already written into the stub.
That proof is authorized to spend tokens; the cost is small against a check that silently stops working.

Build the check on the most structural signal that answers the question, and prefer a kernel or protocol fact over anything a release note could change.
When a rendered surface is genuinely the only source, read more than one independent signal and let any of them carry a positive verdict, so no single vendor string is load-bearing.
Where a surface signal is unavoidable, back it with a guard that fails loudly naming the harness and version rather than degrading quietly.

Every such check needs two tests, because they fail for different reasons:

- A portable regression in `tests/` that pins the logic with real processes and no harness, so CI enforces the classifier everywhere it runs tmux.
Drive the signals apart deliberately and assert the verdict survives losing one; assert the divergence itself so the case cannot go quietly vacuous.
Confirm which signal a given construction actually blinds on each supported platform rather than assuming, because the same trick can break different sources on macOS and Linux.
- A live guard in the `live-harness-optin` family (`bin/fm-test-run.sh`), env-gated and self-skipping, that exercises every INSTALLED harness for real and fails naming the harness and version.
Report an absent harness explicitly rather than passing silently over it, and refuse a pass that checked nothing.
This guard is opt-in and on-demand because standard CI has neither harness binaries nor credentials; run it after every harness upgrade and before trusting refreshed per-harness evidence.

Record the dated per-harness result in `docs/verification/runtime-backends.md`, and point at the live guard as the command that refreshes it, rather than leaving a version-scoped observation to rot into a false claim.

## Documentation change review

For every changed maintained prose surface, identify its inventory audience, authoritative owner, current-behavior relevance, destination for supporting evidence, and any unique safety fact that removal could lose.
Expand All @@ -98,6 +126,7 @@ Run `bin/fm-doc-audience-check.sh`; it enforces classification, README setup rou
- Run `bin/fm-lint.sh` before treating a script change as done; it is the single owner of the lint definition (file set, config, and pinned shellcheck version) that CI and the no-mistakes pre-push gate both invoke, and it refuses to run under any other shellcheck version.
- Colocate tests with the existing pattern in `tests/`, name them `<subject>.test.sh`, and extend an existing script rather than inventing a new runner.
- Tests must exercise behavior through an executable or public interface and must never assert implementation-source bytes, including through parsers, regexes, snapshots, or indirect wrappers.
- Hand every background process a test launches to `fm_test_reap` as soon as its pid is known, because a case that only kills on its happy path orphans that process on every failing or signalled path.
- A maintainer-verification record under `docs/verification/` records active empirical facts, not assumptions or task chronology.
- Include the date, version, exact commands run, and exact output needed to support the current guarantee.
- Keep incident chronology and delivery evidence in private task reports or PR evidence unless a concise rationale is required to maintain a current safety boundary.
Loading