Skip to content

fix(composer): read a blank-padded composer as empty, not as unsent text - #2047

Closed
huynhtandat223 wants to merge 39 commits into
kunchenguid:mainfrom
huynhtandat223:fm/fm-pair-delivery-false-negative-fix
Closed

huynhtandat223 wants to merge 39 commits into
kunchenguid:mainfrom
huynhtandat223:fm/fm-pair-delivery-false-negative-fix

Conversation

@huynhtandat223

Copy link
Copy Markdown

What broke

fm-pair-compose.sh releases both pair barriers through verified fm-send submission. The Claude driver visibly received PAIR READY, but fm-send exited 1 with delivery unconfirmed; verdict=pending. The composer treated that correctly under its contract and aborted before releasing the navigator, so ready was never published and psak-dms-kernel-canonicalization stalled before its plan gate with no navigator coverage.

Trigger, masking condition, symptom

  • Initiating trigger. Claude 2.1.226 pads its empty composer row with U+00A0 after the ❯ glyph. Verified bytes, byte-identical through both the Herdr ANSI read and a plain tmux capture:

    033 [ 0 m 033 [ 3 8 ; 2 ; 1 5 3 ; 1 5 3 ; 1 5 3 m 342 235 257 302 240 033 [ 0 m
    

    The blank is drawn at luminance 153, above the de-emphasis threshold, so fm_composer_strip_ghost correctly keeps it as real text. Every trim in the shared classifier and its callers is ASCII-only, so a genuinely empty Claude composer classified pending.

  • Masking condition. The composer verdict only decides delivery where no stronger signal exists. fm_backend_herdr_send_text_submit confirms from native agent state whenever the target is legibly idle before Enter, and only falls back to reading the composer when the target was already mid-turn. Every ordinary steer took the immune path. The paired-review barrier release takes the other one, because both roles sit in a blocking wait for the release file when it is sent.

  • Visible symptom. verdict=pending on a message that had already landed, and a pair aborted between its two releases.

The post-submit composer is the same padded-empty row plus a dim Press up to edit queued messages placeholder, which ghost stripping already removes correctly. One fault explains the pre-send read, the post-send read, and the retry exhaustion.

Proven path vs failing path

A known-good send and the failing send differ at exactly one branch: the pre-Enter agent-state baseline. idle goes to fm_backend_herdr_wait_for_working (never reads the composer, always worked). working goes to fm_backend_herdr_composer_state (reads the padded row, always failed). Pre-fix, an empty composer and a composer holding real text were indistinguishable - both pending:

PRE-FIX  classify 0 "❯<NBSP>"           = pending     <- empty composer
PRE-FIX  classify 0 "❯<NBSP>PAIR READY" = pending     <- real unsent text

Disconfirming checks run

  • Is the NBSP a Herdr rendering artifact rather than Claude's? No. A plain tmux capture of Claude, with no Herdr in the path, shows 342 235 257 302 240 - one row in the pane, the composer row.
  • Does Claude keep the text in the composer while busy (the OpenCode busy-queue shape)? No. It queues the message, clears the composer, and shows the queued text above. So a busy-pane "still pending" reading is a genuine swallow, which is why this fix must not copy tmux's busy-queue fallback - that would convert this false negative into a false positive.
  • Was my first live guard actually testing a composer? No - it accepted Claude's blank startup screen and its trust dialog, and passed without the fix. Rebuilt so the untouched verdict is only asserted after typed probe text becomes visible in that pane. It now fails on Claude pre-fix.

The fix

Fold non-ASCII Unicode blanks to ASCII space in fm_composer_classify_content - the single fleet-wide owner the tmux, herdr, cmux and orca adapters all delegate to.

This can only move content that is entirely blank from pending to empty. Content carrying any visible character keeps its verdict, so unproven delivery still stops the pair and the away-mode injector still refuses to type over unsent input. Keying on the Unicode space-separator category rather than U+00A0 alone keeps this from recurring when a release picks a different blank.

No change to the pair protocol, and no weakening of delivery verification.

Real-harness evidence

Against real claude 2.1.226 on herdr 0.7.3, through the production adapter function, both directions:

### DIRECTION 1 - accepted while the worker is mid-turn MUST confirm
pre-send: agent_status=working composer_state=empty
send_text_submit VERDICT=empty
### DIRECTION 2 - a genuine swallow MUST still refuse
typed but NOT submitted: composer_state=pending
after composer clear: composer_state=empty

Direction 1's probe was observably delivered - it arrived in the receiving agent's turn - while the pre-fix adapter called the same send unconfirmed.

Coverage

  • tests/fm-composer-lib.test.sh - the fold, every declared blank, the dead-shell refusal, and the divergence assertion (padded-empty is empty while padded-real-text stays pending) so the case cannot go vacuous.
  • tests/fm-backend-herdr.test.sh - the busy-baseline submit path itself, both directions, on captures carrying the real bytes and Claude's real rule-delimited shape. Fails pre-fix.
  • tests/fm-composer-blank-drift-live-e2e.test.sh - new opt-in real-harness guard (FM_COMPOSER_BLANK_DRIFT=1), registered in the live-harness-optin family; bin/fm-composer-lib.sh now selects that family. What a harness draws in an empty composer is a vendor surface, so a fixture can only ever replay a previous release.

Observed: claude 2.1.226 and pi 0.84.1 pass; codex 0.145.0 held a first-run trust screen and is reported unverified; opencode 1.17.20 draws a left-only vertical composer edge the tmux reader refuses as unknown - reported unverified, not failed, since unknown already denies both delivery confirmation and injection.

Compatibility reviewed across all four consuming adapters (suites pass); zellij does not use the classifier. bin/fm-lint.sh clean, bin/fm-doc-audience-check.sh clean.

Notes for the reviewer, outside this change

  • tests/fm-backend-herdr.test.sh already fails on main at two same-labeled home workspaces with no launcher identity must refuse (expected exit 3, got 1). Confirmed pre-existing on a clean HEAD; untouched here.
  • fm-pair-compose.sh releases the driver and navigator in sequence and creates ready only after both. A failure on the second release still leaves a half-released pair needing manual repair. Out of scope here - the proven fix does not require it - but worth its own decision.
  • opencode's composer being unreadable through the tmux reader means firstmate cannot confirm delivery to an opencode worker from composer state on that backend. Fail-safe, pre-existing, and not in this task's scope.

🤖 Generated with Claude Code

huynhtandat223 and others added 30 commits August 8, 2026 00:53
* fix(herdr): admit agy plain-ASCII '> ' composer via native identity

agy's interactive composer draws a bare ASCII "> " prompt with no side
border and no agent-specific glyph, so its empty-composer row is
byte-for-byte the dead-shell prompt the shared classifier in
bin/fm-composer-lib.sh deliberately refuses as unknown (never empty).
Under herdr that made a live-empty agy composer fall to unknown, unlike
bordered/agent-glyph shapes, so away-mode injection never fired.

Close the gap inside the herdr adapter only, mirroring the Pi separator
composer: locate the bottom-most "> " row structurally and admit it as
a real composer solely when native `agent get` corroborates the target
is agy and reports idle/done/blocked. A working agy, a non-agy or
unregistered agent, or an unreadable identity keeps the bare row unknown,
so the shared dead-shell rule is not weakened. Identity is fetched at
most once per call, shared with the Pi block.

* feat: add agy (Antigravity CLI) harness adapter

Wire Google Antigravity CLI (agy 1.1.5, Gemini-backed) as a crewmate harness.

- fm-harness.sh: detect agy via ANTIGRAVITY_AGENT=1 env marker plus
  agy/antigravity process-ancestry match.
- fm-spawn.sh: interactive launch template
  (agy --dangerously-skip-permissions -i "$(cat brief)"); add agy to the
  model flag (--model), effort flag (--effort, ceiling high), and the
  accepted-harness lists.
- fm-tmux-lib.sh / fm-watch.sh: add busy signature "esc to cancel" to the
  default busy regex (daemon inherits it).
- harness-adapters SKILL.md: full agy knowledge section, launch-profile-axes
  row, and no-mistakes invocation note.
- fm-session-start.test.sh: drop ANTIGRAVITY_AGENT in detect_own so the
  suite does not leak the marker.

Verified empirically on agy 1.1.5: env marker, interactive cwd inheritance,
--dangerously-skip-permissions autonomy, busy/idle footer, first-run trust
dialog. Turn-end hook and primary-session guard are NOT yet verified, so agy
is crewmate-only for now (documented in the skill).

* fix: read live pi sessions as live in session-lock liveness

fm_harness_pid_alive joined `basename comm` and `ps -o args=` into one
line before matching FM_HARNESS_RE. That broke every anchored
alternative: a live pi harness joins to "pi pi", which `^pi$` cannot
match, so each live pi session read as stale. bin/fm-lock.sh then let a
second session take the primary's lock, and bin/fm-claude-stop-autoarm.sh
reclaimed a lock whose owner was alive - a whole-home mutual-exclusion
failure that can double-dispatch or double-teardown.

Split the check the way fm_harness_ancestry_pid already does it: match
the basename of comm on its own line, and consult args only when comm is
a bare interpreter (node/python). FM_HARNESS_RE is unchanged, so `^pi$`
stays anchored and pip/pipx/paths containing "pi" keep reading as
not-a-harness.

tests/fm-session-lock-liveness.test.sh covers both directions against
real spawned processes (no fake ps) and fails on the pre-fix library.

* docs: record the gh-axi repo-resolution trap and the turn-boundary stall pattern

* fix: verify pruned resolved captain holds from the done archive

tasks-axi's Done-retention prune archives resolved captain holds out of the
live backlog (closing the originating task auto-prunes), and the completion
gate could then never verify them again: teardown refused, complete failed
with 'absent from data/backlog.md', and the task deadlocked with no
non-bypass exit (hit live with psak-compose-bringup on 2026-07-26).

Prune archives rather than deletes, so the durable resolution record
survives byte-identically in the markdown backend's archive file. Every
read that accepts a Done hold now falls back to that archive: complete,
verify, identical resolve retries, and reopen refusal are order-independent
with respect to pruning. An archived row without the resolution record
still fails the gate, and an active hold is never read from the archive,
so the safety boundary is unchanged - the gate verifies the same evidence
it always required, from where the backend moved it.

Two regression tests reproduce the deadlock sequence end-to-end and prove
the fail-closed path for an archived hold with no resolution record.

* docs: add the paired driver-and-navigator protocol as an agent-only skill

High-blast-radius ship work now runs as a pair launched together, exchanging
directly through a durable shared file rather than through firstmate, which was
the latency bottleneck when it supervises several threads at once. The protocol
body lives in the new agent-only `paired-review` skill; AGENTS.md gains only the
trigger, the narrowed communication rule, and the delivery-path exception.

The second worker is a navigator, not a reviewer: three of the four gates have
no diff to look at, and an agent told it reviews will idle until one exists,
which defeats the design. The plan gate is the cheap one - a wrong-layer
direction error is visible from a file-level plan in about two minutes and
invisible at PR time under hours of work built on it.

The pair may never settle a contract or scope change, anything destructive,
irreversible, or security-sensitive, or a disagreement surviving its one round
trip; removing firstmate from the exchange moves no approval authority to the
crewmates.

Also make the scope and seam statement a default section of every ship and
scout brief, since the direction error that motivated this protocol happened on
a brief that never named the module owning the code.

* docs: exclude diagnosis from pairing and give navigators four structural questions

Three rules the paired-review protocol did not carry, plus the rule that
keeps skill changes written against the skill-writing guide.

A diagnosis is not paired work. Pairing applies to implementation work; a
bug, regression, or crash-loop takes the diagnose-then-fix shape, one
worker following the diagnosing-bugs skill by absolute path. Until a root
cause exists there is no "where" for a navigator to hold a view about,
and the risk a bug carries is a wrong root cause rather than a wrong
layer, so sequence guards it where parallelism cannot.

The root-cause gate: firstmate judges the root cause before any fix
action. Confident, it authorises the fix and reports; otherwise it
escalates and waits. The root-cause document's own standard is the
confidence test. Destructive, irreversible, security-sensitive, and
contract-expanding fixes still go to the captain regardless.

Every navigator brief carries four fixed structural questions alongside
its task-specific checks: was the stated reason delivered, where did the
coupling go, does this add another instance of a shape the repo has been
burned by, and what did the brief not specify that the implementation had
to decide anyway. Each keeps its reason, because a checklist cannot reach
the destination - a navigator is pointed at a diff, and a diff shows what
changed, never whether the change moves toward where the work is going.

firstmate-coding-guidelines gains one requirement: a skill change begins
by loading writing-great-skills and is written against it.

AGENTS.md changes are two edited trigger lines and no new lines; the
rules themselves live in the skills.

* docs: add program orchestration skill

* docs: pin the paired navigator and state paired review's ownership boundary

Replace the paired-review default-routing guidance with a deterministic
navigator rule: every paired navigator runs on pi with the model
cx/gpt-5.6-terra, and only an explicit per-task captain choice of another
navigator runtime or model replaces it. Ordinary dispatch resolution keeps
owning the concrete launch mechanics and carries the pin through instead of
substituting a best-fit alternative.

State the ownership boundary in the same file: paired review is a universal
implementation-task execution protocol owned by whichever firstmate owns the
task, so a standalone paired task needs no persistent orchestrator, while a
program orchestrator dispatches the same protocol unchanged.

Drop the reroute-to-claude clause under the handed-out-skill limitation, which
would otherwise silently replace the pinned navigator runtime.

* feat(skills): add the captain-invoked planner session

A planner session is one temporary worker the captain talks to directly.
Firstmate opens it, points it at a captain-defined project, scope, and
planning question, then leaves the conversation and monitors only whether
the session is alive, waiting, finished, or failed.

The planner investigates before it asks anything, reconstructing both what
the code does today and where the project's accepted decisions say it is
going, and surfaces every conflict between them instead of quietly choosing.
It reads tests as evidence and runs nothing. It then grills the captain one
question at a time, each question carrying its recommended answer and each
prescriptive choice carrying both the simplest workable and the best-practice
case. It may not produce a plan until the captain opens the crystallize gate,
after which it publishes a spec or dependency-aware tickets through the
project's own established conventions. Planning never authorizes
implementation.

Two artifact contracts travel downstream, each with one owner:

- The scope envelope carries the accepted boundary at spec level and per
  ticket, reusing bin/fm-brief.sh's existing scope and seam vocabulary plus
  the two fields a brief cannot carry before the target is chosen.
- The test contract carries acceptance intent, not executable test code, and
  the captain approves it together with scope.

planner owns those artifact fields. program-orchestration owns consuming
them: preserve provenance, revalidate, narrow per worker, never widen,
escalate a stale envelope. bin/fm-brief.sh keeps the final worker statement
and paired-review keeps sharing it with the driver and navigator, plus the
navigator's plan-gate challenge to the approved test contract.

The planning disciplines are vendored from mattpocock/skills at 2ab9580 with
the MIT license and a provenance record, adapted only in paths, tracker
assumptions, and delegation wording. No bundled file is named SKILL.md, so
the bundle stays off every harness's skill index and costs no context; it
also cannot silently clobber a captain's own locally installed copy, which
git overwrites without a conflict on fast-forward when the install path is
locally excluded.

Launch reuses the ordinary brief, spawn, and scout-report mechanics with no
new runtime: pi cx/gpt-5.6-sol at high effort by default, claude
claude-opus-5 when explicitly selected.

Also fixes fm-doc-audience-check.sh treating markdown links inside fenced
code blocks as navigation targets, which a vendored example listing exposed.

* feat(skills): give the navigator a question form and the decision history

A navigator could only assert findings, so a question had to be dressed as a
defect or stay silent, and nothing systematically told it what had already been
decided. One program's navigator invented an undeclared Q entry to work around
exactly this, and another raised a cross-ticket scope question as a finding,
forcing the driver to adjudicate something it had no authority over.

paired-review gains an evidence-backed Q<k> form beside N<k>, on the same one
round trip but with no forced accepted/rejected verdict. The owning firstmate
answers by default and answers it itself, escalating upward only when it lacks
the authority or the knowledge; the driver answers only its own implementation
choices. Navigator briefs carry the decisions bearing on the ticket where a
record exists, and the skill now states that cross-package and whole-solution
direction sits outside one navigator's field of view.

program-orchestration takes the matching duties: maintaining the program's
cross-ticket decision record, handing the bearing decisions to both halves of
every dispatched pair, and holding direction across tickets.

N semantics and the four structural questions are unchanged, and no fifth
question was added. Both briefs are written at dispatch, so this reaches the
next dispatch and leaves a pair already under way on the protocol it launched
with.

* fix(skills): state the Q id discipline the N form already carries

The Q form said a question runs "its own sequence" but never said that
sequence is sequential and never reused, so the identifier rule was
checkable for N and not for Q. Match the N wording.

* fix(herdr): place the lab --session ahead of any -- separator

fm_herdr_lab_raw appended `--session <name>` to the end of the herdr argv.
`--session` is a herdr global option, and `herdr agent start <name> ... -- <argv...>`
hands every token after `--` to the inner command, so for any subcommand with a
separator the flag never reached herdr: the call silently ran against whatever
session herdr resolved by default, and the inner command received two junk
arguments. The reported repro, `agent start envcheck -- bash`, executed as
`bash --session fm-lab-...` and the pane died immediately.

The confirmed hypothesis is argument placement, not content. Placement was
inherited from fm_backend_herdr_cli, whose comment records that leading position
was already verified to work and that trailing was chosen only to keep call-site
diffs append-only; the blind spot in that choice is the `--` separator.

Every rejection rule is unaffected: the caller-supplied --session scan, the
leading-option refusal, the lifecycle and server refusals, and the fresh
refuse-default checks all run in fm_herdr_lab_cli on the caller's arguments
before fm_herdr_lab_raw assembles the herdr argv.

The existing behavior test asserted the trailing position as the contract, so the
suite was green on the defect; its assertions are re-pointed, and a new case
records the herdr argv token by token and asserts the session pair leads and
never lands past a `--`, with the inner command passed through untouched.

The identical latent pattern in fm_backend_herdr_cli is deliberately out of scope
here; no current backend call site uses a `--` separator.

* feat(skills): add the captain-invoked orchestrator session

Open a long-running programme orchestrator the captain talks to directly,
driving an already-authorized spec, ticket set, or GitHub issues to
completion through workers the session dispatches itself.

The skill splits across the branch boundary it actually has: firstmate
reads SKILL.md to open the session, and only the session reads
CONTRACT.md, so the judgment written for the orchestrator costs
firstmate a pointer rather than a context load. program-orchestration
keeps ownership of custody, routing, host ramp, envelope consumption,
and handoff; both files point at it rather than restating it.

AGENTS.md gains the third direct-conversation exception and one
deliverable-classification bullet. The planner test's closed-exception
assertion moves to the new count rather than dropping the invariant.

* fix: reconcile adapter and documentation replay conflicts
agy's interactive composer draws a bare ASCII "> " prompt with no side
border and no agent-specific glyph, so its empty-composer row is
byte-for-byte the dead-shell prompt the shared classifier in
bin/fm-composer-lib.sh deliberately refuses as unknown (never empty).
Under herdr that made a live-empty agy composer fall to unknown, unlike
bordered/agent-glyph shapes, so away-mode injection never fired.

Close the gap inside the herdr adapter only, mirroring the Pi separator
composer: locate the bottom-most "> " row structurally and admit it as
a real composer solely when native `agent get` corroborates the target
is agy and reports idle/done/blocked. A working agy, a non-agy or
unregistered agent, or an unreadable identity keeps the bare row unknown,
so the shared dead-shell rule is not weakened. Identity is fetched at
most once per call, shared with the Pi block.
Wire Google Antigravity CLI (agy 1.1.5, Gemini-backed) as a crewmate harness.

- fm-harness.sh: detect agy via ANTIGRAVITY_AGENT=1 env marker plus
  agy/antigravity process-ancestry match.
- fm-spawn.sh: interactive launch template
  (agy --dangerously-skip-permissions -i "$(cat brief)"); add agy to the
  model flag (--model), effort flag (--effort, ceiling high), and the
  accepted-harness lists.
- fm-tmux-lib.sh / fm-watch.sh: add busy signature "esc to cancel" to the
  default busy regex (daemon inherits it).
- harness-adapters SKILL.md: full agy knowledge section, launch-profile-axes
  row, and no-mistakes invocation note.
- fm-session-start.test.sh: drop ANTIGRAVITY_AGENT in detect_own so the
  suite does not leak the marker.

Verified empirically on agy 1.1.5: env marker, interactive cwd inheritance,
--dangerously-skip-permissions autonomy, busy/idle footer, first-run trust
dialog. Turn-end hook and primary-session guard are NOT yet verified, so agy
is crewmate-only for now (documented in the skill).
fm_harness_pid_alive joined `basename comm` and `ps -o args=` into one
line before matching FM_HARNESS_RE. That broke every anchored
alternative: a live pi harness joins to "pi pi", which `^pi$` cannot
match, so each live pi session read as stale. bin/fm-lock.sh then let a
second session take the primary's lock, and bin/fm-claude-stop-autoarm.sh
reclaimed a lock whose owner was alive - a whole-home mutual-exclusion
failure that can double-dispatch or double-teardown.

Split the check the way fm_harness_ancestry_pid already does it: match
the basename of comm on its own line, and consult args only when comm is
a bare interpreter (node/python). FM_HARNESS_RE is unchanged, so `^pi$`
stays anchored and pip/pipx/paths containing "pi" keep reading as
not-a-harness.

tests/fm-session-lock-liveness.test.sh covers both directions against
real spawned processes (no fake ps) and fails on the pre-fix library.
tasks-axi's Done-retention prune archives resolved captain holds out of the
live backlog (closing the originating task auto-prunes), and the completion
gate could then never verify them again: teardown refused, complete failed
with 'absent from data/backlog.md', and the task deadlocked with no
non-bypass exit (hit live with psak-compose-bringup on 2026-07-26).

Prune archives rather than deletes, so the durable resolution record
survives byte-identically in the markdown backend's archive file. Every
read that accepts a Done hold now falls back to that archive: complete,
verify, identical resolve retries, and reopen refusal are order-independent
with respect to pruning. An archived row without the resolution record
still fails the gate, and an active hold is never read from the archive,
so the safety boundary is unchanged - the gate verifies the same evidence
it always required, from where the backend moved it.

Two regression tests reproduce the deadlock sequence end-to-end and prove
the fail-closed path for an archived hold with no resolution record.
…kill

High-blast-radius ship work now runs as a pair launched together, exchanging
directly through a durable shared file rather than through firstmate, which was
the latency bottleneck when it supervises several threads at once. The protocol
body lives in the new agent-only `paired-review` skill; AGENTS.md gains only the
trigger, the narrowed communication rule, and the delivery-path exception.

The second worker is a navigator, not a reviewer: three of the four gates have
no diff to look at, and an agent told it reviews will idle until one exists,
which defeats the design. The plan gate is the cheap one - a wrong-layer
direction error is visible from a file-level plan in about two minutes and
invisible at PR time under hours of work built on it.

The pair may never settle a contract or scope change, anything destructive,
irreversible, or security-sensitive, or a disagreement surviving its one round
trip; removing firstmate from the exchange moves no approval authority to the
crewmates.

Also make the scope and seam statement a default section of every ship and
scout brief, since the direction error that motivated this protocol happened on
a brief that never named the module owning the code.
…ral questions

Three rules the paired-review protocol did not carry, plus the rule that
keeps skill changes written against the skill-writing guide.

A diagnosis is not paired work. Pairing applies to implementation work; a
bug, regression, or crash-loop takes the diagnose-then-fix shape, one
worker following the diagnosing-bugs skill by absolute path. Until a root
cause exists there is no "where" for a navigator to hold a view about,
and the risk a bug carries is a wrong root cause rather than a wrong
layer, so sequence guards it where parallelism cannot.

The root-cause gate: firstmate judges the root cause before any fix
action. Confident, it authorises the fix and reports; otherwise it
escalates and waits. The root-cause document's own standard is the
confidence test. Destructive, irreversible, security-sensitive, and
contract-expanding fixes still go to the captain regardless.

Every navigator brief carries four fixed structural questions alongside
its task-specific checks: was the stated reason delivered, where did the
coupling go, does this add another instance of a shape the repo has been
burned by, and what did the brief not specify that the implementation had
to decide anyway. Each keeps its reason, because a checklist cannot reach
the destination - a navigator is pointed at a diff, and a diff shows what
changed, never whether the change moves toward where the work is going.

firstmate-coding-guidelines gains one requirement: a skill change begins
by loading writing-great-skills and is written against it.

AGENTS.md changes are two edited trigger lines and no new lines; the
rules themselves live in the skills.
The snapshot passed the parsed backlog and task inventory to jq as --argjson
argv strings. A single argument is capped by MAX_ARG_STRLEN (128 KiB), a
separate and far smaller limit than ARG_MAX, so once a home's backlog grew past
roughly 70 KB every snapshot - and therefore every bearings report - failed with
"Argument list too long". Shortening the other arguments could not help because
only one argument's size is measured.

Add jq_json, which wraps the growable values in one object fed to jq on stdin
and binds each back to its original name, so filter bodies are unchanged. Route
every call site whose value scales with backlog, task, secondmate, or report
volume through it, including the recursive per-home summary, the secondmate
record accumulator, and the final document assembly. Fixed-size values such as
booleans, counts, one path, and one status line stay on argv.

Cover the regression with an oversized-backlog fixture that asserts the full
inventory survives, not merely that the command exits 0.
…undary

Replace the paired-review default-routing guidance with a deterministic
navigator rule: every paired navigator runs on pi with the model
cx/gpt-5.6-terra, and only an explicit per-task captain choice of another
navigator runtime or model replaces it. Ordinary dispatch resolution keeps
owning the concrete launch mechanics and carries the pin through instead of
substituting a best-fit alternative.

State the ownership boundary in the same file: paired review is a universal
implementation-task execution protocol owned by whichever firstmate owns the
task, so a standalone paired task needs no persistent orchestrator, while a
program orchestrator dispatches the same protocol unchanged.

Drop the reroute-to-claude clause under the handed-out-skill limitation, which
would otherwise silently replace the pinned navigator runtime.
A planner session is one temporary worker the captain talks to directly.
Firstmate opens it, points it at a captain-defined project, scope, and
planning question, then leaves the conversation and monitors only whether
the session is alive, waiting, finished, or failed.

The planner investigates before it asks anything, reconstructing both what
the code does today and where the project's accepted decisions say it is
going, and surfaces every conflict between them instead of quietly choosing.
It reads tests as evidence and runs nothing. It then grills the captain one
question at a time, each question carrying its recommended answer and each
prescriptive choice carrying both the simplest workable and the best-practice
case. It may not produce a plan until the captain opens the crystallize gate,
after which it publishes a spec or dependency-aware tickets through the
project's own established conventions. Planning never authorizes
implementation.

Two artifact contracts travel downstream, each with one owner:

- The scope envelope carries the accepted boundary at spec level and per
  ticket, reusing bin/fm-brief.sh's existing scope and seam vocabulary plus
  the two fields a brief cannot carry before the target is chosen.
- The test contract carries acceptance intent, not executable test code, and
  the captain approves it together with scope.

planner owns those artifact fields. program-orchestration owns consuming
them: preserve provenance, revalidate, narrow per worker, never widen,
escalate a stale envelope. bin/fm-brief.sh keeps the final worker statement
and paired-review keeps sharing it with the driver and navigator, plus the
navigator's plan-gate challenge to the approved test contract.

The planning disciplines are vendored from mattpocock/skills at 2ab9580 with
the MIT license and a provenance record, adapted only in paths, tracker
assumptions, and delegation wording. No bundled file is named SKILL.md, so
the bundle stays off every harness's skill index and costs no context; it
also cannot silently clobber a captain's own locally installed copy, which
git overwrites without a conflict on fast-forward when the install path is
locally excluded.

Launch reuses the ordinary brief, spawn, and scout-report mechanics with no
new runtime: pi cx/gpt-5.6-sol at high effort by default, claude
claude-opus-5 when explicitly selected.

Also fixes fm-doc-audience-check.sh treating markdown links inside fenced
code blocks as navigation targets, which a vendored example listing exposed.
…tory

A navigator could only assert findings, so a question had to be dressed as a
defect or stay silent, and nothing systematically told it what had already been
decided. One program's navigator invented an undeclared Q entry to work around
exactly this, and another raised a cross-ticket scope question as a finding,
forcing the driver to adjudicate something it had no authority over.

paired-review gains an evidence-backed Q<k> form beside N<k>, on the same one
round trip but with no forced accepted/rejected verdict. The owning firstmate
answers by default and answers it itself, escalating upward only when it lacks
the authority or the knowledge; the driver answers only its own implementation
choices. Navigator briefs carry the decisions bearing on the ticket where a
record exists, and the skill now states that cross-package and whole-solution
direction sits outside one navigator's field of view.

program-orchestration takes the matching duties: maintaining the program's
cross-ticket decision record, handing the bearing decisions to both halves of
every dispatched pair, and holding direction across tickets.

N semantics and the four structural questions are unchanged, and no fifth
question was added. Both briefs are written at dispatch, so this reaches the
next dispatch and leaves a pair already under way on the protocol it launched
with.
The Q form said a question runs "its own sequence" but never said that
sequence is sequential and never reused, so the identifier rule was
checkable for N and not for Q. Match the N wording.
fm_herdr_lab_raw appended `--session <name>` to the end of the herdr argv.
`--session` is a herdr global option, and `herdr agent start <name> ... -- <argv...>`
hands every token after `--` to the inner command, so for any subcommand with a
separator the flag never reached herdr: the call silently ran against whatever
session herdr resolved by default, and the inner command received two junk
arguments. The reported repro, `agent start envcheck -- bash`, executed as
`bash --session fm-lab-...` and the pane died immediately.

The confirmed hypothesis is argument placement, not content. Placement was
inherited from fm_backend_herdr_cli, whose comment records that leading position
was already verified to work and that trailing was chosen only to keep call-site
diffs append-only; the blind spot in that choice is the `--` separator.

Every rejection rule is unaffected: the caller-supplied --session scan, the
leading-option refusal, the lifecycle and server refusals, and the fresh
refuse-default checks all run in fm_herdr_lab_cli on the caller's arguments
before fm_herdr_lab_raw assembles the herdr argv.

The existing behavior test asserted the trailing position as the contract, so the
suite was green on the defect; its assertions are re-pointed, and a new case
records the herdr argv token by token and asserts the session pair leads and
never lands past a `--`, with the inner command passed through untouched.

The identical latent pattern in fm_backend_herdr_cli is deliberately out of scope
here; no current backend call site uses a `--` separator.
Open a long-running programme orchestrator the captain talks to directly,
driving an already-authorized spec, ticket set, or GitHub issues to
completion through workers the session dispatches itself.

The skill splits across the branch boundary it actually has: firstmate
reads SKILL.md to open the session, and only the session reads
CONTRACT.md, so the judgment written for the orchestrator costs
firstmate a pointer rather than a context load. program-orchestration
keeps ownership of custody, routing, host ramp, envelope consumption,
and handoff; both files point at it rather than restating it.

AGENTS.md gains the third direct-conversation exception and one
deliverable-classification bullet. The planner test's closed-exception
assertion moves to the new count rather than dropping the invariant.
* fix(bin): keep tracked Claude hook entries inert under grok 1.0.0 (kunchenguid#1917)

* fix(hooks): keep tracked Claude entries inert under grok 1.0.0 hooks

Grok loads Claude-compatible settings, so the tracked `.claude/settings.json`
hook entries also fire under Grok. They were meant to be inert there, guarded
by `[ -z "${GROK_AGENT:-}" ] || exit 0`. That guard silently stopped working.

Verified from the live process environment of a wedged grok 1.0.0 Stop hook on
2026-08-07: a grok 1.0.0 HOOK process carries GROK_HOOK_EVENT, GROK_HOOK_NAME,
GROK_SESSION_ID, and GROK_WORKSPACE_ROOT, but no GROK_AGENT. The observed hook
process was labelled `GROK_HOOK_NAME=project/settings:stop[0].hooks[1]`, which
is the Claude-only auto-arm entry.

Consequence: Grok ran `bin/fm-claude-stop-autoarm.sh` synchronously. Grok has
no `asyncRewake`, so it waited on the foregrounded watcher for that entry's
declared 28800-second timeout and the Grok turn never ended - the operator saw
an infinite "Responding".

Widen the guard to `[ -z "${GROK_AGENT:-}${GROK_HOOK_EVENT:-}" ] || exit 0` on
the five entries that have a `.grok/hooks/` counterpart: both Stop entries, the
SessionStart entry, and the two PreToolUse Bash entries.

Two deliberate limits:

- The guard is NOT widened to GROK_SESSION_ID. Grok injects it into every child
  process, so it can survive into a Claude session that Grok launched and would
  silently disable Claude's own watcher continuity. GROK_HOOK_EVENT is
  per-hook-invocation and does not leak that way.
- `bin/fm-subagent-pretool-check.sh` stays unguarded on purpose. It is the one
  tracked entry with no `.grok/hooks/` counterpart, so guarding it would remove
  the guard from Grok entirely rather than deduplicate it. The new test asserts
  it stays unguarded so the exception cannot be closed silently, and
  docs/subagent-guard.md is honest that the coverage it leaves is partial.

`bin/fm-harness.sh` corrects a comment that presented GROK_AGENT as reliably
present; it is a fast path only, and the ancestry walk is what actually
guarantees grok identification.

tests/fm-turnend-guard.test.sh adds test_tracked_claude_entries_inert_under_grok,
which runs every tracked entry under a real grok 1.0.0 hook environment, a
legacy GROK_AGENT environment, and a native Claude environment.

* no-mistakes(document): docs: sync grok hook-marker guard facts to owners

* no-mistakes(review): docs: state grok guard criterion by event coverage

* feat(startup-network): record per-step elapsed times for the deferred stage (kunchenguid#1918)

The deferred network stage published one aggregate started/finished pair, so
a run that took a minute could not be attributed to a phase, a host, or a
clone without re-running it by hand under manual tracing.

Add bin/fm-timing-lib.sh as the single owner of elapsed-time records, and
bracket each network owner with one: the gh auth probe, the secondmate
liveness sweep, secondmate convergence, pending handoff delivery, and the
project clone refresh, plus one record per secondmate for the remote-touching
steps (id and host) and one per project clone. Each record carries a start
offset from one shared origin, so the artifact reads as a timeline.

The stage publishes them beside its report as state/.startup-network.timings,
for a timed-out or failed run too, where the partial record is the answer.
Only the on-demand `report` command prints them: `harvest` composes the
session-start digest, so its output, the wake cadence, and every other part
of a normal session start are unchanged.

Recording is inert unless a run asks for it, so nothing else that sources
these scripts pays for it. Details are identities only - a detail carrying
whitespace is refused rather than cleaned up, which is what keeps a command
line, an environment dump, or a captured error out of the file.

Split two per-item loop bodies into their own functions so each iteration can
be timed; every `continue` became a `return 0` with the same meaning, and the
sweeps still run directly, in the same order, returning the same results.

* feat(stow): cascade the internal /stow to every registered secondmate (kunchenguid#1928)

* feat(stow): cascade the internal /stow to every registered secondmate

Invoked in a primary home, /stow now sweeps every registered secondmate
after the primary's own required pass, enforcing the same startup-memory
threshold in each home against that home's own allowance rather than a
fleet total.

bin/fm-stow-cascade.sh owns the mechanical inputs: it enumerates each
registered secondmate exactly once from data/secondmates.md, reports that
home's own budget accounting, and resolves how the sweep reaches it. A
live agent sweeps its own home so its uncaptured session knowledge is
captured too; a local home without one is curated in place; a remote home
without one is accounted read-only and deferred, because there is no
generic remote write path for a home's own memory files. Every host-
crossing step and each home's accounting runs under one hard bound, so a
slow or unreachable home reports an exception and the sweep continues.

Nothing changes until /stow is invoked: no new notification, digest
section, or background work. The public skills/stow skill is untouched.

* no-mistakes(review): fix(stow): extend cascade --help range to include full exit-code contract

* fix(remote-job): stop workers abandoned by a pruned code root (kunchenguid#1927)

29 fm-remote-job-worker.sh processes were found running at ppid 1, 1-2 days
old, each still polling and appending to a log inside a no-mistakes gate
worktree that had already been returned.

Three things combined to make that possible:

- The recorded worker.pid is the serving child, not the restart supervisor
  above it, so a teardown that stops that one pid only makes the supervisor
  respawn. The Linux start path also left the worker tree in the launching
  command's process group, so there was no group to signal instead.
- Neither the serving loop nor the supervisor ever rechecked whether its
  configured FM_ROOT still existed, so a worker launched from a worktree
  outlived that worktree indefinitely.
- The supervisor restarted a failing child with a fixed 0.1s delay and no
  bound, which is what grew the logs (~66MB/day measured).

The Linux start path now puts the worker tree in its own process group, and
fm_remote_job_stop_worker_tree signals that whole group - refusing any group
whose leader is not itself a worker, so a worker from an older build or from
launchd's own session is still stopped safely as a single process. The worker
stops itself once its code root stops being a Firstmate checkout, confirmed
across a grace window so an ordinary transient cannot stop a healthy worker.
The supervisor backs off and gives up rather than restarting forever.

bin/fm-remote-job-reap-orphans.sh is the belt-and-suspenders sweep for workers
already orphaned that way, wired into fm-teardown.sh. Its reap condition is
exactly "the code root named in the worker's own command line is gone", which
is why the account's healthy LaunchAgent worker and every live remote
secondmate worker are never candidates.

The two suites that leaked these in the first place now stop the worker tree
rather than the recorded pid alone.

* feat(bin): lint only the changed shard locally, full lint in CI (kunchenguid#1925)

* fix(bin): lint only the changed shard locally, full lint in CI

Two ships hitting fm-lint.sh at once could spike CPU to 190% and load
to 8.58 on a captain's Mac, even though each run finishes quickly.
fm-lint.sh now defaults to linting only the canonical-set files
changed since the merge-base with origin/main (including uncommitted
edits) on an ordinary local branch, using plain local git with no
network calls. It still lints the full canonical set in CI
(GITHUB_ACTIONS=true or CI=true), on the main branch, or whenever no
merge-base can be found, so CI coverage never depends on a local diff.
Explicit paths keep bypassing this selection entirely.

* no-mistakes: apply CI fixes

* fix(herdr): admit agy plain-ASCII '> ' composer via native identity

agy's interactive composer draws a bare ASCII "> " prompt with no side
border and no agent-specific glyph, so its empty-composer row is
byte-for-byte the dead-shell prompt the shared classifier in
bin/fm-composer-lib.sh deliberately refuses as unknown (never empty).
Under herdr that made a live-empty agy composer fall to unknown, unlike
bordered/agent-glyph shapes, so away-mode injection never fired.

Close the gap inside the herdr adapter only, mirroring the Pi separator
composer: locate the bottom-most "> " row structurally and admit it as
a real composer solely when native `agent get` corroborates the target
is agy and reports idle/done/blocked. A working agy, a non-agy or
unregistered agent, or an unreadable identity keeps the bare row unknown,
so the shared dead-shell rule is not weakened. Identity is fetched at
most once per call, shared with the Pi block.

* feat: add agy (Antigravity CLI) harness adapter

Wire Google Antigravity CLI (agy 1.1.5, Gemini-backed) as a crewmate harness.

- fm-harness.sh: detect agy via ANTIGRAVITY_AGENT=1 env marker plus
  agy/antigravity process-ancestry match.
- fm-spawn.sh: interactive launch template
  (agy --dangerously-skip-permissions -i "$(cat brief)"); add agy to the
  model flag (--model), effort flag (--effort, ceiling high), and the
  accepted-harness lists.
- fm-tmux-lib.sh / fm-watch.sh: add busy signature "esc to cancel" to the
  default busy regex (daemon inherits it).
- harness-adapters SKILL.md: full agy knowledge section, launch-profile-axes
  row, and no-mistakes invocation note.
- fm-session-start.test.sh: drop ANTIGRAVITY_AGENT in detect_own so the
  suite does not leak the marker.

Verified empirically on agy 1.1.5: env marker, interactive cwd inheritance,
--dangerously-skip-permissions autonomy, busy/idle footer, first-run trust
dialog. Turn-end hook and primary-session guard are NOT yet verified, so agy
is crewmate-only for now (documented in the skill).

* fix: read live pi sessions as live in session-lock liveness

fm_harness_pid_alive joined `basename comm` and `ps -o args=` into one
line before matching FM_HARNESS_RE. That broke every anchored
alternative: a live pi harness joins to "pi pi", which `^pi$` cannot
match, so each live pi session read as stale. bin/fm-lock.sh then let a
second session take the primary's lock, and bin/fm-claude-stop-autoarm.sh
reclaimed a lock whose owner was alive - a whole-home mutual-exclusion
failure that can double-dispatch or double-teardown.

Split the check the way fm_harness_ancestry_pid already does it: match
the basename of comm on its own line, and consult args only when comm is
a bare interpreter (node/python). FM_HARNESS_RE is unchanged, so `^pi$`
stays anchored and pip/pipx/paths containing "pi" keep reading as
not-a-harness.

tests/fm-session-lock-liveness.test.sh covers both directions against
real spawned processes (no fake ps) and fails on the pre-fix library.

* docs: record the gh-axi repo-resolution trap and the turn-boundary stall pattern

* fix: verify pruned resolved captain holds from the done archive

tasks-axi's Done-retention prune archives resolved captain holds out of the
live backlog (closing the originating task auto-prunes), and the completion
gate could then never verify them again: teardown refused, complete failed
with 'absent from data/backlog.md', and the task deadlocked with no
non-bypass exit (hit live with psak-compose-bringup on 2026-07-26).

Prune archives rather than deletes, so the durable resolution record
survives byte-identically in the markdown backend's archive file. Every
read that accepts a Done hold now falls back to that archive: complete,
verify, identical resolve retries, and reopen refusal are order-independent
with respect to pruning. An archived row without the resolution record
still fails the gate, and an active hold is never read from the archive,
so the safety boundary is unchanged - the gate verifies the same evidence
it always required, from where the backend moved it.

Two regression tests reproduce the deadlock sequence end-to-end and prove
the fail-closed path for an archived hold with no resolution record.

* docs: add the paired driver-and-navigator protocol as an agent-only skill

High-blast-radius ship work now runs as a pair launched together, exchanging
directly through a durable shared file rather than through firstmate, which was
the latency bottleneck when it supervises several threads at once. The protocol
body lives in the new agent-only `paired-review` skill; AGENTS.md gains only the
trigger, the narrowed communication rule, and the delivery-path exception.

The second worker is a navigator, not a reviewer: three of the four gates have
no diff to look at, and an agent told it reviews will idle until one exists,
which defeats the design. The plan gate is the cheap one - a wrong-layer
direction error is visible from a file-level plan in about two minutes and
invisible at PR time under hours of work built on it.

The pair may never settle a contract or scope change, anything destructive,
irreversible, or security-sensitive, or a disagreement surviving its one round
trip; removing firstmate from the exchange moves no approval authority to the
crewmates.

Also make the scope and seam statement a default section of every ship and
scout brief, since the direction error that motivated this protocol happened on
a brief that never named the module owning the code.

* docs: exclude diagnosis from pairing and give navigators four structural questions

Three rules the paired-review protocol did not carry, plus the rule that
keeps skill changes written against the skill-writing guide.

A diagnosis is not paired work. Pairing applies to implementation work; a
bug, regression, or crash-loop takes the diagnose-then-fix shape, one
worker following the diagnosing-bugs skill by absolute path. Until a root
cause exists there is no "where" for a navigator to hold a view about,
and the risk a bug carries is a wrong root cause rather than a wrong
layer, so sequence guards it where parallelism cannot.

The root-cause gate: firstmate judges the root cause before any fix
action. Confident, it authorises the fix and reports; otherwise it
escalates and waits. The root-cause document's own standard is the
confidence test. Destructive, irreversible, security-sensitive, and
contract-expanding fixes still go to the captain regardless.

Every navigator brief carries four fixed structural questions alongside
its task-specific checks: was the stated reason delivered, where did the
coupling go, does this add another instance of a shape the repo has been
burned by, and what did the brief not specify that the implementation had
to decide anyway. Each keeps its reason, because a checklist cannot reach
the destination - a navigator is pointed at a diff, and a diff shows what
changed, never whether the change moves toward where the work is going.

firstmate-coding-guidelines gains one requirement: a skill change begins
by loading writing-great-skills and is written against it.

AGENTS.md changes are two edited trigger lines and no new lines; the
rules themselves live in the skills.

* fix: keep whole-fleet JSON off the jq command line in the fleet snapshot

The snapshot passed the parsed backlog and task inventory to jq as --argjson
argv strings. A single argument is capped by MAX_ARG_STRLEN (128 KiB), a
separate and far smaller limit than ARG_MAX, so once a home's backlog grew past
roughly 70 KB every snapshot - and therefore every bearings report - failed with
"Argument list too long". Shortening the other arguments could not help because
only one argument's size is measured.

Add jq_json, which wraps the growable values in one object fed to jq on stdin
and binds each back to its original name, so filter bodies are unchanged. Route
every call site whose value scales with backlog, task, secondmate, or report
volume through it, including the recursive per-home summary, the secondmate
record accumulator, and the final document assembly. Fixed-size values such as
booleans, counts, one path, and one status line stay on argv.

Cover the regression with an oversized-backlog fixture that asserts the full
inventory survives, not merely that the command exits 0.

* docs: add program orchestration skill

* docs: pin the paired navigator and state paired review's ownership boundary

Replace the paired-review default-routing guidance with a deterministic
navigator rule: every paired navigator runs on pi with the model
cx/gpt-5.6-terra, and only an explicit per-task captain choice of another
navigator runtime or model replaces it. Ordinary dispatch resolution keeps
owning the concrete launch mechanics and carries the pin through instead of
substituting a best-fit alternative.

State the ownership boundary in the same file: paired review is a universal
implementation-task execution protocol owned by whichever firstmate owns the
task, so a standalone paired task needs no persistent orchestrator, while a
program orchestrator dispatches the same protocol unchanged.

Drop the reroute-to-claude clause under the handed-out-skill limitation, which
would otherwise silently replace the pinned navigator runtime.

* feat(skills): add the captain-invoked planner session

A planner session is one temporary worker the captain talks to directly.
Firstmate opens it, points it at a captain-defined project, scope, and
planning question, then leaves the conversation and monitors only whether
the session is alive, waiting, finished, or failed.

The planner investigates before it asks anything, reconstructing both what
the code does today and where the project's accepted decisions say it is
going, and surfaces every conflict between them instead of quietly choosing.
It reads tests as evidence and runs nothing. It then grills the captain one
question at a time, each question carrying its recommended answer and each
prescriptive choice carrying both the simplest workable and the best-practice
case. It may not produce a plan until the captain opens the crystallize gate,
after which it publishes a spec or dependency-aware tickets through the
project's own established conventions. Planning never authorizes
implementation.

Two artifact contracts travel downstream, each with one owner:

- The scope envelope carries the accepted boundary at spec level and per
  ticket, reusing bin/fm-brief.sh's existing scope and seam vocabulary plus
  the two fields a brief cannot carry before the target is chosen.
- The test contract carries acceptance intent, not executable test code, and
  the captain approves it together with scope.

planner owns those artifact fields. program-orchestration owns consuming
them: preserve provenance, revalidate, narrow per worker, never widen,
escalate a stale envelope. bin/fm-brief.sh keeps the final worker statement
and paired-review keeps sharing it with the driver and navigator, plus the
navigator's plan-gate challenge to the approved test contract.

The planning disciplines are vendored from mattpocock/skills at 2ab9580 with
the MIT license and a provenance record, adapted only in paths, tracker
assumptions, and delegation wording. No bundled file is named SKILL.md, so
the bundle stays off every harness's skill index and costs no context; it
also cannot silently clobber a captain's own locally installed copy, which
git overwrites without a conflict on fast-forward when the install path is
locally excluded.

Launch reuses the ordinary brief, spawn, and scout-report mechanics with no
new runtime: pi cx/gpt-5.6-sol at high effort by default, claude
claude-opus-5 when explicitly selected.

Also fixes fm-doc-audience-check.sh treating markdown links inside fenced
code blocks as navigation targets, which a vendored example listing exposed.

* feat(skills): give the navigator a question form and the decision history

A navigator could only assert findings, so a question had to be dressed as a
defect or stay silent, and nothing systematically told it what had already been
decided. One program's navigator invented an undeclared Q entry to work around
exactly this, and another raised a cross-ticket scope question as a finding,
forcing the driver to adjudicate something it had no authority over.

paired-review gains an evidence-backed Q<k> form beside N<k>, on the same one
round trip but with no forced accepted/rejected verdict. The owning firstmate
answers by default and answers it itself, escalating upward only when it lacks
the authority or the knowledge; the driver answers only its own implementation
choices. Navigator briefs carry the decisions bearing on the ticket where a
record exists, and the skill now states that cross-package and whole-solution
direction sits outside one navigator's field of view.

program-orchestration takes the matching duties: maintaining the program's
cross-ticket decision record, handing the bearing decisions to both halves of
every dispatched pair, and holding direction across tickets.

N semantics and the four structural questions are unchanged, and no fifth
question was added. Both briefs are written at dispatch, so this reaches the
next dispatch and leaves a pair already under way on the protocol it launched
with.

* fix(skills): state the Q id discipline the N form already carries

The Q form said a question runs "its own sequence" but never said that
sequence is sequential and never reused, so the identifier rule was
checkable for N and not for Q. Match the N wording.

* fix(herdr): place the lab --session ahead of any -- separator

fm_herdr_lab_raw appended `--session <name>` to the end of the herdr argv.
`--session` is a herdr global option, and `herdr agent start <name> ... -- <argv...>`
hands every token after `--` to the inner command, so for any subcommand with a
separator the flag never reached herdr: the call silently ran against whatever
session herdr resolved by default, and the inner command received two junk
arguments. The reported repro, `agent start envcheck -- bash`, executed as
`bash --session fm-lab-...` and the pane died immediately.

The confirmed hypothesis is argument placement, not content. Placement was
inherited from fm_backend_herdr_cli, whose comment records that leading position
was already verified to work and that trailing was chosen only to keep call-site
diffs append-only; the blind spot in that choice is the `--` separator.

Every rejection rule is unaffected: the caller-supplied --session scan, the
leading-option refusal, the lifecycle and server refusals, and the fresh
refuse-default checks all run in fm_herdr_lab_cli on the caller's arguments
before fm_herdr_lab_raw assembles the herdr argv.

The existing behavior test asserted the trailing position as the contract, so the
suite was green on the defect; its assertions are re-pointed, and a new case
records the herdr argv token by token and asserts the session pair leads and
never lands past a `--`, with the inner command passed through untouched.

The identical latent pattern in fm_backend_herdr_cli is deliberately out of scope
here; no current backend call site uses a `--` separator.

* feat(skills): add the captain-invoked orchestrator session

Open a long-running programme orchestrator the captain talks to directly,
driving an already-authorized spec, ticket set, or GitHub issues to
completion through workers the session dispatches itself.

The skill splits across the branch boundary it actually has: firstmate
reads SKILL.md to open the session, and only the session reads
CONTRACT.md, so the judgment written for the orchestrator costs
firstmate a pointer rather than a context load. program-orchestration
keeps ownership of custody, routing, host ramp, envelope consumption,
and handoff; both files point at it rather than restating it.

AGENTS.md gains the third direct-conversation exception and one
deliverable-classification bullet. The planner test's closed-exception
assertion moves to the new count rather than dropping the invariant.

* fix: resolve inherited test-run conflict marker

* ci: remove inherited no-mistakes requirement

* fix: restore scoped reconciliation regressions

* ci: update snapshot compatibility count

---------

Co-authored-by: Jay Park <jay.jongcheol.park@gmail.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…ncestry

Restore accepted reconciliation ancestry
* vendor Matt productivity and engineering skills snapshot

* classify vendored Matt documentation surfaces
* feat(policy): add planner policy launcher and focused test

* fix(policy): keep recommendation path singular by default

* refactor(planner): remove stale planner surfaces

* docs(agents): retain only policy pointer

* fix(agents): restore upstream surface for policy pointer

* test(planner): accept static instruction contract
Move the paired implement-and-check protocol to custom-skills/paired-review as
the sole custom owner, adapted from the upstream-path owner with the accepted
hardening: Herdr pair composition in adjacent panes with direct communication
supplementing the canonical pair-log, navigator immediate-stop, the safe-boundary
hold, explicit no-shared-Git and no-new-transport limits, and a bounded
cross-owner pointer to /orchestrator and program-orchestration. Delete the old
.agents/skills/paired-review owner so no stale duplicate remains. Add first-match
role=driver and role=navigator trigger rows to custom-skills/policy before the
generic implement and review routes, move the agent-runtime inventory entry to
the custom owner path, and add a colocated focused test proving the policy route
and path resolution including the .claude/skills symlink. No core scripts changed.
* chore: remove stale orchestrator test from test runner

* fix: revert firstmate-coding-guidelines to upstream
huynhtandat223 and others added 9 commits August 9, 2026 00:59
The send subcommand injected messages into the recipient Herdr composer,
whose success response only proves text placement, not a submitted turn.
Resolve the role task id from recovery.json and submit through fm-send,
the verified submission owner, so an unconfirmed send fails loudly without
retry, Herdr injection, or fallback message. Tests cover successful and
failed submission through the public interface.
…se send (#30)

The single live-signal API for paired roles is fm-pair-compose.sh send
<recovery.json> <driver|navigator> "<signal>", which resolves the
recipient task from recovery evidence and submits through fm-send.sh so
the signal is Enter-verified. Role instructions now require this method
for PLAN READY, MILESTONE, STOP, ACK STOP, and direct pointers, and no
longer direct a worker to a Herdr agent target. Pair bootstrap PAIR READY
uses the same verified path instead of raw herdr agent send; Herdr keeps
topology/pane placement only. Focused tests prove role instructions name
the method and runtime delivery never uses raw herdr agent send or pane
send-text.

Co-authored-by: firstmate-worker <firstmate@localhost>
… a secondmate (#32)

`/orchestrator` opened its programme session as an ordinary scout, which cannot
supervise children, and the surrounding docs described it as a persistent
secondmate it never was. Issue #31 accepts a third shape: one bounded temporary
supervisor with a firstmate home of its own, which dispatches and supervises
ordinary workers itself and closes when its programme ends.

The whole lifecycle lives in one new file, bin/fm-supervisor-lib.sh, so the
footprint in the rest of bin/ is a set of one-line delegations that
`grep -rn fm_supervisor bin/` lists in full - the seam can be removed by
deleting that file and those call sites, or carried across an upstream upgrade
by re-applying only the call sites.

Runtime:
- bin/fm-supervisor-lib.sh: the lease, home identity, isolation assertions,
  relaunch reuse, and the self-supervising kind set.
- bin/fm-spawn.sh --supervisor: leases an isolated firstmate worktree with
  `treehouse get --lease`, marks it, records home= and parent_home=, and
  launches the session in it. No project is cloned there; the parent home's
  existing clones stay read-only allocation sources. A relaunch returns to the
  recorded home rather than leasing a second one, so a stopped supervisor comes
  back to its own worker inventory and cannot re-dispatch duplicates. A failed
  launch gives the fresh lease back. No secondmate registry, charter, inherited
  material, or liveness sweep is involved.
- bin/fm-primary-scope-lib.sh: the .fm-supervisor-home marker puts that home in
  scope for its own hooks, alongside the existing secondmate marker.
- bin/fm-watch.sh, bin/fm-crew-state.sh: a report that runs its own firstmate
  session is read from its status writes, not its pane, so an idle supervisor
  is healthy rather than stale.
- bin/fm-teardown.sh: cleanup refuses until the child records, the home's own
  unlanded work, the unresolved-decision gate, and the programme report are all
  reconciled, then retires the home and releases its lease.

Policy stays in custom-skills/: the orchestrator skill now opens, relaunches,
and closes a temporary supervisor; program-orchestration no longer routes
through secondmate provisioning; and fm-supervisor-brief.sh scaffolds the
session's contract there rather than in core.

AGENTS.md, .gitignore, and bin/fm-brief.sh are deliberately untouched: the
marker is excluded per worktree, the brief is programme policy, and the skill
is reachable through the policy router AGENTS.md already points at.

Verification: tests/fm-temporary-supervisor.test.sh (13 cases: home isolation,
no project cloning, primary scope, child dispatch, same-home relaunch,
duplicate prevention, every cleanup refusal, lease release),
custom-skills/tests/fm-orchestrator-policy.test.sh (7 cases), plus the
secondmate, watcher-wake-lock, and backend-dispatch families and bin/fm-lint.sh
clean.

Closes #31
…pane ids (#34)

Herdr assigns a pane a NEW pane id on every move, so the pane id recorded at
launch is a hint that stops naming any pane once the pair is arranged. The
coordinator carried that recorded id through its own workspace and split moves
and then asserted adjacency against it.

`pane neighbor` compounded it: Herdr reports the queried pane in `pane_id` and
the adjacent pane in `neighbor_pane_id`, and the reciprocal check read the
former. It therefore compared each pane against itself and refused with
`reciprocal adjacency` for every pair, including one whose roles were live,
correctly labelled, and adjacent.

Both composition and the new `recover` subcommand now resolve each role's
current pane, workspace, and tab from Herdr's own agent inventory, keyed by the
stable role agent name and proven against that role's isolated copy. A topology
generation is recorded only once both roles are proven distinct, in one tab, and
reciprocal neighbours; a role identity that is missing, duplicated, or bound to
another copy refuses, records no generation, and never releases the pair.

fm-send remains the only pair delivery path; Herdr is still read for topology
only and is never used to inject text.

The fake Herdr in the behavior tests encoded the same wrong response shapes as
the helper, which is why the defect passed CI. It now models the real CLI - pane
ids reassigned on move, `neighbor_pane_id` distinct from the queried pane, agent
names as the stable identity - and covers stale launch ids, a moved pair
recovered into a new generation, recovery idempotence, and refusal on missing,
duplicated, copy-mismatched, and split-tab identities.
…up error (#35)

Follow-up accuracy pass on #34. The behaviour is unchanged; the shipped
explanation of the mechanism was imprecise and the regression pinned the wrong
failure.

Herdr does not stop resolving a superseded pane id. It accepts the old id as a
lookup key and answers with the pane's CURRENT id, so a recorded id keeps
working for lookups and silently stops matching any answer compared against it.
The header and inline comments said the id stops naming a pane, which reads as a
lookup failure and points a future reader at the wrong class of defect.

The fake Herdr had the same gap: a superseded id resolved to nothing, so the
pre-fix helper died at `pane get` instead of reaching the refusal that was
actually reported. It now keeps each pane's former ids and resolves them to the
current pane, and the launch-time drift is produced by a real move rather than a
synthetic unknown id. Against the pre-fix helper the suite now fails with
`reciprocal adjacency` - the reported symptom.

The superseded-id case also asserts its own divergence: the recorded id must
still resolve while no longer being the role's current pane, so the case cannot
go quietly vacuous if the fake's aliasing is ever dropped.

Verified read-only against the running Herdr and the live pair whose
composition failed:

  $ herdr pane neighbor --direction right --pane w4B:p2X   # superseded id
  {"neighbor":{"pane_id":"w5J:p1","neighbor_pane_id":"w5J:p2",...}}
Claude 2.1.226 pads its EMPTY composer row with U+00A0 after the `❯` prompt
glyph. The blank is drawn at luminance 153, so ghost stripping correctly keeps
it, and every trim in the shared classifier and its callers is ASCII-only - so
a genuinely empty Claude composer classified `pending`, as if it still held
unsubmitted text.

Nothing failed loudly, because the composer verdict only decides delivery where
no stronger signal exists. An ordinary steer to an idle worker is confirmed
from Herdr's native agent state and never reads the composer at all. The one
path with no stronger signal is Herdr's submit confirmation for a target that
was ALREADY mid-turn before Enter, and that is exactly the path a paired-review
barrier release takes, because both roles sit in a blocking wait when the
release is sent. fm-send reported `delivery unconfirmed; verdict=pending` for a
message the driver had already received, and the composer correctly refused to
release the navigator, leaving the pair without coverage before its plan gate.

Fold non-ASCII Unicode blanks to ASCII space in fm_composer_classify_content,
the one fleet-wide owner every adapter delegates to. This can only move content
that is ENTIRELY blank from `pending` to `empty`; content carrying any visible
character keeps its verdict, so a genuinely swallowed Enter still refuses and
the away-mode injector still declines to type over unsent input. Keying on the
Unicode space-separator category rather than on U+00A0 alone keeps the fix from
depending on which blank a given release happens to pick.

Verified against real claude 2.1.226 on herdr 0.7.3 through the production
adapter, both directions: a message accepted by a mid-turn worker now confirms
(and was observably delivered), while text typed but never submitted still
reports pending.

Regressions:
- tests/fm-composer-lib.test.sh pins the fold, every declared blank, and the
  divergence - padded-empty is empty while padded-real-text stays pending, so
  the case cannot go vacuous - plus the unchanged dead-shell refusal.
- tests/fm-backend-herdr.test.sh pins the busy-baseline submit path this fault
  actually reached, in both directions, on captures carrying the real bytes.
- tests/fm-composer-blank-drift-live-e2e.test.sh is the opt-in real-harness
  guard, since what a harness draws in an empty composer is a vendor surface a
  fixture can only replay from a previous release. It records the untouched
  verdict, then asserts it only after probe text typed into that pane becomes
  visible, so a blank startup screen or a trust dialog is reported unverified
  instead of passing. It fails on claude without this fix and passes with it.

Compatibility: the shared owner is consumed by the tmux, herdr, cmux and orca
adapters; all four suites pass, and zellij does not use it.
@huynhtandat223

Copy link
Copy Markdown
Author

Closing this accidentally opened upstream PR. The work must be delivered through the huynhtandat223/firstmate fork PR path.

@huynhtandat223

Copy link
Copy Markdown
Author

Closing: this was opened against the wrong repository. The work belongs on my fork and is being raised there instead. Apologies for the noise.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant