Skip to content

Rebase custom Firstmate capabilities onto upstream main - #3

Merged
huynhtandat223 merged 15 commits into
mainfrom
fm/fm-upstream-rebase-82-upgrade-only
Aug 7, 2026
Merged

huynhtandat223 merged 15 commits into
mainfrom
fm/fm-upstream-rebase-82-upgrade-only

Conversation

@huynhtandat223

Copy link
Copy Markdown
Owner

Rebases the retained captain custom capabilities directly onto current upstream/main.\n\nValidated locally: lint, documentation audience check, session-lock liveness, fleet snapshot, planner, and orchestrator focused suites.

agy's interactive composer draws a bare ASCII "> " prompt with no side
border and no agent-specific glyph, so its empty-composer row is
byte-for-byte the dead-shell prompt the shared classifier in
bin/fm-composer-lib.sh deliberately refuses as unknown (never empty).
Under herdr that made a live-empty agy composer fall to unknown, unlike
bordered/agent-glyph shapes, so away-mode injection never fired.

Close the gap inside the herdr adapter only, mirroring the Pi separator
composer: locate the bottom-most "> " row structurally and admit it as
a real composer solely when native `agent get` corroborates the target
is agy and reports idle/done/blocked. A working agy, a non-agy or
unregistered agent, or an unreadable identity keeps the bare row unknown,
so the shared dead-shell rule is not weakened. Identity is fetched at
most once per call, shared with the Pi block.
Wire Google Antigravity CLI (agy 1.1.5, Gemini-backed) as a crewmate harness.

- fm-harness.sh: detect agy via ANTIGRAVITY_AGENT=1 env marker plus
  agy/antigravity process-ancestry match.
- fm-spawn.sh: interactive launch template
  (agy --dangerously-skip-permissions -i "$(cat brief)"); add agy to the
  model flag (--model), effort flag (--effort, ceiling high), and the
  accepted-harness lists.
- fm-tmux-lib.sh / fm-watch.sh: add busy signature "esc to cancel" to the
  default busy regex (daemon inherits it).
- harness-adapters SKILL.md: full agy knowledge section, launch-profile-axes
  row, and no-mistakes invocation note.
- fm-session-start.test.sh: drop ANTIGRAVITY_AGENT in detect_own so the
  suite does not leak the marker.

Verified empirically on agy 1.1.5: env marker, interactive cwd inheritance,
--dangerously-skip-permissions autonomy, busy/idle footer, first-run trust
dialog. Turn-end hook and primary-session guard are NOT yet verified, so agy
is crewmate-only for now (documented in the skill).
fm_harness_pid_alive joined `basename comm` and `ps -o args=` into one
line before matching FM_HARNESS_RE. That broke every anchored
alternative: a live pi harness joins to "pi pi", which `^pi$` cannot
match, so each live pi session read as stale. bin/fm-lock.sh then let a
second session take the primary's lock, and bin/fm-claude-stop-autoarm.sh
reclaimed a lock whose owner was alive - a whole-home mutual-exclusion
failure that can double-dispatch or double-teardown.

Split the check the way fm_harness_ancestry_pid already does it: match
the basename of comm on its own line, and consult args only when comm is
a bare interpreter (node/python). FM_HARNESS_RE is unchanged, so `^pi$`
stays anchored and pip/pipx/paths containing "pi" keep reading as
not-a-harness.

tests/fm-session-lock-liveness.test.sh covers both directions against
real spawned processes (no fake ps) and fails on the pre-fix library.
tasks-axi's Done-retention prune archives resolved captain holds out of the
live backlog (closing the originating task auto-prunes), and the completion
gate could then never verify them again: teardown refused, complete failed
with 'absent from data/backlog.md', and the task deadlocked with no
non-bypass exit (hit live with psak-compose-bringup on 2026-07-26).

Prune archives rather than deletes, so the durable resolution record
survives byte-identically in the markdown backend's archive file. Every
read that accepts a Done hold now falls back to that archive: complete,
verify, identical resolve retries, and reopen refusal are order-independent
with respect to pruning. An archived row without the resolution record
still fails the gate, and an active hold is never read from the archive,
so the safety boundary is unchanged - the gate verifies the same evidence
it always required, from where the backend moved it.

Two regression tests reproduce the deadlock sequence end-to-end and prove
the fail-closed path for an archived hold with no resolution record.
…kill

High-blast-radius ship work now runs as a pair launched together, exchanging
directly through a durable shared file rather than through firstmate, which was
the latency bottleneck when it supervises several threads at once. The protocol
body lives in the new agent-only `paired-review` skill; AGENTS.md gains only the
trigger, the narrowed communication rule, and the delivery-path exception.

The second worker is a navigator, not a reviewer: three of the four gates have
no diff to look at, and an agent told it reviews will idle until one exists,
which defeats the design. The plan gate is the cheap one - a wrong-layer
direction error is visible from a file-level plan in about two minutes and
invisible at PR time under hours of work built on it.

The pair may never settle a contract or scope change, anything destructive,
irreversible, or security-sensitive, or a disagreement surviving its one round
trip; removing firstmate from the exchange moves no approval authority to the
crewmates.

Also make the scope and seam statement a default section of every ship and
scout brief, since the direction error that motivated this protocol happened on
a brief that never named the module owning the code.
…ral questions

Three rules the paired-review protocol did not carry, plus the rule that
keeps skill changes written against the skill-writing guide.

A diagnosis is not paired work. Pairing applies to implementation work; a
bug, regression, or crash-loop takes the diagnose-then-fix shape, one
worker following the diagnosing-bugs skill by absolute path. Until a root
cause exists there is no "where" for a navigator to hold a view about,
and the risk a bug carries is a wrong root cause rather than a wrong
layer, so sequence guards it where parallelism cannot.

The root-cause gate: firstmate judges the root cause before any fix
action. Confident, it authorises the fix and reports; otherwise it
escalates and waits. The root-cause document's own standard is the
confidence test. Destructive, irreversible, security-sensitive, and
contract-expanding fixes still go to the captain regardless.

Every navigator brief carries four fixed structural questions alongside
its task-specific checks: was the stated reason delivered, where did the
coupling go, does this add another instance of a shape the repo has been
burned by, and what did the brief not specify that the implementation had
to decide anyway. Each keeps its reason, because a checklist cannot reach
the destination - a navigator is pointed at a diff, and a diff shows what
changed, never whether the change moves toward where the work is going.

firstmate-coding-guidelines gains one requirement: a skill change begins
by loading writing-great-skills and is written against it.

AGENTS.md changes are two edited trigger lines and no new lines; the
rules themselves live in the skills.
…undary

Replace the paired-review default-routing guidance with a deterministic
navigator rule: every paired navigator runs on pi with the model
cx/gpt-5.6-terra, and only an explicit per-task captain choice of another
navigator runtime or model replaces it. Ordinary dispatch resolution keeps
owning the concrete launch mechanics and carries the pin through instead of
substituting a best-fit alternative.

State the ownership boundary in the same file: paired review is a universal
implementation-task execution protocol owned by whichever firstmate owns the
task, so a standalone paired task needs no persistent orchestrator, while a
program orchestrator dispatches the same protocol unchanged.

Drop the reroute-to-claude clause under the handed-out-skill limitation, which
would otherwise silently replace the pinned navigator runtime.
A planner session is one temporary worker the captain talks to directly.
Firstmate opens it, points it at a captain-defined project, scope, and
planning question, then leaves the conversation and monitors only whether
the session is alive, waiting, finished, or failed.

The planner investigates before it asks anything, reconstructing both what
the code does today and where the project's accepted decisions say it is
going, and surfaces every conflict between them instead of quietly choosing.
It reads tests as evidence and runs nothing. It then grills the captain one
question at a time, each question carrying its recommended answer and each
prescriptive choice carrying both the simplest workable and the best-practice
case. It may not produce a plan until the captain opens the crystallize gate,
after which it publishes a spec or dependency-aware tickets through the
project's own established conventions. Planning never authorizes
implementation.

Two artifact contracts travel downstream, each with one owner:

- The scope envelope carries the accepted boundary at spec level and per
  ticket, reusing bin/fm-brief.sh's existing scope and seam vocabulary plus
  the two fields a brief cannot carry before the target is chosen.
- The test contract carries acceptance intent, not executable test code, and
  the captain approves it together with scope.

planner owns those artifact fields. program-orchestration owns consuming
them: preserve provenance, revalidate, narrow per worker, never widen,
escalate a stale envelope. bin/fm-brief.sh keeps the final worker statement
and paired-review keeps sharing it with the driver and navigator, plus the
navigator's plan-gate challenge to the approved test contract.

The planning disciplines are vendored from mattpocock/skills at 2ab9580 with
the MIT license and a provenance record, adapted only in paths, tracker
assumptions, and delegation wording. No bundled file is named SKILL.md, so
the bundle stays off every harness's skill index and costs no context; it
also cannot silently clobber a captain's own locally installed copy, which
git overwrites without a conflict on fast-forward when the install path is
locally excluded.

Launch reuses the ordinary brief, spawn, and scout-report mechanics with no
new runtime: pi cx/gpt-5.6-sol at high effort by default, claude
claude-opus-5 when explicitly selected.

Also fixes fm-doc-audience-check.sh treating markdown links inside fenced
code blocks as navigation targets, which a vendored example listing exposed.
…tory

A navigator could only assert findings, so a question had to be dressed as a
defect or stay silent, and nothing systematically told it what had already been
decided. One program's navigator invented an undeclared Q entry to work around
exactly this, and another raised a cross-ticket scope question as a finding,
forcing the driver to adjudicate something it had no authority over.

paired-review gains an evidence-backed Q<k> form beside N<k>, on the same one
round trip but with no forced accepted/rejected verdict. The owning firstmate
answers by default and answers it itself, escalating upward only when it lacks
the authority or the knowledge; the driver answers only its own implementation
choices. Navigator briefs carry the decisions bearing on the ticket where a
record exists, and the skill now states that cross-package and whole-solution
direction sits outside one navigator's field of view.

program-orchestration takes the matching duties: maintaining the program's
cross-ticket decision record, handing the bearing decisions to both halves of
every dispatched pair, and holding direction across tickets.

N semantics and the four structural questions are unchanged, and no fifth
question was added. Both briefs are written at dispatch, so this reaches the
next dispatch and leaves a pair already under way on the protocol it launched
with.
The Q form said a question runs "its own sequence" but never said that
sequence is sequential and never reused, so the identifier rule was
checkable for N and not for Q. Match the N wording.
fm_herdr_lab_raw appended `--session <name>` to the end of the herdr argv.
`--session` is a herdr global option, and `herdr agent start <name> ... -- <argv...>`
hands every token after `--` to the inner command, so for any subcommand with a
separator the flag never reached herdr: the call silently ran against whatever
session herdr resolved by default, and the inner command received two junk
arguments. The reported repro, `agent start envcheck -- bash`, executed as
`bash --session fm-lab-...` and the pane died immediately.

The confirmed hypothesis is argument placement, not content. Placement was
inherited from fm_backend_herdr_cli, whose comment records that leading position
was already verified to work and that trailing was chosen only to keep call-site
diffs append-only; the blind spot in that choice is the `--` separator.

Every rejection rule is unaffected: the caller-supplied --session scan, the
leading-option refusal, the lifecycle and server refusals, and the fresh
refuse-default checks all run in fm_herdr_lab_cli on the caller's arguments
before fm_herdr_lab_raw assembles the herdr argv.

The existing behavior test asserted the trailing position as the contract, so the
suite was green on the defect; its assertions are re-pointed, and a new case
records the herdr argv token by token and asserts the session pair leads and
never lands past a `--`, with the inner command passed through untouched.

The identical latent pattern in fm_backend_herdr_cli is deliberately out of scope
here; no current backend call site uses a `--` separator.
Open a long-running programme orchestrator the captain talks to directly,
driving an already-authorized spec, ticket set, or GitHub issues to
completion through workers the session dispatches itself.

The skill splits across the branch boundary it actually has: firstmate
reads SKILL.md to open the session, and only the session reads
CONTRACT.md, so the judgment written for the orchestrator costs
firstmate a pointer rather than a context load. program-orchestration
keeps ownership of custody, routing, host ramp, envelope consumption,
and handoff; both files point at it rather than restating it.

AGENTS.md gains the third direct-conversation exception and one
deliverable-classification bullet. The planner test's closed-exception
assertion moves to the new count rather than dropping the invariant.
@huynhtandat223
huynhtandat223 merged commit 09e79da into main Aug 7, 2026
@huynhtandat223
huynhtandat223 deleted the fm/fm-upstream-rebase-82-upgrade-only branch August 7, 2026 17:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant