Rebase custom Firstmate capabilities onto upstream main - #3
Merged
Merged
Conversation
agy's interactive composer draws a bare ASCII "> " prompt with no side border and no agent-specific glyph, so its empty-composer row is byte-for-byte the dead-shell prompt the shared classifier in bin/fm-composer-lib.sh deliberately refuses as unknown (never empty). Under herdr that made a live-empty agy composer fall to unknown, unlike bordered/agent-glyph shapes, so away-mode injection never fired. Close the gap inside the herdr adapter only, mirroring the Pi separator composer: locate the bottom-most "> " row structurally and admit it as a real composer solely when native `agent get` corroborates the target is agy and reports idle/done/blocked. A working agy, a non-agy or unregistered agent, or an unreadable identity keeps the bare row unknown, so the shared dead-shell rule is not weakened. Identity is fetched at most once per call, shared with the Pi block.
Wire Google Antigravity CLI (agy 1.1.5, Gemini-backed) as a crewmate harness. - fm-harness.sh: detect agy via ANTIGRAVITY_AGENT=1 env marker plus agy/antigravity process-ancestry match. - fm-spawn.sh: interactive launch template (agy --dangerously-skip-permissions -i "$(cat brief)"); add agy to the model flag (--model), effort flag (--effort, ceiling high), and the accepted-harness lists. - fm-tmux-lib.sh / fm-watch.sh: add busy signature "esc to cancel" to the default busy regex (daemon inherits it). - harness-adapters SKILL.md: full agy knowledge section, launch-profile-axes row, and no-mistakes invocation note. - fm-session-start.test.sh: drop ANTIGRAVITY_AGENT in detect_own so the suite does not leak the marker. Verified empirically on agy 1.1.5: env marker, interactive cwd inheritance, --dangerously-skip-permissions autonomy, busy/idle footer, first-run trust dialog. Turn-end hook and primary-session guard are NOT yet verified, so agy is crewmate-only for now (documented in the skill).
fm_harness_pid_alive joined `basename comm` and `ps -o args=` into one line before matching FM_HARNESS_RE. That broke every anchored alternative: a live pi harness joins to "pi pi", which `^pi$` cannot match, so each live pi session read as stale. bin/fm-lock.sh then let a second session take the primary's lock, and bin/fm-claude-stop-autoarm.sh reclaimed a lock whose owner was alive - a whole-home mutual-exclusion failure that can double-dispatch or double-teardown. Split the check the way fm_harness_ancestry_pid already does it: match the basename of comm on its own line, and consult args only when comm is a bare interpreter (node/python). FM_HARNESS_RE is unchanged, so `^pi$` stays anchored and pip/pipx/paths containing "pi" keep reading as not-a-harness. tests/fm-session-lock-liveness.test.sh covers both directions against real spawned processes (no fake ps) and fails on the pre-fix library.
tasks-axi's Done-retention prune archives resolved captain holds out of the live backlog (closing the originating task auto-prunes), and the completion gate could then never verify them again: teardown refused, complete failed with 'absent from data/backlog.md', and the task deadlocked with no non-bypass exit (hit live with psak-compose-bringup on 2026-07-26). Prune archives rather than deletes, so the durable resolution record survives byte-identically in the markdown backend's archive file. Every read that accepts a Done hold now falls back to that archive: complete, verify, identical resolve retries, and reopen refusal are order-independent with respect to pruning. An archived row without the resolution record still fails the gate, and an active hold is never read from the archive, so the safety boundary is unchanged - the gate verifies the same evidence it always required, from where the backend moved it. Two regression tests reproduce the deadlock sequence end-to-end and prove the fail-closed path for an archived hold with no resolution record.
…kill High-blast-radius ship work now runs as a pair launched together, exchanging directly through a durable shared file rather than through firstmate, which was the latency bottleneck when it supervises several threads at once. The protocol body lives in the new agent-only `paired-review` skill; AGENTS.md gains only the trigger, the narrowed communication rule, and the delivery-path exception. The second worker is a navigator, not a reviewer: three of the four gates have no diff to look at, and an agent told it reviews will idle until one exists, which defeats the design. The plan gate is the cheap one - a wrong-layer direction error is visible from a file-level plan in about two minutes and invisible at PR time under hours of work built on it. The pair may never settle a contract or scope change, anything destructive, irreversible, or security-sensitive, or a disagreement surviving its one round trip; removing firstmate from the exchange moves no approval authority to the crewmates. Also make the scope and seam statement a default section of every ship and scout brief, since the direction error that motivated this protocol happened on a brief that never named the module owning the code.
…ral questions Three rules the paired-review protocol did not carry, plus the rule that keeps skill changes written against the skill-writing guide. A diagnosis is not paired work. Pairing applies to implementation work; a bug, regression, or crash-loop takes the diagnose-then-fix shape, one worker following the diagnosing-bugs skill by absolute path. Until a root cause exists there is no "where" for a navigator to hold a view about, and the risk a bug carries is a wrong root cause rather than a wrong layer, so sequence guards it where parallelism cannot. The root-cause gate: firstmate judges the root cause before any fix action. Confident, it authorises the fix and reports; otherwise it escalates and waits. The root-cause document's own standard is the confidence test. Destructive, irreversible, security-sensitive, and contract-expanding fixes still go to the captain regardless. Every navigator brief carries four fixed structural questions alongside its task-specific checks: was the stated reason delivered, where did the coupling go, does this add another instance of a shape the repo has been burned by, and what did the brief not specify that the implementation had to decide anyway. Each keeps its reason, because a checklist cannot reach the destination - a navigator is pointed at a diff, and a diff shows what changed, never whether the change moves toward where the work is going. firstmate-coding-guidelines gains one requirement: a skill change begins by loading writing-great-skills and is written against it. AGENTS.md changes are two edited trigger lines and no new lines; the rules themselves live in the skills.
…undary Replace the paired-review default-routing guidance with a deterministic navigator rule: every paired navigator runs on pi with the model cx/gpt-5.6-terra, and only an explicit per-task captain choice of another navigator runtime or model replaces it. Ordinary dispatch resolution keeps owning the concrete launch mechanics and carries the pin through instead of substituting a best-fit alternative. State the ownership boundary in the same file: paired review is a universal implementation-task execution protocol owned by whichever firstmate owns the task, so a standalone paired task needs no persistent orchestrator, while a program orchestrator dispatches the same protocol unchanged. Drop the reroute-to-claude clause under the handed-out-skill limitation, which would otherwise silently replace the pinned navigator runtime.
A planner session is one temporary worker the captain talks to directly. Firstmate opens it, points it at a captain-defined project, scope, and planning question, then leaves the conversation and monitors only whether the session is alive, waiting, finished, or failed. The planner investigates before it asks anything, reconstructing both what the code does today and where the project's accepted decisions say it is going, and surfaces every conflict between them instead of quietly choosing. It reads tests as evidence and runs nothing. It then grills the captain one question at a time, each question carrying its recommended answer and each prescriptive choice carrying both the simplest workable and the best-practice case. It may not produce a plan until the captain opens the crystallize gate, after which it publishes a spec or dependency-aware tickets through the project's own established conventions. Planning never authorizes implementation. Two artifact contracts travel downstream, each with one owner: - The scope envelope carries the accepted boundary at spec level and per ticket, reusing bin/fm-brief.sh's existing scope and seam vocabulary plus the two fields a brief cannot carry before the target is chosen. - The test contract carries acceptance intent, not executable test code, and the captain approves it together with scope. planner owns those artifact fields. program-orchestration owns consuming them: preserve provenance, revalidate, narrow per worker, never widen, escalate a stale envelope. bin/fm-brief.sh keeps the final worker statement and paired-review keeps sharing it with the driver and navigator, plus the navigator's plan-gate challenge to the approved test contract. The planning disciplines are vendored from mattpocock/skills at 2ab9580 with the MIT license and a provenance record, adapted only in paths, tracker assumptions, and delegation wording. No bundled file is named SKILL.md, so the bundle stays off every harness's skill index and costs no context; it also cannot silently clobber a captain's own locally installed copy, which git overwrites without a conflict on fast-forward when the install path is locally excluded. Launch reuses the ordinary brief, spawn, and scout-report mechanics with no new runtime: pi cx/gpt-5.6-sol at high effort by default, claude claude-opus-5 when explicitly selected. Also fixes fm-doc-audience-check.sh treating markdown links inside fenced code blocks as navigation targets, which a vendored example listing exposed.
…tory A navigator could only assert findings, so a question had to be dressed as a defect or stay silent, and nothing systematically told it what had already been decided. One program's navigator invented an undeclared Q entry to work around exactly this, and another raised a cross-ticket scope question as a finding, forcing the driver to adjudicate something it had no authority over. paired-review gains an evidence-backed Q<k> form beside N<k>, on the same one round trip but with no forced accepted/rejected verdict. The owning firstmate answers by default and answers it itself, escalating upward only when it lacks the authority or the knowledge; the driver answers only its own implementation choices. Navigator briefs carry the decisions bearing on the ticket where a record exists, and the skill now states that cross-package and whole-solution direction sits outside one navigator's field of view. program-orchestration takes the matching duties: maintaining the program's cross-ticket decision record, handing the bearing decisions to both halves of every dispatched pair, and holding direction across tickets. N semantics and the four structural questions are unchanged, and no fifth question was added. Both briefs are written at dispatch, so this reaches the next dispatch and leaves a pair already under way on the protocol it launched with.
The Q form said a question runs "its own sequence" but never said that sequence is sequential and never reused, so the identifier rule was checkable for N and not for Q. Match the N wording.
fm_herdr_lab_raw appended `--session <name>` to the end of the herdr argv. `--session` is a herdr global option, and `herdr agent start <name> ... -- <argv...>` hands every token after `--` to the inner command, so for any subcommand with a separator the flag never reached herdr: the call silently ran against whatever session herdr resolved by default, and the inner command received two junk arguments. The reported repro, `agent start envcheck -- bash`, executed as `bash --session fm-lab-...` and the pane died immediately. The confirmed hypothesis is argument placement, not content. Placement was inherited from fm_backend_herdr_cli, whose comment records that leading position was already verified to work and that trailing was chosen only to keep call-site diffs append-only; the blind spot in that choice is the `--` separator. Every rejection rule is unaffected: the caller-supplied --session scan, the leading-option refusal, the lifecycle and server refusals, and the fresh refuse-default checks all run in fm_herdr_lab_cli on the caller's arguments before fm_herdr_lab_raw assembles the herdr argv. The existing behavior test asserted the trailing position as the contract, so the suite was green on the defect; its assertions are re-pointed, and a new case records the herdr argv token by token and asserts the session pair leads and never lands past a `--`, with the inner command passed through untouched. The identical latent pattern in fm_backend_herdr_cli is deliberately out of scope here; no current backend call site uses a `--` separator.
Open a long-running programme orchestrator the captain talks to directly, driving an already-authorized spec, ticket set, or GitHub issues to completion through workers the session dispatches itself. The skill splits across the branch boundary it actually has: firstmate reads SKILL.md to open the session, and only the session reads CONTRACT.md, so the judgment written for the orchestrator costs firstmate a pointer rather than a context load. program-orchestration keeps ownership of custody, routing, host ramp, envelope consumption, and handoff; both files point at it rather than restating it. AGENTS.md gains the third direct-conversation exception and one deliverable-classification bullet. The planner test's closed-exception assertion moves to the new count rather than dropping the invariant.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rebases the retained captain custom capabilities directly onto current upstream/main.\n\nValidated locally: lint, documentation audience check, session-lock liveness, fleet snapshot, planner, and orchestrator focused suites.