Skip to content

feat(orchestrator): open /orchestrator as a temporary supervisor, not a secondmate - #32

Merged
huynhtandat223 merged 1 commit into
mainfrom
fm/fm-temporary-supervisor-issue31
Aug 9, 2026
Merged

huynhtandat223 merged 1 commit into
mainfrom
fm/fm-temporary-supervisor-issue31

Conversation

@huynhtandat223

Copy link
Copy Markdown
Owner

What

/orchestrator opened its programme session as an ordinary scout, which cannot supervise children, while /orchestrator and program-orchestration described that session as a persistent secondmate it never was. Issue #31 accepts a third shape: one bounded temporary supervisor with a firstmate home of its own, which dispatches and supervises ordinary workers itself and closes when its programme ends.

Keeping the core seam small and removable

The captain's constraint on this work was to limit changes in bin/ so upstream upgrades stay easy, and to keep whatever lands there easy to remove or migrate.

The whole lifecycle is in one new file, bin/fm-supervisor-lib.sh. Every other core file gets one-line delegations, and grep -rn fm_supervisor bin/ lists that footprint in full:

File Change
bin/fm-supervisor-lib.sh new - lease, home identity, isolation assertions, relaunch reuse, self-supervising kind set
bin/fm-primary-scope-lib.sh generalized its marker reader; one || in the scope predicate
bin/fm-watch.sh 5 predicate swaps + 1 import
bin/fm-crew-state.sh 1 predicate swap + 1 import
bin/fm-spawn.sh the --supervisor kind
bin/fm-teardown.sh the supervisor cleanup gates
bin/fm-test-run.sh 6-line changed-file → family map entry

To drop the feature: delete bin/fm-supervisor-lib.sh and those call sites. To carry it across an upstream sync: re-apply the call sites and leave that file alone.

Three things were deliberately kept out of core:

  • AGENTS.md unchanged. The skill is reachable through the custom-skills/policy/SKILL.md router that AGENTS.md line 7 already points at, so the always-loaded surface pays nothing for a feature most sessions never use.
  • .gitignore unchanged. The .fm-supervisor-home marker is excluded through the worktree's own info/exclude, the same mechanism spawn already uses for its hook tokens.
  • bin/fm-brief.sh unchanged. The supervisor's brief is programme policy, so custom-skills/orchestrator/fm-supervisor-brief.sh writes it.

The design

  • Leased home. A fresh spawn takes treehouse get --lease. The lease is the whole point: a leased worktree is never handed out again and never pruned, so the home - and with it the supervisor's own record of its workers - survives a dead agent.
  • No cloning. No product project is created in that home. The parent home's existing clones stay read-only allocation sources for the isolated worker copies the supervisor allocates itself.
  • One direct report. Main firstmate manages the supervisor and nothing below it. No data/secondmates.md line, no charter, no inherited local material, no liveness sweep.
  • Relaunch, not restart. Re-running the same spawn returns into the recorded home and reconciles the workers already there. A recorded home that is missing refuses loudly rather than quietly starting a second one, because duplicating a ticket's worker is the failure this design exists to prevent.
  • Idle is healthy. A report that runs its own firstmate session is read from its status writes rather than its pane, so a supervisor waiting on its children is not surfaced as stale.
  • Cleanup refuses on four counts - a child worker record left in the home, unlanded work in the home, an unresolved captain decision, or a missing programme report - then retires the home and releases the lease.

Policy stays in custom-skills/: the orchestrator skill now opens, relaunches, and closes a temporary supervisor, and program-orchestration no longer routes through secondmate provisioning. Those documents were rewritten against custom-skills/matt/productivity/writing-for-agents/SKILL.md.

Verification

  • tests/fm-temporary-supervisor.test.sh - 13 new cases through the real executables: home isolation, no project cloning, primary-scope marker, child dispatch into the supervisor's own home, same-home relaunch, single-lease/no-duplicate, id reuse refusal, mode/positional refusals, each cleanup refusal, and lease release.
  • custom-skills/tests/fm-orchestrator-policy.test.sh - 7 new cases on the skill/runtime contract.
  • Families: secondmate 18/18, watcher-wake-lock 12/12, backend-dispatch 12/12, plus fm-teardown, fm-brief, fm-crew-state, fm-turnend-guard, fm-subagent-pretool-check, fm-sessionstart-nudge, fm-documentation-audiences, fm-lint, fm-test-run.
  • bin/fm-lint.sh clean; bin/fm-test-run.sh --check-coverage ok; bin/fm-doc-audience-check.sh now passes - it was already red on main because the three orchestrator prose surfaces were unclassified, and this adds their agent-runtime entries.

One local note for the reviewer: tests/fm-backend-herdr.test.sh fails in my worktree only, because that pooled worktree carries a stale gitignored .fm-secondmate-home from a previous occupant, which changes the Herdr workspace label the fixture expects. The same suite passes 12/12 in a clean clone of the same base commit with this branch applied.

Delivery

direct-PR. Do not merge without the captain's explicit approval.

Closes #31

… a secondmate

`/orchestrator` opened its programme session as an ordinary scout, which cannot
supervise children, and the surrounding docs described it as a persistent
secondmate it never was. Issue #31 accepts a third shape: one bounded temporary
supervisor with a firstmate home of its own, which dispatches and supervises
ordinary workers itself and closes when its programme ends.

The whole lifecycle lives in one new file, bin/fm-supervisor-lib.sh, so the
footprint in the rest of bin/ is a set of one-line delegations that
`grep -rn fm_supervisor bin/` lists in full - the seam can be removed by
deleting that file and those call sites, or carried across an upstream upgrade
by re-applying only the call sites.

Runtime:
- bin/fm-supervisor-lib.sh: the lease, home identity, isolation assertions,
  relaunch reuse, and the self-supervising kind set.
- bin/fm-spawn.sh --supervisor: leases an isolated firstmate worktree with
  `treehouse get --lease`, marks it, records home= and parent_home=, and
  launches the session in it. No project is cloned there; the parent home's
  existing clones stay read-only allocation sources. A relaunch returns to the
  recorded home rather than leasing a second one, so a stopped supervisor comes
  back to its own worker inventory and cannot re-dispatch duplicates. A failed
  launch gives the fresh lease back. No secondmate registry, charter, inherited
  material, or liveness sweep is involved.
- bin/fm-primary-scope-lib.sh: the .fm-supervisor-home marker puts that home in
  scope for its own hooks, alongside the existing secondmate marker.
- bin/fm-watch.sh, bin/fm-crew-state.sh: a report that runs its own firstmate
  session is read from its status writes, not its pane, so an idle supervisor
  is healthy rather than stale.
- bin/fm-teardown.sh: cleanup refuses until the child records, the home's own
  unlanded work, the unresolved-decision gate, and the programme report are all
  reconciled, then retires the home and releases its lease.

Policy stays in custom-skills/: the orchestrator skill now opens, relaunches,
and closes a temporary supervisor; program-orchestration no longer routes
through secondmate provisioning; and fm-supervisor-brief.sh scaffolds the
session's contract there rather than in core.

AGENTS.md, .gitignore, and bin/fm-brief.sh are deliberately untouched: the
marker is excluded per worktree, the brief is programme policy, and the skill
is reachable through the policy router AGENTS.md already points at.

Verification: tests/fm-temporary-supervisor.test.sh (13 cases: home isolation,
no project cloning, primary scope, child dispatch, same-home relaunch,
duplicate prevention, every cleanup refusal, lease release),
custom-skills/tests/fm-orchestrator-policy.test.sh (7 cases), plus the
secondmate, watcher-wake-lock, and backend-dispatch families and bin/fm-lint.sh
clean.

Closes #31
@huynhtandat223
huynhtandat223 merged commit 8b5a038 into main Aug 9, 2026
9 of 12 checks passed
huynhtandat223 added a commit that referenced this pull request Aug 10, 2026
The gotmp fixture builds a fake FM_HOME and symlinks each library
bin/fm-teardown.sh sources into it. The orchestrator work (#32) added
bin/fm-supervisor-lib.sh, made teardown source it, and left the fixture's
symlink list untouched, so teardown died inside the test on a missing sibling:

  .../bin/fm-supervisor-lib.sh: No such file or directory

Real teardown was never affected - the file exists in a real bin/. Only the
fake root was incomplete, and this is fork-only: upstream has no
fm-supervisor-lib.sh and no source line for it, which is why upstream stays
green here.

fm-supervisor-lib.sh sources fm-primary-scope-lib.sh in turn, so both are
needed; symlinking only the first moves the same failure one library along.

This failure has been red on main since #32 landed. A permanently red CI stops
being able to report the next real regression, which is the actual cost.
huynhtandat223 added a commit that referenced this pull request Aug 10, 2026
…ext (#36)

* fix(composer): read a blank-padded composer as empty, not as unsent text

Claude 2.1.226 pads its EMPTY composer row with U+00A0 after the `❯` prompt
glyph. The blank is drawn at luminance 153, so ghost stripping correctly keeps
it, and every trim in the shared classifier and its callers is ASCII-only - so
a genuinely empty Claude composer classified `pending`, as if it still held
unsubmitted text.

Nothing failed loudly, because the composer verdict only decides delivery where
no stronger signal exists. An ordinary steer to an idle worker is confirmed
from Herdr's native agent state and never reads the composer at all. The one
path with no stronger signal is Herdr's submit confirmation for a target that
was ALREADY mid-turn before Enter, and that is exactly the path a paired-review
barrier release takes, because both roles sit in a blocking wait when the
release is sent. fm-send reported `delivery unconfirmed; verdict=pending` for a
message the driver had already received, and the composer correctly refused to
release the navigator, leaving the pair without coverage before its plan gate.

Fold non-ASCII Unicode blanks to ASCII space in fm_composer_classify_content,
the one fleet-wide owner every adapter delegates to. This can only move content
that is ENTIRELY blank from `pending` to `empty`; content carrying any visible
character keeps its verdict, so a genuinely swallowed Enter still refuses and
the away-mode injector still declines to type over unsent input. Keying on the
Unicode space-separator category rather than on U+00A0 alone keeps the fix from
depending on which blank a given release happens to pick.

Verified against real claude 2.1.226 on herdr 0.7.3 through the production
adapter, both directions: a message accepted by a mid-turn worker now confirms
(and was observably delivered), while text typed but never submitted still
reports pending.

Regressions:
- tests/fm-composer-lib.test.sh pins the fold, every declared blank, and the
  divergence - padded-empty is empty while padded-real-text stays pending, so
  the case cannot go vacuous - plus the unchanged dead-shell refusal.
- tests/fm-backend-herdr.test.sh pins the busy-baseline submit path this fault
  actually reached, in both directions, on captures carrying the real bytes.
- tests/fm-composer-blank-drift-live-e2e.test.sh is the opt-in real-harness
  guard, since what a harness draws in an empty composer is a vendor surface a
  fixture can only replay from a previous release. It records the untouched
  verdict, then asserts it only after probe text typed into that pane becomes
  visible, so a blank startup screen or a trust dialog is reported unverified
  instead of passing. It fails on claude without this fix and passes with it.

Compatibility: the shared owner is consumed by the tmux, herdr, cmux and orca
adapters; all four suites pass, and zellij does not use it.

* test(composer): evidence every installed harness, and stop running agents in the tree

The blank fold needed per-harness proof, not just proof for the harness that
provoked it. All five installed harnesses were measured against their own live
binary, idle and holding typed-but-unsent text, and the byte-exact captures are
kept under tests/fixtures/composer-captures/ and replayed through the real
Herdr composer reader.

Measured against the pre-fix classifier, exactly one of the ten verdicts moves:

  harness   version    idle(pre)  idle(now)  typed     pads a Unicode blank
  claude    2.1.226    pending    empty      pending   yes, U+00A0
  codex     0.145.0    empty      empty      pending   no
  pi        0.84.1     empty      empty      pending   no
  opencode  1.17.20    unknown    unknown    unknown   no
  agy       1.1.11     unknown    unknown    unknown   no

Claude is the only one of the five that pads its composer at all, so the other
four are evidence that the fold changed nothing for them, not evidence that it
repaired them. Recording that distinction matters more than a green row.

Two seams are unsupported on BOTH backends and are recorded rather than papered
over. opencode draws a composer row with a left-only edge glyph, matching
neither the bordered shape nor a bare agent prompt. agy draws a bare `>`,
byte-identical to a dead shell prompt, and the rule that would admit it on
native identity requires that row to sit below all separator activity while
real agy draws a rule underneath its prompt. Both refuse as `unknown`, which
denies delivery confirmation and away-mode injection alike, and neither is
affected by the fold. Reading either composer is a separate change with its own
safety argument and is not attempted here.

agy was also missing from the live guard's harness loop entirely, so the
agy-specific prompt rule had never been exercised against the real binary.

The guard no longer launches harnesses inside the checkout. During this work a
real harness launched in the firstmate worktree modified a tracked file with no
prompt ever submitted to it; the change was reverted and did not reach a commit.
A guard that runs live agents must not run them in the tree it is testing, so
launches now default to a scratch directory, and FM_COMPOSER_BLANK_DRIFT_DIR
takes an already-trusted disposable directory for fuller coverage and refuses
this checkout. The honest cost is coverage: from a scratch directory only pi
draws a composer here, because the others gate on directory trust first.

* test(gotmp): give the fake root the libraries teardown actually sources

The gotmp fixture builds a fake FM_HOME and symlinks each library
bin/fm-teardown.sh sources into it. The orchestrator work (#32) added
bin/fm-supervisor-lib.sh, made teardown source it, and left the fixture's
symlink list untouched, so teardown died inside the test on a missing sibling:

  .../bin/fm-supervisor-lib.sh: No such file or directory

Real teardown was never affected - the file exists in a real bin/. Only the
fake root was incomplete, and this is fork-only: upstream has no
fm-supervisor-lib.sh and no source line for it, which is why upstream stays
green here.

fm-supervisor-lib.sh sources fm-primary-scope-lib.sh in turn, so both are
needed; symlinking only the first moves the same failure one library along.

This failure has been red on main since #32 landed. A permanently red CI stops
being able to report the next real regression, which is the actual cost.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a lightweight temporary supervisor for /orchestrator

1 participant