Skip to content

fix(bin): speed up session start and merge upstream firstmate main - #6176

Closed
keenvc wants to merge 51 commits into
kunchenguid:mainfrom
keenvc:fm/fm-upstream-sync
Closed

keenvc wants to merge 51 commits into
kunchenguid:mainfrom
keenvc:fm/fm-upstream-sync

Conversation

@keenvc

@keenvc keenvc commented Sep 30, 2026

Copy link
Copy Markdown

Intent

Clean up the firstmate herdr instance, which has become useless for getting work done. Review https://github.com/kunchenguid/kun and https://github.com/kunchenguid/firstmate, debug the current system, and fix the issues so it is back to being a fast agentic software factory that can be used and that does not crash and have issues all the time.

Context from the diagnosis: this home runs a private fork (keenvc/firstmate, main at 9928ab1) that is 164 commits behind upstream kunchenguid/firstmate main and carries about 30 fork-only commits. Upstream has since fixed watcher stalls and re-arm failures (#6103, #5941, #5594, #5362, #5732), reclaim of tasks whose herdr endpoint was destroyed (#5007), teardown retiring watcher markers, orphan journals and wake rows (#5997, #5390), session-start fixes (#6125), and bounded status classification (#5383). Session start currently exceeds its 120s bound, 47 of 99 task records read as "unreadable runs table", and a straight merge of upstream into the fork reports 32 conflicts.

What Changed

  • Merged upstream kunchenguid/firstmate main into the fork. This brings in the watcher stall and re-arm fixes in bin/fm-watch.sh, and reclaim of tasks whose herdr endpoint was destroyed. It also brings teardown cleanup of watcher markers, orphan journals and wake rows in bin/fm-teardown.sh. New upstream scripts arrive with it: fm-claim.sh, fm-claim-lib.sh, fm-hold-reverify.sh, fm-provider-lib.sh and fm-provider-load.sh. Cline and OpenHands harness adapters, docs and verification notes are added, and fm-spawn.sh, fm-busy-lib.sh, fm-composer-lib.sh and fm-fleet-snapshot.sh are extended.
  • Bounded herdr CLI calls in bin/backends/herdr.sh and bin/fm-backend.sh with a timeout cap, and adjusted bin/fm-session-start.sh and bin/fm-crew-state.sh, so session start and status classification no longer hang past their bounds. The fleet-snapshot ARG_MAX file transport is restored after the merge. The timeout cap is documented in docs/herdr-backend.md and docs/configuration.md.
  • Added and updated shell tests for the merged behavior: claim, hold-reverify, provider lane cap, the herdr probe timeout, teardown, watch triage, and the Cline and OpenHands harnesses. The .gitignore and CI workflow are also updated.

Risk Assessment

✅ Low: The only change since the prior review is a one-line removal of a duplicate .omc/ entry from .gitignore, and .omc/ is still ignored once.

Testing

I ran the session-start, herdr probe-timeout, teardown, watcher-arm and herdr-lab test files. Two real-Herdr e2e tests and a live lab provision/run/teardown also passed. The session-start, probe-timeout, teardown and watcher-arm checks used fakes or fixtures, so they are recorded as untested against the live product. I did not launch a real primary or spawn a lane in the lab, so the intent's "session start under 120s" and "unreadable runs table" symptoms were not measured against real fleet state. Watch triage was not re-run. Running bin/fm-herdr-lab.sh prepare first made the following provision refuse an existing tripwire, so I ran provision alone.

  • Live validation: ✅ go - 2 of 6 scenarios driven live against the product
Scenario Result Live Evidence
A hung or killed herdr endpoint read during session start is bounded, reported as that task's error line, and leaves no stray herdr process ⏸️ untested no The prior payload only ran the repo's test with a fake herdr, not a real herdr server or live session start, so it did not establish a live result. To cover it, run session start against a real primar…
A hung herdr CLI call is bounded and reaped by the shared read owner, while the long-lived server launch stays exempt ⏸️ untested no The prior payload only ran the repo's test with a fake hung herdr, so it did not establish a live result. To cover it, drive a real herdr CLI call that hangs.
Teardown reclaims tasks, retires watcher markers, and refuses unsafe slot ownership ⏸️ untested no The prior payload only ran fixture-driven repo tests, not a live fleet, so it did not establish a live result. To cover it, run teardown against a live fleet with real tasks.
Watcher arm, re-arm and recovery keep wakes durable and quarantine malformed state ⏸️ untested no The prior payload only ran a fixture-driven repo test, not a live fleet, so it did not establish a live result. To cover it, arm and re-arm a real watcher on a live fleet.
Live Herdr lab: provision an fm-lab-* session, run a herdr command, and tear it down ✅ pass live herdr-lab-live.log
Real Herdr e2e: session cleanup and respawn idempotency ✅ pass live fm-herdr-session-cleanup-e2e.log, fm-backend-herdr-respawn-idem-e2e.log
Evidence: herdr lab live transcript

Source: herdr lab live transcript

== fm-lab-gate-2190546-13671
provision=0
{"id":"cli:workspace:list","result":{"type":"workspace_list","workspaces":[]}}
run=0
teardown=0
Evidence: session-start test log

Source: session-start test log

ok - context digest distinguishes ABSENT, empty-but-present, and populated files
ok - a lock refusal prints a loud read-only banner, skips every mutating step, and still completes the digest
ok - session start stays read-only when lock ownership cannot be published
ok - locked session start freezes trace context and lock refusal leaves it unchanged
ok - concurrent session-lock acquisition admits exactly one live harness
ok - digest sections are ordered safety-preamble first, live fleet state before curated memory
ok - the read-once contract is stated once, ahead of the sources it governs
ok - session start: configured and auto-detected Herdr homes never require tmux
ok - session start: an absent recorded tmux window relaunches its Pi secondmate exactly once, off the blocking path
ok - session start: a deferred relaunch is always reported, so the digest's stale endpoint record cannot stand
ok - session start: inactive reconciliation runs after the digest and retains its durable wake
ok - session start: an unreachable host delays a reported check, not the digest
ok - session start: a deferred result the digest outran still reaches the agent as a wake
ok - session start: a read-only session declares its skipped network checks rather than dropping them
ok - session start: the tasks-axi compatibility verdict is computed once and reused
ok - session start: an existing ambiguous Pi process prevents duplicate recovery
ok - session start: transient tmux unreadability never licenses a relaunch
ok - session start: the proven bare-shell recovery path remains intact
ok - session start: a confirmed Herdr husk is closed and relaunched
ok - status tail is bounded to the configured line count, with the full log path always printed
ok - status tail lines are capped with a truncation marker while the full log stays reachable
ok - orphan status logs are printed once with bounded tails
ok - tmux endpoint liveness is reported per task: alive for a live window, dead for a gone one
ok - herdr endpoint liveness is reported per task: alive, dead for exit 1, dead for any other probe status
ok - a killed per-task endpoint read becomes that task's error line and the digest completes
ok - a hung per-task endpoint read hits its configured bound, reports the task, and leaves nothing stuck
ok - a padded-zero per-read bound falls back to the 10s default instead of removing the bound
ok - the perl timeout fallback reports a signal death as a nonzero status
Terminated
ok - a digest child killed mid-stage is bannered by the parent, which still exits 0
ok - fm-session-start.sh composes the real fm-lock.sh, fm-bootstrap.sh, and fm-wake-drain.sh output verbatim
ok - locked Pi session start replays leading routine outcomes, preserves the captain barrier, and sweeps only dead leases
ok - non-Pi session start neither sweeps nor replays Pi branch state
ok - session start seeds an existing outcome store's absent display tail copy while away, moving no marker
ok - compatible tasks-axi backlog rendering drops done rows and keeps every in-flight, held, and blocked row
ok - the startup backlog bound cuts only dispatchable queued rows and discloses the remainder exactly
ok - manual backlog rendering drops done rows, keeps every held or blocked title line, and bounds the rest
ok - unavailable or incompatible tasks-axi falls back to compact manual backlog rendering
ok - an empty fleet reports (none) for in-flight tasks and an absent AFK flag
ok - session start emits X-mode cadence guidance in the harness supervision block
ok - next step delegates watcher ownership to the AFK daemon
ok - next step delegates watcher ownership to the daemon in quiet mode, distinctly from away mode
ok - the AFK digest reads a quiet record as a present captain holding nothing, and an away record as hold-for-return
ok - a legacy empty .afk flag (written before mode existed) still reads as away mode
ok - session start emits exactly one detected harness block and reports Pi extension load state
ok - session start preserves pi-signed primary identity while applying Pi extension guarantees
ok - session start rejects stale Pi loaded markers
ok - session start rejects a Pi watcher generation left in handoff
ok - session start accepts current Pi markers written before lock acquisition
ok - session start emits the omp block and reports omp extension load state
ok - session start accepts current omp markers written before lock acquisition
ok - session start rejects Pi sessions missing the turn-end guard marker
ok - session start rejects Pi loaded markers from previous sessions
ok - the pure-Bash watchdog bounds session start, kills its hung grandchild, and emits the truncation contract
ok - the portable timeout path force-kills a command that ignores TERM
ok - a session start inside its budget prints no truncation banner
ok - the runtime bound leaves enough ancestry headroom for a deeply nested session to take the lock
ok - --reemit reprints the digest without repeating startup's mutating sweeps and still drains queued wakes
ok - true-start AGENTS baselines stay immutable while every drifted Pi compact re-emits the current contract
ok - read-only Pi compact refreshes against the rebuilding session identity without mutation
ok - Codex reset sources do not claim an unavailable instruction-refresh channel
ok - instruction baselines require SHA-256 and successful startup completion
ok - --reemit re-verifies lock ownership and keeps repair ownership with whoever holds it
# fm-session-start.test.sh: all assertions passed
exit=0
Evidence: herdr probe-timeout test log

Source: herdr probe-timeout test log

ok - capture read: a TERM-ignoring hung herdr is bounded and its process is reaped
ok - composer_state probe: a TERM-ignoring hung herdr is bounded and its process is reaped
ok - fm_backend_herdr_cli: the shared read owner bounds and reaps a hung herdr
ok - fm_backend_herdr_cli: the long-lived server launch stays exempt from the bound
exit=0
Evidence: teardown test log

Source: teardown test log

ok - a missing teardown startup source refuses before cleanup
ok - an unreadable teardown startup source refuses before cleanup
ok - a missing adapter sibling refuses before cleanup
ok - a forced descendant with a missing adapter sibling refuses before cleanup
ok - a forced secondmate with a missing own adapter sibling refuses before child cleanup
ok - present required sources still reach the ordinary teardown refusal
ok - local-only worktree with HEAD on a fork remote is torn down and the home summary is refreshed
ok - teardown closes its own backlog item before reporting success
ok - teardown honors config/backlog-backend=manual and still finishes cleanly
ok - local-only worktree with truly unpushed work is refused (safety preserved)
ok - local-only worktree with work merged into local main is torn down (no regression)
ok - no-mistakes worktree with HEAD on origin is torn down (no regression)
ok - no-mistakes worktree with genuinely unlanded work is refused (safety preserved)
ok - local-only worktree with unpushed work is torn down under --force (escape hatch)
ok - fm-pr-check publishes the PR-ready line on a secondmate's parent channel once
ok - a secondmate home's teardown delivers the child's final line or refuses until it can
ok - teardown completes when an exact busy-state sidecar is already absent
ok - herdr teardown removes pane-owned escalation dedupe state
ok - herdr flat teardown refuses before returning the isolated copy under lock contention and the retry completes cleanly
ok - herdr flat teardown never erases records when pane presence is unparseable
ok - herdr flat teardown preflight refuses before every destructive change
ok - forced secondmate teardown preflights every Herdr child before cleanup mutation
ok - forced secondmate teardown holds every descendant lifecycle and metadata lock
ok - forced secondmate teardown retains Herdr child identity until exact pane disappearance
ok - forced teardown retains a nested secondmate home and its grandchild's Herdr identity when the grandchild close is unconfirmed
ok - herdr projection teardown retires its journal only after confirming the exact recorded pane is gone
ok - herdr projection teardown retains every record when post-close presence is unknown
ok - herdr projection teardown surfaces failed focus restoration without turning confirmed cleanup into a hard failure
ok - a projected teardown removes the workspace its task-pane close left behind, without calling workspace close
ok - a projected teardown refuses and retains every record while its workspace cannot be confirmed gone
ok - teardown retires the task's own watcher markers and orphaned presentation journal, leaving other tasks' markers alone
ok - teardown retains a presentation journal bound to a pane other than the closed endpoint
ok - teardown retires a v1 presentation journal once its token workspace is confirmed gone
ok - teardown retains a v1 presentation journal while its token workspace is still present
ok - teardown retains a v1 presentation journal when the workspace query is ambiguous
ok - squash-merged + deleted-branch worktree (PR merged) is torn down (the fix)
ok - squash-merged PR accepts a local HEAD that is an ancestor of the final PR head
ok - teardown discovers a merged PR by branch name and tears down when no pr= was ever recorded
ok - squash-merged PR accepts replayed unpushed local patches contained in the PR head
ok - merged PR does not allow teardown after a later local commit
ok - squash-merged task whose local branch followed the pipeline rebase is torn down
ok - squash-merged same-path different content still refuses
ok - squash-merged rebased local still refuses a genuinely unlanded follow-up commit
ok - squash-merged stale local still refuses when the forge is unreachable
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
ok - fm-pr-check does not refresh PR head after HEAD moves
fm-contributions: data directory unavailable
contributions: observation not armed; coverage is unconfirmed
ok - fm-pr-check records the remote PR head when the local worktree lags
ok - worktree whose content already landed in the default branch is torn down (content fallback)
ok - content fallback refreshes origin default before comparing trees
ok - dirty worktree is refused even when its committed work has landed (dirty always wins)
ok - gh lookup error with content not in default refuses (fail-safe)
ok - a record predating spawn_gen refuses teardown until --legacy-record is passed
ok - a windowless leftover with no spawn_gen and no worktree tears down without --legacy-record
ok - a windowless leftover with no spawn_gen also tears down when --legacy-record is passed
ok - a windowless leftover still refuses while its worktree holds unlanded work
ok - a windowless record with a spawn_gen, a non-tmux backend or endpoint identity, no backlog validation, or ambiguous, foreign, or malformed identity still refuses
ok - a windowless leftover retries its retained legacy stamp without --legacy-record
ok - a landed legacy record with a dead endpoint tears down and logs its accepted incarnation
ok - --legacy-record never relaxes the unlanded-work refusal
ok - an endpoint that cannot be confidently read as dead refuses --legacy-record teardown
ok - --legacy-record teardown rolls its stamp back when the close marker write fails
ok - an explicitly cleared endpoint tears down without --force or --legacy-record
ok - an explicitly cleared endpoint never relaxes the unlanded-work refusal
ok - a missing window without an endpoint_cleared stamp still refuses
ok - a legacy stamp a failed rollback left behind still faces the endpoint gate
ok - a corrupt spawn_gen is never accepted as a legacy record
ok - provably-stale worktree index.lock (old, no live holder) is cleared and teardown succeeds
ok - live-held worktree index.lock is never removed and teardown refuses
ok - lsof errors leave worktree index.lock in place and refuse teardown
ok - stale lock cleanup rechecks and refuses dirty worktree before return
ok - normal repo index.lock is resolved from the worktree and cleared when stale
ok - lock mtime read failures leave worktree index.lock in place and refuse teardown
ok - transient index.lock cleared after first failed return is retried successfully without force-remove
ok - persistent index.lock exhausts retries and refuses without force-removing the lock
ok - empty retry wait overrides use the default without aborting teardown
ok - fractional legacy retry wait remains supported without arithmetic
ok - a task's own parked no-mistakes run is aborted, not orphaned, before the worker is removed
ok - a run that lands on passed-with-override after abort is still recognized as terminal
ok - a run that lands on passed-with-skips after abort is still recognized as terminal
ok - a parked run the pipeline advanced past the task copy is still concluded from the runs ledger, not orphaned
ok - a ledger row for a different head never authorizes a parked-run abort
ok - a malformed ledger row never authorizes a parked-run abort
ok - an impossible ledger date never authorizes a parked-run abort
ok - a terminal status with a stale gate never reaches ledger cleanup
ok - an advanced head present locally aborts through the strict rule alone - the ledger fallback stays dormant
ok - an unresolvable active row with no same-branch anchor is never concluded (conservative refusal)
ok - an ancestor-only anchor never binds an advanced parked run to this task
ok - a terminal unfetched-head row is stale history and never concludes a run
ok - a terminal newest row anchored at this worktree's head never authorizes an abort
ok - a resolvable diverged newer same-branch row makes every older row stale history; no run is concluded
ok - consecutive unresolvable rows are ambiguous and never conclude a run
ok - a ledger-proven continuation is still left alone while the run is autonomously active
ok - teardown refuses before reap or removal when a task-owned run remains parked
ok - a different run cannot confirm the targeted abort
ok - empty post-abort status is not accepted as confirmation
ok - the CLI's exact run-not-found signal confirms completion
ok - a parked run on another branch is never aborted by this task's teardown (ownership is precise)
ok - a task-owned autonomous running step is left alone rather than aborted
ok - a leaked descendant process rooted under the task's worktree is reaped by teardown, not left surviving
ok - a leaked descendant process rooted under the task's per-task tasktmp is reaped by teardown too
ok - missing lsof falls back to reaping the tmux pane process group
ok - an erroring lsof scan refuses teardown and preserves the task
ok - a reused pid with a different start time is never force-killed
ok - an exec change preserves birth identity and the process is reaped
ok - a process spawned during grace is reaped on a later pass
ok - persistent leaked processes refuse teardown after bounded retries
ok - a process exiting during identity lookup does not block teardown
ok - the run abort and the leaked-process reap both complete before the destructive worktree return
exit=0
- Outcome: 🔧 3 issues found → auto-fixed ✅ across 2 runs (1h2m24s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ℹ️ .gitignore:8 - .gitignore now lists .omc/ twice (lines 6 and 8). The second entry is a redundant leftover from the merge and changes nothing; remove one.

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Test** - 3 issues found → auto-fixed ✅
  • ⚠️ bin/fm-session-start.sh - Merging the fork's bounded-herdr-CLI probes (fix(bin): bound herdr CLI probes so a hung read cannot wedge a supervisor #4988) with the session-start endpoint bound broke tests/fm-session-start.test.sh. The backend CLI runs in its own process group under a 10s bound, so when session start's per-read bound (for example 2s) killed its own group, a hung herdr survived up to 10s. The test's fake herdr also killed the wrong ancestor because the bounded wrapper added extra process hops. Fixed both: bin/fm-session-start.sh now caps FM_BACKEND_HERDR_CLI_TIMEOUT at the per-read bound, and the deadly-read fake in tests/fm-session-start.test.sh walks ancestry to the read's own shell. The full session-start file now passes.
  • ℹ️ tests/fm-watch-triage.test.sh - tests/fm-watch-triage.test.sh ran 101 passing assertions with no failures but did not finish inside the 900s cap. I could not confirm full completion.
  • ⚠️ live validation verdict: inconclusive (0 of 7 scenarios were driven live against the product); untested: Session start digest survives a killed endpoint read and reports it as that task's error line, A hung herdr endpoint read is bounded by the per-read timeout and leaves no stray herdr process, Teardown reclaims tasks, retires watcher markers, and refuses unsafe slot ownership, Watcher re-arm and stop work, A hung herdr CLI call is bounded and reaped, Watch triage classifies wedged panes without stalling, Live Herdr lab: real primary starts, spawns a lane, and tears it down
  • Live validation: ⚠️ inconclusive - 0 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Session start digest survives a killed endpoint read and reports it as that task's error line ⏸️ untested no The prior payload only ran tests/fm-session-start.test.sh and did not drive this against a live running product, so no live pass is established.
A hung herdr endpoint read is bounded by the per-read timeout and leaves no stray herdr process ⏸️ untested no The prior payload only ran the hung-read and padded-zero-bound unit tests, not a live herdr, so no live result is established.
Teardown reclaims tasks, retires watcher markers, and refuses unsafe slot ownership ⏸️ untested no The prior payload only ran tests/fm-teardown.test.sh and tests/fm-teardown-endpoint-safety.test.sh, not a live teardown, so no live result is established.
Watcher re-arm and stop work ⏸️ untested no The prior payload only ran tests/fm-watch-arm.test.sh, not a live watcher, so no live result is established.
A hung herdr CLI call is bounded and reaped ⏸️ untested no The prior payload only ran tests/fm-backend-herdr-probe-timeout.test.sh, not a live herdr CLI, so no live result is established.
Watch triage classifies wedged panes without stalling ⏸️ untested no The suite is long-running and hit the 900s cap. A longer timeout on an idle host is needed to see it finish.
Live Herdr lab: real primary starts, spawns a lane, and tears it down ⏸️ untested no No live Herdr lab was driven in this run. Running one needs bin/fm-herdr-lab.sh with a named fm-lab-* session on a host where herdr can be stood up.
  • bash tests/fm-session-start.test.sh (failed on the merge, passes after the fix; the hung-read and padded-zero-bound tests passed in isolation)
  • bash tests/fm-teardown.test.sh
  • bash tests/fm-teardown-endpoint-safety.test.sh
  • bash tests/fm-backend-herdr-probe-timeout.test.sh
  • bash tests/fm-watch-arm.test.sh
  • bash tests/fm-watch-triage.test.sh (101 ok, no failures, hit the 900s cap)
  • checked the base commit a774c448 in a temporary worktree: session-start had no failures there, so the failure came from this change

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 2 of 6 scenarios driven live against the product
Scenario Result Live Evidence
A hung or killed herdr endpoint read during session start is bounded, reported as that task's error line, and leaves no stray herdr process ⏸️ untested no The prior payload only ran the repo's test with a fake herdr, not a real herdr server or live session start, so it did not establish a live result. To cover it, run session start against a real primar…
A hung herdr CLI call is bounded and reaped by the shared read owner, while the long-lived server launch stays exempt ⏸️ untested no The prior payload only ran the repo's test with a fake hung herdr, so it did not establish a live result. To cover it, drive a real herdr CLI call that hangs.
Teardown reclaims tasks, retires watcher markers, and refuses unsafe slot ownership ⏸️ untested no The prior payload only ran fixture-driven repo tests, not a live fleet, so it did not establish a live result. To cover it, run teardown against a live fleet with real tasks.
Watcher arm, re-arm and recovery keep wakes durable and quarantine malformed state ⏸️ untested no The prior payload only ran a fixture-driven repo test, not a live fleet, so it did not establish a live result. To cover it, arm and re-arm a real watcher on a live fleet.
Live Herdr lab: provision an fm-lab-* session, run a herdr command, and tear it down ✅ pass live herdr-lab-live.log
Real Herdr e2e: session cleanup and respawn idempotency ✅ pass live fm-herdr-session-cleanup-e2e.log, fm-backend-herdr-respawn-idem-e2e.log
  • tests/fm-session-start.test.sh (exit 0)
  • tests/fm-backend-herdr-probe-timeout.test.sh (exit 0)
  • tests/fm-teardown-endpoint-safety.test.sh (exit 0)
  • tests/fm-teardown.test.sh (107 ok, exit 0)
  • tests/fm-watch-arm.test.sh (exit 0)
  • tests/fm-herdr-lab.test.sh (exit 0)
  • tests/fm-herdr-session-cleanup-e2e.test.sh (real Herdr, exit 0)
  • tests/fm-backend-herdr-respawn-idem-e2e.test.sh (real Herdr, exit 0)
  • bin/fm-herdr-lab.sh provision, run workspace list, then teardown on a fm-lab-gate-* session (all exit 0)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

keenvc and others added 30 commits September 16, 2026 16:15
…spatch (#1)

* feat(harness): add the cline crewmate/scout adapter with ClinePass dispatch

Add Cline CLI 3.0.62 as a verified crewmate/scout harness following the agy
pattern: ancestry detection on the native .cline process, a launch-then-send
TUI launch, per-task .cline/hooks busy/turn-end wiring under a new cline-hook
busy source, Escape interrupt, /exit, and control-plane tables. Wire the
ClinePass open-weights pool into the crew-dispatch example and document the
known composer-empty gap (placeholder luminance above the shared ghost
ceiling) with a tmux live guard as the refresh command.

* no-mistakes(review): fix(docs,quota): correct cline resume grouping and cline-pass family id

* no-mistakes(document): docs: cover cline in tmux liveness/anchoring list and configuration.md secondmate-refusal note

* no-mistakes(lint): {"summary": "lint: no code changes needed, fm-lint.sh passes with shellcheck on PATH"}

* chore(gitignore): drop the stray .omc handoff artifact and ignore .omc/

The no-mistakes gate agent's Claude Code oh-my-claudecode plugin writes
.omc/handoffs/last-session-end.md into the run worktree at session end, and a
later pipeline step committed it into this branch. Remove the committed file and
ignore .omc/ so a home-environment handoff artifact can never ride into a PR.

---------

Co-authored-by: firstmate-worker <worker@local>
Two Claude subscriptions need to run concurrently across lanes without
moving every claude spawn onto one account. --claude-config-dir picks the
CLAUDE_CONFIG_DIR one claude spawn's pane resolves into: validated before
any worktree or endpoint exists, recorded in the task's own meta, reused
unchanged on --relaunch, and threaded through both the pre-launch trust
registration and the launch's own environment so the two halves can never
land in different stores. A spawn naming no seat is byte-identical to
before.
Count live crewmate, scout, and local secondmate lanes grouped by the
billing provider a candidate actually draws on, expose the load for
dispatch intake, and refuse a spawn that would push a provider past its
configured cap.

Provider identity comes from the resolved model string, not the harness
name: a provider-qualified prefix or a model-id pattern decides the pool,
and the harness table is only the fallback. That keeps two models on one
pool counting together while a different pool stays separate. The mapping
lives once in bin/fm-provider-lib.sh, whose fallback reuses the existing
quota tables rather than restating them.

A lane occupies a seat unless its recorded endpoint is provably dead or
missing, so a cleared seat is visible before the next dispatch. The cap
comes from providerCaps in config/crew-dispatch.json (per provider, else
default), falling back to 4; bootstrap now rejects a malformed
providerCaps instead of silently ignoring it.

bin/fm-provider-load.sh prints the current per-provider used/cap for
intake. Stranded-record detection stays with fm-lane-account-dead-records;
this counter reads the current endpoint classifier.
Deliver bin/fm-hold-reverify.sh, an armed watcher check that re-checks each
captain hold past an age threshold against shipped reality and reports it as
dead, still_live, not_a_decision, or unestablishable - the reconciliation
vocabulary captain-hold-lifecycle already owns.

It reports only: it never calls answer and never closes or annotates a call,
so only the captain's own words or an explicit evidence-backed reconciliation
can resolve one. Dead is never inferred from absence or an unreadable source.

Aged holds come from the canonical local backlog projection
(fm-fleet-snapshot.sh --contribution-input); recorded pull requests are read
through fm-pr-lib.sh. Each sweep writes a docket and prints one line only when
the finding set changes, with a report record keyed on that set.
…rget

A firstmate home had no way to see that another home was already working a
shared external target, so the main home and a secondmate could both arm to
land the same PR with nothing to stop a double merge.

bin/fm-claim.sh records, releases, and inspects a work claim on a shared
external target - a PR, an issue id, or a declared file area. The store is a
machine-wide directory (FM_CLAIM_ROOT, default
${XDG_STATE_HOME:-$HOME/.local/state}/firstmate/claims), the sibling of the
existing process-event source claim root, because one owner per canonical
target cannot live inside a single home. Local homes share one filesystem; a
remote secondmate is a separate host and stays outside the mechanism.

Acquire is atomic and fails closed: a second home's live claim refuses rather
than racing. A claim is released on cleanup or reclaimed only when its holder
is provably gone (its home directory is absent, or its task record is absent
past FM_CLAIM_PENDING_GRACE). Any uncertainty keeps the claim.

bin/fm-spawn.sh --claim records the claim before any endpoint or task record
exists and refuses the spawn on conflict, recording the canonical keys on the
task as claims=; fm-teardown.sh releases them on cleanup. The flag is refused
on --secondmate, --relaunch, and a batch dispatch.

Tests: tests/fm-claim.test.sh drives the real CLI across two simulated homes
sharing one claim root, plus a real spawn that records its claim and a second
dispatch that is refused.
Merge authority was decided once at intake and never revisited, so a task
dispatched yolo=on kept that authority even after firstmate held it for the
captain, and the recorded authority and the merge path could disagree.

Fold the captain-hold predicate into fm_merge_authority_resolve, the single
owner of a task's standing merge authority, so a held task resolves to
captain-hold (or hold-unreadable) whatever its yolo posture or away grants say.
Remove bin/fm-pr-merge.sh's duplicate require_released_captain_hold and fold its
refusal into the shared gate, and have bin/fm-merge-local.sh share the same
predicate instead of repeating it.
A worker parked on a provider quota wall kept a live, painting harness
while its turn could not advance, so every one of them read as working
from its semantic busy record. The measured fleet incident had all of
one provider's workers stalled at the same weekly limit while
supervision saw a healthy fleet.

Recognize the wall from the pane text the busy reader already inspects.
The signal is built from two independent rendered families - a
wall-shaped limit phrase and a scheduled retry/reset phrase - within the
last few non-empty lines, so no single vendor string is load-bearing and
ordinary worker prose does not match. A busy verdict over that wall
reports `quota` instead of busy.

fm-crew-state.sh surfaces it as its own `state: quota` rather than
collapsing it into working or a declared pause, because a quota-killed
worker cannot be relaunched in place; the recovery skill now states that
preserve-and-replace under a new id is the path.

The portable regression pins the logic and its divergence cases over
synthetic transcripts. The live guard drives the real installed OpenCode
TUI against a local 429 stub so its own retry modal renders with no
model tokens spent, and proves the same task reads working before the
wall and quota after it.
Teardown refused any record whose endpoint was already cleared, so a lane
could never be retired once its window was gone and it kept occupying an
in-flight row. Accept an explicit endpoint_cleared stamp as stronger
agent-less evidence than a dead window, with no flag and no --force, while
keeping the unlanded-work refusal unchanged.

A projected Herdr teardown confirmed only the task pane was gone, so a
workspace whose recorded pane vanished before its close survived for a
restart to restore as a live agent in the primary clone. Remove the
workspace's remaining panes through the same focus-preserving pane close and
require the workspace gone before retiring the journal.

A dead pane whose display redrew re-alarmed on every new hash, a supervision
tax that grew with each dead lane. Absorb a redrawn dead display against the
existing once-record, while a relaunched agent re-arms the incarnation and
its own death still reports in full.
keenvc and others added 21 commits September 20, 2026 02:22
… alongside the reliability batch

# Conflicts:
#	bin/fm-bootstrap.sh
#	bin/fm-control-lib.sh
#	bin/fm-quota-choose.sh
#	docs/configuration.md
#	docs/examples/crew-dispatch.json
… fork merge

Committed only to let a clean merge proceed without touching this work;
not otherwise reviewed or authored by this session.
…s the preserved cline adapter

# Conflicts:
#	AGENTS.md
Make the Fireworks DeepSeek dispatch lane usable: openhands is now a
verified crewmate/scout harness with launch, readiness, busy, interrupt,
and exit mechanics, so config/crew-dispatch.json no longer fails as an
unverified adapter.
# Conflicts:
#	.agents/skills/harness-adapters/SKILL.md
#	AGENTS.md
#	bin/fm-agent-process-lib.sh
#	bin/fm-bootstrap.sh
#	bin/fm-busy-lib.sh
#	bin/fm-composer-lib.sh
#	bin/fm-control-lib.sh
#	bin/fm-harness.sh
#	bin/fm-spawn.sh
#	docs/agent-control.md
#	docs/architecture.md
#	docs/configuration.md
#	docs/trace-context.md
#	docs/verification/runtime-backends.md
#	tests/fm-control.test.sh
# Conflicts:
#	.gitignore
#	bin/fm-spawn.sh
* fix(bin): read aged holds via backlog-json, not contribution-input

The hold re-verify sweep only needs canonical backlog rows; contribution-input
also walks every task meta for merge-authority resolution, which took ~96s at
this fleet size and always exceeded the five-second projection bound. Add
fm-fleet-snapshot.sh --backlog-json for that narrower read and point the sweep
at it so failures stay loud without raising the timeout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): apply review decisions for aged hold re-verify sweep

Sort aged captain holds oldest-first before the per-sweep cap, let
FM_HOLD_REVERIFY_BUDGET_SECS govern the backlog projection bound, clamp
forge probes to remaining budget, skip probes for predetermined
not-a-decision rows, drop the classify subcommand, and document
--backlog-json.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): use backlog_json output mode label for shellcheck

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs: fold pipeline document-step hold-reverify pointers into fork tip

Adds the toolbelt row and prose alignment from the failed run's document
step without rebasing onto upstream main.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(document): correct hold-reverify docket contents description

* chore: gitignore .omc session state and document in AGENTS.md

Apply captain inbox 006 on the fork publication branch without rebasing onto upstream main.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: expect 19 snapshot/fleet-view tests under stock macOS Bash

The fork's tests/fm-fleet-snapshot-view.test.sh carries
test_large_payloads_compose_through_files, the regression for
bin/fm-fleet-snapshot.sh composing payloads above 128KB through files
instead of argv. It landed in d9356ca together with that snapshot
change and passes under /bin/bash 3.2 on the macOS runner, which
already counted 19 ok lines, so the hard-coded 18 in the macOS job was
the only thing left behind.

* docs: declare the cline-pass provider on the documented cline profiles

The resolver refuses docs/examples/crew-dispatch.json with "profiles
whose harness lacks one authoritative provider family require
provider: cline", so the documented-example check in
tests/fm-dispatch-resolve.test.sh has failed since the cline profiles
were added to the example.

The example is the wrong side. The resolver's provider is the quota-axi
provider family it ranks candidates by, not the launch prefix cline
reads from its model id, and the single-provider table in
docs/configuration.md leaves cline out on purpose: the model prefix,
not the harness, decides who bills the run, exactly as for pi and omp.
The Pi default in the same example already declares provider: claude
for that reason. Add provider: cline-pass to the four cline profiles,
correct the one sentence in docs/configuration.md that claimed no field
was needed, and give the test's canned Choice answer the fourth rule
the example now has, since the resolver checks the answer against the
full option set.

* test: corrupt the claim record under test, not the first one find returns

The corrupt-record check picked its victim with find | head -n 1 while
three claims exist, so on a filesystem whose directory order differs
from the author's it corrupted o/r#7 or repos/example#9 and then asked
about owner/repo#11, whose intact record answered "held" with exit 0.
CI's stdout showed exactly that line. Select the record by its
documented key= line instead. The guard itself is intact: corrupting
the right record returns exit 5 on the same inputs.

* test: give the restart watchers a refresh bound a slow runner can meet

In the watcher-restart section of tests/fm-home-summary-refresh.test.sh
the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of
reports 999999 for a missing ledger, so its detached refresh fires on
every one-second poll. After the lock holder is killed that refresh
steals the dead lock ahead of the test's idle-only refresh, which then
returns without publishing, and with the section's
FM_HOME_SUMMARY_TIMEOUT=2 a slow runner kills the watcher's attempt
before it publishes. The next poll dies the same way and the ledger
never appears, which is the "a dead publication lock wedged
publication" failure in CI run 35683307194.

Raise the three restart watchers' bound to 30 seconds, the deadline
this file already uses for its accumulated-home publication. Under a
30% CPU quota the unchanged test fails on exactly that line and the
changed one passes all 21 checks. The idle-only refresh, the dead-lock
reclamation and the 10-second wait are untouched.
…wner (adopt kunchenguid#5993) (#7)

* Fix reassigned teardown slot collisions

(cherry picked from commit 115333b)

* no-mistakes(document): Document claim-first pool-slot teardown behavior

(cherry picked from commit dcbe51a)

* no-mistakes(document): Confirm teardown documentation reflects slot ownership behavior

(cherry picked from commit bdcedcf)

* fix(teardown): avoid presentation-lock races after claim-first slot cleanup

Reassigned-slot teardown with a dead Herdr husk no longer takes the shared
presentation session lock, restoring the pre-5993 contention profile for stale
records while keeping live-slot protection intact. Herdr presentation recovery
spawns now wait up to 120s for that lock and release it on abort so concurrent
cross-home recovery cannot fail the 5s try loop after a legitimate holder.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: QIanGua <15757826110@163.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…in config (#9)

* Fix OpenCode v2 worker launch to use --standalone and config model.

OpenCode 2.x removed the interactive --model flag; carry the resolved model in OPENCODE_CONFIG_CONTENT and launch with --standalone so the model and permission block are honored off the shared service.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(bin): honor opencode v2 top-level model and gate --standalone

Always write the resolved model as OPENCODE_CONFIG_CONTENT top-level model on 2.x, drop unverified agent.build variant JSON, and keep the 1.x --model launch shape when opencode --version reports major 1.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test: stub opencode --version in shared spawn fakebin

Spawn tests prepend a fakebin to PATH; fm-spawn now probes opencode
--version for the v1/v2 launch gate, so every spawn fakebin must answer it.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Brings the fork up to date with upstream, including the watcher stall and
re-arm fixes, teardown retiring watcher markers, orphan journals and wake
rows, herdr endpoint reclaim, the shared secondmate liveness library, the
worker account pin, the Devin adapter, and the launch-prompt backstop.

Conflict policy: upstream's version wins wherever it fixes or supersedes a
fork change; the fork's adapters, features and fixes that upstream does not
ship are kept on top of it.

- Merge authority: upstream's away-record model replaces the fork's
  captain-hold-revokes-yolo change, so that change and its tests are dropped.
  Upstream's pr-merge and merge-local already refuse a held task.
- The WIP commit of uncommitted local edits is not carried over.
- OpenCode: 1.x keeps --model and the effort variant from upstream; 2.x keeps
  the fork's top-level model plus --standalone launch.
- A per-lane Claude config-dir seat and the per-home worker account pin both
  choose the Claude store, so fm-spawn refuses the combination.
- Definition of done: the evidence-pair block is rendered by a wrapper around
  upstream's reworked per-forge blocks.
…m merge

Merge took upstream's --arg jq transport and dropped the fork's --rawfile /
--slurpfile staging for >128KB status folds, crew-state detail, secondmate
row composition, and contribution-input. CI failed with jq Argument list too
long on portable serial 8 and Stock macOS Bash. Port the fork transport onto
the merged schema while keeping upstream age_seconds / observed_age.
@keenvc

keenvc commented Sep 30, 2026

Copy link
Copy Markdown
Author

Closed: this PR is intended for the downstream fork keenvc/firstmate (keenvc#11), created against upstream by mistake during automated sync.

@keenvc keenvc closed this Sep 30, 2026
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 0/5

[Medium risk] Merges upstream changes and adds new agent harness adapters.

The PR is not safe to merge until claim exclusivity, provider-cap admission, and hold-sweep coverage are corrected.

Reviews (1) · Last reviewed commit: "no-mistakes(document): Document herdr CL..."

Comment thread bin/fm-claim.sh
Comment on lines +302 to +304
fm_claim_read "$path" || continue
if [ "$FM_CLAIM_HOME" = "$HOME" ] && [ "$FM_CLAIM_TASK" = "$TASK" ]; then
if rm -f -- "$path"; then

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Cleanup can delete another claim
When teardown or an aborted spawn releases a task's claims, release-task checks each claim and removes its file without taking the per-target lock. If another home reclaims a stale claim between those steps, cleanup deletes the new home's live claim. A third home can then claim the same target and dispatch duplicate work.

Comment thread bin/fm-claim.sh
Comment on lines +283 to +288
if [ "$FM_CLAIM_HOME" != "$HOME" ] || [ "$FM_CLAIM_TASK" != "$TASK" ]; then
echo "error: release refused - $KEY is held by home $FM_CLAIM_HOME for task $FM_CLAIM_TASK, not by home $HOME for task $TASK" >&2
return 3
fi
begin_mutex "${PATH_CLAIM}.lock.d"
rm -f -- "$PATH_CLAIM" || die "could not remove the claim at $PATH_CLAIM"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Release uses stale ownership
release checks who owns a claim before taking the per-target lock, then removes the file without checking again. If another home replaces a stale claim in between, release deletes that home's live claim and allows conflicting work to be dispatched.

Comment thread bin/fm-claim-lib.sh
Comment on lines +262 to +267
while [ "$path" != "${path#./}" ]; do path=${path#./}; done
path=$(printf '%s' "$path" | tr -s '/')
path=${path#/}
while [ "$path" != "${path%/}" ]; do path=${path%/}; done
[ -n "$path" ] || return 1
printf 'area:%s:%s\n' "$project" "$path"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Equivalent areas get different claims
Area normalization leaves internal .. segments intact. area:project:src/../docs and area:project:docs name the same directory but get different claim files, so two homes can claim and work that area without detecting the conflict.

Comment thread bin/fm-claim-lib.sh
Comment on lines +142 to +154
if [ -e "$FM_CLAIM_HOME/state/$FM_CLAIM_TASK.meta" ] || [ -L "$FM_CLAIM_HOME/state/$FM_CLAIM_TASK.meta" ]; then
return 1
fi
case "$FM_CLAIM_CREATED" in
'' | *[!0-9]*) return 0 ;;
esac
now=$(date +%s) || return 1
grace=$(fm_claim_pending_grace)
case "$grace" in
'' | *[!0-9]*) grace=300 ;;
esac
age=$((now - FM_CLAIM_CREATED))
[ "$age" -ge "$grace" ]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Active spawns lose claims
Spawn acquires a claim before setup but publishes the task record much later. If setup takes longer than the 300-second grace period, this check treats the missing record as proof that the claim is stale, even though spawn is still running. Another home can take the claim, leaving both homes free to dispatch against the same target.

Comment thread bin/fm-spawn.sh
Comment on lines +2614 to +2626
if [ -n "$HARNESS" ]; then
LANE_CAP_MODEL=$MODEL
LANE_CAP_EXCLUDE=
if [ "$RELAUNCH" -eq 1 ]; then
LANE_CAP_EXCLUDE=$ID
[ -n "$LANE_CAP_MODEL" ] || LANE_CAP_MODEL=$(fm_meta_get "$RELAUNCH_META" model)
fi
fm_provider_cap_refuse "$STATE" "$CONFIG" "$HARNESS" "$LANE_CAP_MODEL" "$LANE_CAP_EXCLUDE" || exit 1
fi
if [ "$HARNESS" = openhands ]; then
if [ -z "$MODEL" ] || [ "$MODEL" = default ]; then
if [ -n "${LLM_MODEL:-}" ]; then
MODEL=$LLM_MODEL

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 OpenHands checks wrong provider cap
When an OpenHands spawn gets its model from LLM_MODEL rather than --model, the cap check runs before that model is assigned. Admission checks the openhands bucket, but the task is later recorded and counted under the model's provider. These spawns can therefore exceed that provider's configured cap.

Comment thread bin/fm-hold-reverify.sh
Comment on lines +418 to +420
if [ "$EXAMINED" -ge "$MAX_HOLDS" ] || budget_exhausted; then
DEFERRED=$((DEFERRED + 1))
continue

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Deferred holds never get checked
Every sweep sorts aged holds oldest-first and checks only the first MAX_HOLDS, without saving a position for the next sweep. If more than the default 12 aged holds remain, each sweep checks the same 12 and defers the rest indefinitely. Those holds never receive a verdict in the docket.

Comment thread bin/fm-spawn.sh
Comment on lines +5583 to +5586
{
printf 'LLM_API_KEY=%s\n' "$(shell_quote "$OPENHANDS_API_KEY")"
printf 'LLM_MODEL=%s\n' "$(shell_quote "$MODEL")"
} > "$OPENHANDS_ENV_FILE" || {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 security Failed spawn retains API key
OpenHands writes LLM_API_KEY to a per-task env file before checking whether the pane starts processing its brief. If that readiness check fails, spawn closes the pane and exits without removing the file. The credential remains in state/ until a separate teardown, so abort cleanup should remove it.

How this was verified: The readiness-failure path reaches endpoint and abort cleanup, neither of which removes the secret-bearing env file.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant