Skip to content

fix(bin): sync upstream batch 16 — test-run refusal, captain-hold bounding, Pi provider resolution - #50

Merged
zeeshaanahmad merged 6 commits into
mainfrom
fm/upstream-batch-16-to-tip
Sep 7, 2026
Merged

zeeshaanahmad merged 6 commits into
mainfrom
fm/upstream-batch-16-to-tip

Conversation

@zeeshaanahmad

Copy link
Copy Markdown
Owner

Intent

Captain's standing policy for this fork (2026-09-05, verbatim): "the main goal for firstmate is to become sync with upstream. once we reach that state, we avoid introducing more changes and stay aligned with the upstream. until then you can keep adding work that helps aligning with the upstream easier and reaches the intended state quickly."

Captain's conflict rule (2026-09-05, verbatim): "if there is something upstream and our code both worked on and fixed but approaches differ, take the upstream change, however, if our change is better than upstream, file a PR. for other items, take the upstream changes."

Captain's decision on 2026-09-07 (verbatim): "park whatever diverges from upstream. our goal was to sync with upstream for firstmate. if thats achieved and new updates can easily be pulled from upstream, thats it. no further work on firstmate." Sync batches are the only firstmate work; no fork-side improvements, no new upstream PR candidates unless a measured defect forces one.

The ask this task serves: upstream sync batch 16 - bring the fork's origin/main (now bb429c7: batch 15 landed as #49, waypoint 5592cb6) up to upstream's current tip, waypoint 0b9f518 ("feat(pi): resolve extension-registered providers in the supervision branch (kunchenguid#3871)"). Three upstream commits in range 5592cb6..0b9f518, 17 files, +783/-48:

What Changed

  • bin/fm-spawn.sh now exports FM_TASK_ID to ship/scout panes over the same channel as GOTMPDIR, and bin/fm-test-run.sh refuses to run any executing test mode when that marker is set and the runner resolves to the repository's primary checkout (comparing git-dir against git-common-dir), while inspection-only modes remain unaffected; tests/lib.sh gained the runner's library needed by the new tests/fm-test-run.test.sh coverage.
  • bin/fm-captain-hold.sh open gained an --identity flag that prints a call's lifecycle identity (hold-set timestamp + resolution-record count) on an open captain hold, and bin/fm-watch.sh uses it to bound repeated stale-wait alarms per distinct identity/status hash within the resurface window, covering both plain stale waits and captain-held waits; bin/fm-teardown.sh and doc references were updated to match, and tests/fm-watch-triage.test.sh / tests/fm-backend-herdr-presentation-e2e.test.sh were expanded and isolated accordingly.
  • .pi/extensions/fm-branch-supervision.ts captures main's ModelRegistry alongside its model, and copies extension-registered providers (e.g. a runtime-registered "devin" provider) into the isolated branch ModelRuntime when a model can't otherwise be resolved, refreshing auth for the copied providers so branch model resolution and availability checks see them; covered by the new tests/fm-pi-branch-extension.test.sh.
  • Updated docs/architecture.md, docs/configuration.md, docs/scripts.md, docs/verification/trace-context.md, and docs/pi-supervision-branch.md to describe the new task-marker refusal, captain-hold identity/bounding behavior, and extension provider resolution.

Risk Assessment

✅ Low: Clean upstream sync; the one real merge conflict in bin/fm-watch.sh was independently traced and does not reintroduce the original bug it appears to remove protection for.

Testing

Two of the three targeted suites for batch 16 (primary-checkout refusal and Pi extension provider resolution) ran to completion with all cases passing, directly demonstrating the fixed/added behavior end-to-end via real fixture repos and worktrees; the third suite (stale-alarm bounding for captain holds) was still executing — with all observed assertions passing — when this report was required, so its final pass/fail status is not yet confirmed.

Evidence: fm-test-run.test.sh full transcript (primary-checkout refusal, 34/34 passed)

Source: fm-test-run.test.sh full transcript (primary-checkout refusal, 34/34 passed)

FM_TEST_BEGIN 2026-09-07T22:18:22Z tests/fm-test-run.test.sh family=pure-contract-unit expected_gate_skip=none
ok - exact suite coverage: --all lists every tests/*.test.sh once
ok - family selection returns a proper subset of the suite
ok - single-script selection lists exactly that path
ok - changed-file selection stays conservative (never silent full suite)
ok - a task marker refuses execution in the primary checkout and leaves worktrees and inspection alone
ok - runner and its documentation surfaces select their curated family, not just their contract owners
ok - shell line-ending policy selects runner coverage
fm-test-run: no tests selected for changes vs HEAD (map is conservative; use --all for the complete suite)
ok - changed selection covers dependents, fails closed for live unmapped source, and accepts retired unconsumed source
ok - a bin reference selects the referencing scripts, and consumers still select their curated families
ok - changed defaults to bounded automatic scheduling with serial override
ok - Windows emulation exempts only synthetic POSIX modes
ok - a plain script list defaults to bounded automatic concurrency without an automatic timeout
ok - family proofs run concurrently only within separate family phases
ok - empty changed selection emits deterministic text and JSON summaries
ok - timing markers and JSON artifact are valid
ok - aggregate exit reflects any script failure
ok - gate-skip accounting is honest and non-failing
ok - a gate skip records why it skipped
ok - a script that actually ran records no skip reason
ok - live guards are recorded as a capability class, not a bare env opt-in
ok - fail-on-gate-skip converts herdr-not-found into a hard failure
ok - exclude-family drops the named primary family after selection
ok - portable shard union, disjointness, and coverage guard hold
ok - portable serial shards are a deterministic disjoint cover of the serial lane
ok - coverage guard reports and bounds the unmeasured portable serial share
ok - portable serial shard lanes refuse mismatched, out-of-range, and countless names
ok - --jobs refuses non-proven / stateful selections
ok - --jobs admits and schedules a family with a recorded concurrent proof
ok - an unclassified new test stays serial while the proven residual family runs concurrently
ok - a concurrent run starts the longest-hint script first
ok - --per-script-timeout-secs turns a hung script into a bounded failure
ok - --max-wall-ms fails an over-budget run and refuses a malformed budget
ok - jobs scheduler runs proven scripts; failure propagates; non-proven refused
ok - both lanes scrub the same inherited fleet-home overrides and keep the rest of the environment
ok - Herdr CI family-run step times out at 20 min under a 75 min job backstop
ok - aggregate-json merges lane timing artifacts
FM_TEST_END 2026-09-07T22:23:13Z tests/fm-test-run.test.sh exit=0 duration_ms=290209 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=290675
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=290209 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-test-run.test.sh duration_ms=290209
fm-test-run: wrote timing artifact: /tmp/fm-test-run-result.json

[exited with code 0]
Evidence: fm-pi-branch-extension.test.sh full transcript (extension provider resolution, 42/42 passed)

Source: fm-pi-branch-extension.test.sh full transcript (extension provider resolution, 42/42 passed)

FM_TEST_BEGIN 2026-09-07T22:20:28Z tests/fm-pi-branch-extension.test.sh family=standalone expected_gate_skip=none
ok - fm_branch_outcomes hides through ToolExecutionComponent while Calm-off and HTML export stay stock
ok - the installed Pi still bounds the picker's list and ranks its search
ok - branch owns accepted wakes with a stable prefix and deterministic verdict-driven delivery
ok - a captain outcome reaches main's model as one typed, sequence-keyed processing request while routine notes stay plain
ok - requested and unsolicited healthy outcomes keep distinct delivery and event ownership
ok - captain outcomes are exact and exactly once across crash, reload, busy main, compaction, and an unrelated assistant response
ok - a captain outcome opens one sequence-keyed processing turn, survives empty and unrelated answers, is re-presented at run end and session start, and closes only on its acknowledgement
ok - scopeForUnreadWake excludes every main-only class without vetoing eligible task-local rows, and writes the eligible snapshot
ok - branch prompt_cache_key is stable per home across sessions and distinct between homes
ok - branch default-on eligibility (task-scoped, heartbeat, afk) binds and a broken branch rejects to watcher fallback
ok - a heartbeat review survives a check row arriving before its drain
ok - fm_branch_report refuses a task the wake did not name, fleet included, while a heartbeat is unscoped
ok - pre-drain eligibility re-check excludes a newly main-owned row without deferring eligible work
ok - a co-present needs-decision row neither vetoes nor falsely settles routine branch delivery
ok - a settled branch turn without a durable outcome falls back and releases its grant for main replay
ok - provider-error latches cool down, re-probe once with backoff, and recover through a durable report
ok - selection changes preserve in-flight transcript ownership and reset provider-error streaks
ok - a stale main claim returns the durable wake to watcher delivery
ok - pre-drain eligibility re-check no-ops an already-drained wake
ok - dialog mirror filters tool and operational traffic, lands before wakes, and keeps a durable cursor
ok - the dialog mirror re-anchors for each session's new branch conversation and stays incremental within it
ok - every main session start begins a new branch conversation while one session keeps its own
ok - the current pin state binds every branch build, and clearing it returns the branch to main's model
ok - unpinned branches follow main model changes live while pinned branches stay fixed
ok - supervision-model command persists the captain's pick and rebinds the live branch
ok - supervision-model opens a bounded searchable list, follow main first, and pins the branch alone
ok - branch model picker keeps follow main first and filters the eligible catalog
ok - the effort pin binds every branch build, and clearing it returns the branch to main's effort
ok - unpinned branches follow main effort changes live while pinned branches stay fixed
ok - an extension-registered provider resolves in the isolated branch runtime
ok - supervision-model runs an effort picker after the model picker and persists both independently
ok - an unusable model pin rejects to watcher fallback and an unparseable one is treated as no pin
ok - replacement activation cleans old branch leases and retries failed cleanup
ok - branch activates on a cold start once the lock is acquired, never before
ok - queued wakes and mirrors stop mutating branch state after lock ownership is lost
ok - stale reports, shells, mirrors, cursors, leases, and prompts perform no side effects
ok - a Pi session that does not own the lock accepts nothing and mutates no branch state
ok - an extension rebind re-mirrors undelivered dialog instead of dropping it
ok - outcome delivery keeps the event loop running and interleaved reports stay ordered and exactly once
ok - a session replaced mid-delivery cancels cleanly and the stored outcome still arrives exactly once
ok - a failing store script surfaces to the branch and its outcome is neither lost nor delivered twice
ok - a failed cursor write re-delivers a routine note exactly once more while a captain outcome stays deduplicated
FM_TEST_END 2026-09-07T22:25:10Z tests/fm-pi-branch-extension.test.sh exit=0 duration_ms=281686 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=282205
FM_TEST_SUMMARY_FAMILY family=standalone count=1 duration_ms=281686 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-pi-branch-extension.test.sh duration_ms=281686
fm-test-run: wrote timing artifact: /tmp/fm-pi-branch-result.json

[exited with code 0]
Evidence: fm-watch-triage.test.sh partial live transcript (stale-alarm bounding, in progress at report time)

Source: fm-watch-triage.test.sh partial live transcript (stale-alarm bounding, in progress at report time)

ok - status_span_has_actionable: benign absorbed, captain events surfaced, classified events not re-fired
ok - an actionable event is not hidden by later routine appends, and is named as itself
ok - span classification retires closed decisions and surfaces rejected transitions for reconciliation
ok - a malformed seen signature causes the whole status log to be classified
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, paused is not captain-relevant, and the two declared-wait verbs stay separable
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - crew_worktree_written_since: real writes are evidence; no worktree, no anchor, quiet trees, .git churn and a mate's own home are not
ok - an empty FM_WORKTREE_WRITE_PRUNE widens the probe to the whole depth-bounded tree instead of disabling it
ok - an empty FM_WORKTREE_WRITE_PRUNE exported into the environment prunes nothing, widening the probe
ok - the worktree write probe is wall-clock bounded, and hitting the bound reads as no write evidence
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a secondmate's status signal is never absorbed as provably working; crewmates are unaffected
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a bare turn-end from a pane that churned since the previous poll is absorbed
ok - pane churn starts a fresh stale-classification interval before a stopped render returns
ok - pane churn resets prior wedge escalation state before the stale-path poll
ok - a bare turn-end from a pane unchanged since the previous poll still surfaces
ok - a bare turn-end backed by a malformed prior hash surfaces
ok - a bare turn-end backed by a newline-terminated prior hash surfaces
ok - a churning secondmate turn-end surfaces without a stale resurface path
ok - a turn-end whose marker key matches another recorded endpoint surfaces
ok - two metadata records sharing one endpoint make churn evidence ambiguous
ok - a batch may satisfy positive evidence independently per task
ok - per-task evidence composition stays off until the home opts in
ok - a status-bearing batch never falls through to pane-churn evidence
ok - pane-churn turn-end absorb is off until a home opts in
ok - a perpetually churning pane surfaces once its bounded deferral window is spent
ok - a second pane-churn absorb keeps supervising instead of killing the watcher
ok - signal_turnend_panes_churned guards every possibly-empty array expansion (missing_keys, created_keys)
ok - an unrecordable pane-churn deadline surfaces the turn-end
ok - an invalid pane-churn bound surfaces the turn-end
ok - an oversized pane-churn bound surfaces the turn-end
ok - invalid existing pane-churn deadlines surface without mutation
ok - a surfaced batch opens no partial pane-churn deadline
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - a secondmate's status note surfaces even while its own agent is busy
ok - a self-announced close never wakes its own home, and the next real note still does
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a needs-decision signal row's queued payload is marked needs-decision: for branch exclusion
ok - a reconciliation-required needs-decision row's queued payload is still marked needs-decision:
ok - a captain-held signal stays actionable while the crew is still working
ok - a pending-reply second-mate escalation is marked for main-only routing
ok - an ordinary blocked event remains branch-eligible
ok - a routine event containing a needs-decision phrase keeps its ordinary payload, unmarked
ok - a captain event hidden behind a later routine append is still surfaced (queue + exit)
ok - a finished release reported before routine cleanup chatter is still surfaced
ok - a routine append after an already-classified event is absorbed (no re-wake)
ok - unreadable status reports are bounded without advancing classification
ok - permission recovery surfaces content from the unadvanced position
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
- Outcome: ⚠️ 1 warning across 1 run (12m10s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Step was skipped.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 warning
  • ⚠️ tests/fm-watch-triage.test.sh - The targeted test suite for the batch's stale-alarm bounding change (d4eb228, exercised by tests/fm-watch-triage.test.sh, +272 lines) did not finish running before this report had to be produced. It builds real fixture repos and appears to use wall-clock timers (multiple minutes elapsed with minimal CPU), similar to the other two suites which each took ~5 minutes. Live partial output (captured to evidence) showed dozens of 'ok' assertions covering stale-classification, wedge-escalation, and captain-hold absorb logic with zero failures observed up to that point, but the suite had not reached its FM_TEST_END/exit status. The other two targeted suites (tests/fm-test-run.test.sh for 6d396da + the d6190d3 fixture fix, and tests/fm-pi-branch-extension.test.sh for 0b9f518) both completed with full passes. Recommend re-running ./bin/fm-test-run.sh tests/fm-watch-triage.test.sh to completion and confirming its final exit=0 before treating batch 16 as fully validated.
  • ./bin/fm-test-run.sh tests/fm-test-run.test.sh --json /tmp/fm-test-run-result.json — 34/34 cases passed (exit=0), including 'a task marker refuses execution in the primary checkout and leaves worktrees and inspection alone' (the case d6190d3 fixed by using install_runner in the linked-worktree fixture)
  • ./bin/fm-test-run.sh tests/fm-pi-branch-extension.test.sh --json /tmp/fm-pi-branch-result.json — 42/42 cases passed (exit=0), including 'an extension-registered provider resolves in the isolated branch runtime' (the 0b9f5186 feature)
  • ./bin/fm-test-run.sh tests/fm-watch-triage.test.sh --json /tmp/fm-watch-triage-result.json — started but did not complete before this report; live partial transcript showed dozens of passing stale-alarm/wedge-escalation/captain-hold assertions with zero observed failures
  • git status --short — confirmed the worktree is clean, no transient test artifacts left behind
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

3264studios and others added 6 commits September 6, 2026 23:17
… is set (kunchenguid#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>
)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
…ranch (kunchenguid#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy
… bounded captain-hold stale alarms, Pi extension providers

Upstream range 5592cb6..0b9f518, three commits, brought in with one merge
commit so the waypoint stays in ancestry:
  6d396da fix(bin): refuse test runs in the primary checkout when a task marker is set (kunchenguid#3891)
  d4eb228 fix(bin): bound stale alarms for backlog captain holds (kunchenguid#3842)
  0b9f518 feat(pi): resolve extension-registered providers in the supervision branch (kunchenguid#3871)

Two files conflicted.

1. bin/fm-watch.sh, one hunk, resolved to UPSTREAM.

This is the same problem the fork's PR 18 (5e1e655, "stop false stale wakes on
healthy paused crew lanes") and upstream kunchenguid#3842 (d4eb228) both fixed with
different approaches: the fork anchored a one-shot on the status-file mtime
(paused_gate_needs_surface), upstream bounds the alarm on an open backlog
captain hold (task_captain_call_open / stale_wait_declaration /
captain_call_declaration / stale_wait_throttled / stale_wait_record /
captain_call_stale_bound). The captain's standing conflict rule is to take
upstream where both sides fixed the same thing differently, so upstream's block
is kept whole and paused_gate_needs_surface is deleted.

git auto-merged the fork's call site beside upstream's, leaving a half-spliced
hybrid: the fork function was gone from the conflict resolution but still had a
caller. Four fork-side remnants of PR 18's approach were therefore resolved to
upstream's shape as well, so the main loop is upstream's at every line:
  - the non-terminal-stale paused) branch: the fork's
    `if paused_gate_needs_surface ... surface_nonterminal_stale ... else
    handle_paused_stale` restored to upstream's plain `handle_paused_stale`.
  - pause_state_class's header comment: the fork's PR-18 expansion (which
    pointed at the deleted function) restored to upstream's four-line text.
  - the secondmate stale case: the fork's three-way
    `paused) / working) clear_pause_tracking / *) rm -f "$ssf" "$ewf"`
    restored to upstream's `paused) / *) clear_pause_tracking`.
  - the busy-pane and hash-change pause-bookkeeping clears: the fork's removal
    of `[ "$n" -ge 2 ] ||` and its added `afk_present ||` both restored. Both
    existed only to protect PR 18's declaration one-shot from being cleared.

After the resolution, bin/fm-watch.sh differs from upstream/main in 9 hunks,
each belonging to a fork watcher fix upstream never touched:
  - sourcing bin/fm-liveness-lib.sh; liveness_defers_wedge; the wedge_timer_check
    deferral and its comment; the fm_task_script_snapshot_* naming (PR 1,
    f9a59bb, declared-liveness wedge anchoring).
  - WATCHER_CLEANUP_LOCK_TICKS and watcher_cleanup's bounded lock wait (PR 7,
    2386360, deadlock-safe signal handling).
  - the Bash 3.2 guarded array expansions in signal_turnend_panes_churned;
    watcher_close_has_nothing_to_recover and the release-lock-quiet transition
    (PR 47, 9e69eed, silent watcher deaths and empty recovery wakes).
No surviving hunk belongs to PR 18. All 132 non-empty lines upstream added to
bin/fm-watch.sh in this range are present in the merged file.
grep -c 'the guard is progress, not depth' bin/fm-wake-lib.sh = 1 (PR 47 intact).

2. docs/scripts.md, one hunk, adjacency only.

Upstream reworded the fm-test-run.sh row for kunchenguid#3891 while the fork had added a
fm-test-env-lib.sh row (a fork-only script from PR #41) directly beneath it.
Resolved to upstream's row text plus the fork's additive row.

Test fallout.

tests/fm-watch-triage.test.sh auto-merged and carried both sides (98 base + 5
fork-only + 4 upstream-new = 107 cases). One fork case was deleted:

  - test_declared_pause_survives_benign_pane_repaint (from PR 18): its phase C
    asserts that a one-poll busy blip must not clear a parked lane's pause
    bookkeeping, which is exactly the behaviour upstream's `[ "$n" -ge 2 ]`
    clear restores. Phases A and B pass unchanged; only the dropped approach's
    assertion fails.

No other case was touched. The two remaining PR-18-era cases
(test_live_declared_pause_gate_surfaces_once_per_declaration,
test_live_pause_flag_absorbs_when_authoritative_state_falls_back) pass on
upstream's content-keyed throttle and were kept, as were PR 47's two cases. One
stale comment naming the deleted function was corrected to describe the
throttle that actually governs that fixture. Merged suite: 106 cases, 0
failures.

Silent-splice sweep: every non-empty line upstream added since 5592cb6 is
present in the merged tree for all 14 other touched files (fm-spawn.sh 17,
fm-test-run.sh 40, fm-teardown.sh 3, tests/lib.sh 5, fm-test-run.test.sh 57,
fm-kimi-harness.test.sh 2, fm-pi-branch-extension.test.sh 106,
fm-backend-herdr-presentation-e2e.test.sh 16, fm-branch-supervision.ts 54,
architecture.md 8, configuration.md 5, scripts.md 1, pi-supervision-branch.md 2,
verification/trace-context.md 1; 0 missing each). kunchenguid#3891's marker lands whole:
FM_TASK_ID is exported to ship and scout panes by bin/fm-spawn.sh and refused
against the primary checkout by bin/fm-test-run.sh. bin/fm-captain-hold.sh
auto-merged clean and carries upstream's new `open <task> --identity`
predicate; its only difference from upstream is the fork's additive
precondition-reminder feature. bin/fm-test-run.sh keeps the fork's family
registrations: --check-coverage reports total=201 with every tests/*.test.sh in
a family. .github/workflows/ is byte-identical to upstream.
…eckout fixture

Upstream's kunchenguid#3891 test `test_task_marker_refuses_the_primary_checkout` builds a
synthetic repo plus linked worktree and copies `bin/fm-test-run.sh` into each
with a bare `cp`. Upstream's runner starts from that alone; the fork's does not.

`bin/fm-test-env-lib.sh` is fork-only (PR #41, the owner of the fleet-home
overrides a test must never inherit) and `bin/fm-test-run.sh` sources it,
refusing with "missing bin/fm-test-env-lib.sh beside this runner; a fixture that
copies the runner must copy it too" when it is absent. So the fixture's linked
worktree could not start the runner at all, and the case failed on its second
phase ("the runner must still run in a linked task worktree") while its first
phase, the refusal this upstream commit adds, passed correctly.

This file already owns the fix: `install_runner()` was added for exactly this,
with the comment "A fixture that copies only the runner leaves it unable to
start." The new fixture now uses it instead of the bare `cp`. No runner
behaviour changes, and the case exercises upstream's contract as written.

Verified: reproduced the failure standalone against the merged tree (the linked
worktree exited 2 with the missing-library diagnostic and never wrote its `ran`
marker), then re-ran tests/fm-test-run.test.sh, which now passes end to end.
@zeeshaanahmad
zeeshaanahmad merged commit c442cd0 into main Sep 7, 2026
26 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants