Skip to content

Merge upstream/main into fork main (8-commit parity sync) - #38

Merged
adibirzu merged 9 commits into
mainfrom
fm/upstream-sync-r7
Sep 8, 2026
Merged

adibirzu merged 9 commits into
mainfrom
fm/upstream-sync-r7

Conversation

@adibirzu

@adibirzu adibirzu commented Sep 8, 2026 •

Copy link
Copy Markdown
Owner

What this is

Brings adibirzu/firstmate current with kunchenguid/firstmate at 891dc517.
Scope is deliberately narrow: the audited 8-commit parity merge and nothing else.
No fork features are bundled here.

The branches share merge-base 6d396da7, so this is an ordinary merge — no rebase,
no cherry-pick, no force, nothing discarded. After the merge,
git rev-list --count upstream/main ^HEAD is 0 and all 8 upstream commits are ancestors.

Commit Subject
891dc517 derive watcher beacon staleness grace from poll cadence (kunchenguid#3946)
72bfdd00 carry attribution-off policy in every claude launch (kunchenguid#3945)
98b37d40 use system stat for Darwin BSD formats (kunchenguid#3305)
3af74fe8 make every counted wake queue row presentable or retired (kunchenguid#3950)
36fd955b harden the export-DOM render step, record Pi 0.85.1 evidence (kunchenguid#3952)
ffd2c899 gate secondmate wake-loop stall alerts on real queue no-progress (kunchenguid#3943)
0b9f5186 resolve extension-registered providers in the supervision branch (kunchenguid#3871)
d4eb2280 bound stale alarms for backlog captain holds (kunchenguid#3842)

Conflicts and how each was resolved

Five files conflicted, exactly as the pre-merge audit predicted. Each was resolved on its
own merits — no blanket "ours" or "theirs".

bin/fm-wake-lib.sh — union. The only conflicting production file. Both sides appended
an independent function at the same insertion point: the fork's
fm_lock_resolve_path/fm_lock_paths_equal (symlink-divergent lock identity) and upstream's
fm_poll_derived_grace (poll-cadence guard grace). Different names, different conditions, no
competing behavior for the same condition, and both have live callers — so both are kept.
Verified fm_poll_derived_grace still computes max(300, poll+60) after the merge.

tests/fm-watch-triage.test.sh — kept the fork's watch_marker_key helper. This is the
d4eb2280 cherry-pick artifact the audit flagged. Upstream's inline tr ':/.' '___' is the
v1 marker-key format; the fork has since moved production to the v2 v2-<hex> format in
bin/fm-marker-lib.sh. The two are not equivalent — they produce
test_fm-held-merge vs v2-746573743a666d2d68656c642d6d65726765. Taking upstream's line would
have made the fixture look for a marker file the merged watcher never writes.

tests/fm-calm-pi-extension.test.sh — took upstream. 36fd955b extracts the inline Chrome
--dump-dom block into render_export_dom with bounded retries and a diagnostic report; the
fork's side is the superseded inline form. The local declaration hunk follows the same body:
it drops the four inline-only Chrome variables and adds chrome_report.

docs/configuration.md — took upstream. ffd2c899 changes
FM_SECONDMATE_WAKE_STALL_SECS from a 60s row-age floor to a 180s no-progress interval, and
the auto-merged bin/fm-watch.sh already enforces 180. Keeping the fork's line would have
documented a default the code no longer has.

docs/fm-test-portable-shards.md — fork side, then regenerated. Both sides carry derived
figures that the merge invalidates, so neither could be right as written. Kept the fork's newer
measurement pass, then recomputed against the merged runner: the serial lane is now 183
scripts over 175 hints with 8 unhinted
, shards 36/37/37/36/37 at ~16.81 min, 17 ms imbalance.
bin/fm-test-run.sh --check-coverage passes on the merged tree
(total=220 parallel=24 serial=183 serial_shards=5 serial_unhinted=8 herdr=13).

Also verified: no d4eb2280 content is applied twice, and every line the merged tree removes
relative to upstream/main in the watcher files is a fork refactor with a live replacement, not
a lost upstream change.

Pre-existing failures — please do not attribute these to this sync

Baselined before merging, from CI on the pre-merge commit 833c9001 (run
34185010583):

Lane Status on main before this PR Failing assertion
Behavior portable parallel 2 already red an exact non-active projection close should succeed after restoring focus
Behavior tests (Herdr) already red projected teardown changed active workspace/tab from w3/w3:t1 to w8/w8:t2

Both are Herdr presentation-space projection teardown failures and are unrelated to the
watcher/wake machinery this sync touches.

Confirmed after this PR's CI run (34187972666):
both lanes fail on the byte-identical assertions as the baseline, same families, same failed=1.

parallel 2  not ok - an exact non-active projection close should succeed after restoring focus:
                     warning: herdr presentation pane close did not restore the exact prior workspace and tab
            FM_TEST_SUMMARY_FAMILY family=backend-dispatch count=3 failed=1
Herdr       not ok - projected teardown changed active workspace/tab from w3/w3:t1 to w8/w8:t2
            FM_TEST_SUMMARY_FAMILY family=real-herdr-gated count=13 failed=1

A third red: Behavior portable serial 4 was killed on the job timeout

It reports as cancelled, which is how GitHub renders a timeout-minutes kill — there is no
failing assertion. Its 36 scripts all passed (FM_TEST_SUMMARY total=36 failed=0); the job
simply ran out of wall clock at exactly 20:00.

This lane was already at the cliff edge on main before this merge:

Run Shard-4 step time Scripts Test failures Outcome
baseline 833c9001 19.67 min 37 0 passed, by ~20 s
this PR 19.75 min 36 0 killed at the 20 min cap

Five seconds of runner variance separates them, and this PR's shard 4 carries one fewer
script than the baseline's. So this is a pre-existing capacity cliff the merge made more likely
to trip, not a defect introduced by it — exactly the risk
docs/fm-test-portable-shards.md already names ("the lane grows by scripts rather than by
minutes per script, so the next few additions are the ones to watch for a shard-count
increase").

The fix belongs in its own PR, because it is a coordinated two-file change outside this
sync's audited scope: bin/fm-test-run.sh owns the shard count n and .github/workflows/ci.yml
derives the same n from strategy.job-total, so 5 → 6 has to move in both at once or the lane
fails loudly. Raising timeout-minutes instead would weaken a deliberate hang tripwire.

Local test results

Run on macOS (arm64), which matters for two of them.

  • tests/fm-claude-stop-autoarm.test.sh — pass, including upstream's new
    a long FM_POLL with FM_GUARD_GRACE unset reaches fm-watch-arm.sh with the derived grace,
    which is the merge's own proof that fm_poll_derived_grace is wired correctly.

  • bin/fm-lint.sh — pass (ShellCheck 0.11.0, actionlint 1.7.12, 3 workflows valid).

  • bin/fm-test-run.sh --check-coverage — pass.

  • tests/fm-watch-triage.test.sh (the union arbiter, full file) — 101 assertions, 0 failures,
    exit=0, three runs out of three.
    This is the suite that proves the bin/fm-wake-lib.sh
    union and the v2 marker-key resolution are behaviorally correct and not merely textually
    clean. The captain-held churn section is a race, so it was repeated deliberately — a single
    green run would have proved little (the PR feat(fleet): add per-crew usage rows #36 work measured 3/4 and 3/3 here).

    Run Assertions Failures Exit Duration
    1 101 0 0 1210961 ms
    2 101 0 0 1209570 ms
    3 101 0 0 1205028 ms
  • tests/fm-wake-queue.test.sh — 1 failure,
    a foreign queue with no progress did not alert.

  • tests/fm-turnend-guard.test.sh — 1 failure,
    a beacon older than the poll-derived grace must still block: expected exit 2, got 0.

Those last two are not caused by this merge. They reproduce identically on a pristine
upstream/main tree with no fork code present:

$ git archive upstream/main | tar -x -C /tmp/up && cd /tmp/up
$ bin/fm-test-run.sh tests/fm-wake-queue.test.sh tests/fm-turnend-guard.test.sh
not ok - a foreign queue with no progress did not alert: checkpoint: no actionable wake within 1s
not ok - a beacon older than the poll-derived grace must still block: expected exit 2, got 0
FM_TEST_SUMMARY total=2 failed=2

Root cause for the turn-end one is a macOS portability gap in upstream's own new test:
891dc517 added three touch -d "@$beat" calls to tests/fm-turnend-guard.test.sh, and BSD
touch rejects -d @epoch (touch: out of range or illegal time specification). The beacon
keeps a fresh mtime, so the guard correctly declines to block and the test sees 0 instead of
2. It is the same Darwin class 98b37d40 fixed for stat, just missed for touch. Linux CI
is unaffected, which is why these lanes are green upstream. Left unfixed here on purpose to
keep this PR to the audited parity merge; worth a small follow-up.

Not done here

Not merged — leaving that to the captain.

mremond and others added 9 commits September 7, 2026 03:26
)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in kunchenguid#3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.
…ranch (kunchenguid#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy
…gress (kunchenguid#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
…idence (kunchenguid#3952)

* fix(tests): make the Calm export-DOM render step retry and report

The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.

What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.

Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.

The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo

* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change

* no-mistakes(review): tests: drop inert chrome flags, report render timeouts

* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…unchenguid#3950)

* fix(bin): make every counted wake queue row presentable or retired

A wake row could be counted as queued while no drain would ever present
it, leaving the operator told to "drain them before anything else" by a
command that printed nothing and offered no acknowledgement.

Two independent paths produced that state.
A row reserved by a live supervision-branch grant is excluded from a main
drain by design, but fm-guard.sh counted the whole queue, so main was
warned about rows only the branch could present, on every guarded command
for as long as the grant was held.
A row that lost its five appended fields or its numeric sequence can never
be claimed, presented, or named by an --ack-through cutoff, yet it still
counted as queued, wedging the queue permanently.

The guard now counts only the rows the calling actor can itself present or
retire, and a main drain retires unusable rows under the queue lock,
reporting them in bounded escaped form before removal so the evidence
survives for the separate row-generating defects. A retirement failure is
reported loudly and never suppresses unrelated consumable work. A main
drain whose remaining rows are all branch-held says so in one bounded line
instead of exiting silently. Grant row-list and owner-record reads move
into fm-wake-lib.sh so the drain, the grant publisher, and the guard share
one implementation.

Ownership is unchanged: a branch drain still touches nothing outside its
grant and never retires a row, and main still cannot present or acknowledge
an active grant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1

* no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass

* no-mistakes(review): add guard advisory for branch-held wake rows

* no-mistakes(document): document per-actor wake counting and unusable-row retirement

* no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing

* fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main

* test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state

* ci: re-trigger after Chrome headless timeout in calm HTML export test

* ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test

* test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat

* no-mistakes(document): Refresh stat and shard docs
…henguid#3945)

The captain's attribution policy (no Co-Authored-By trailer, no
Claude-Session link, no generated-with line) lives in Claude Code's `user`
settings scope. A spawned worker's settings sources are not guaranteed to
load that scope, so a launched worker could write attribution trailers into
its commits and PR bodies regardless of the captain's own configuration.

launch_template()'s claude case now carries the same policy
("attribution": {"commit": "", "pr": "", "sessionUrl": false}) directly in
its inline --settings JSON, so every claude launch keeps attribution off
independent of which settings scopes end up loaded. Tests assert the policy
on the rendered launch command for both a crewmate and a secondmate spawn.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nchenguid#3946)

* fix(bin): derive the away-mode beacon grace from the poll cadence

fm-turnend-guard.sh's away-mode branch required the watcher beacon to be
fresh within the flat FM_GUARD_GRACE default (300s), but the daemon starts a
fresh one-shot watcher only after it finishes handling the previous wake, and
that handling can legitimately outrun a fixed 300s window under load (a slow
registered check, a busy supervisor pane) with the daemon perfectly healthy
throughout. That misread a live, correctly-cycling daemon as down and blocked
the turn.

Add fm_poll_derived_grace, the single owner of the max(300, FM_POLL + 60)
formula, and have the away-mode branch, fm-claude-stop-autoarm.sh, and
fm-watch.sh's own runtime beacon-staleness check all derive their default
grace from it instead of the flat default. A dead daemon pid or a beacon
older than that grace still blocks, so a genuinely lapsed away mode still
alarms; every other check is unchanged.

fm-claude-stop-autoarm.sh computed the derived grace into GRACE but its two
fm-watch-arm.sh invocations called the wrapper bare, so the wrapper fell back
to its own flat 300s default and could reject a healthy long-poll watcher.
Both invocations now pass FM_GUARD_GRACE="$GRACE" through explicitly, and a
new test proves a long FM_POLL with FM_GUARD_GRACE unset reaches
fm-watch-arm.sh with the derived value.

Also drops fm_last_activity_age, added alongside the derivation but never
called anywhere in the tree; fm-inactive-reconcile.sh already owns that
computation.

* no-mistakes(review): Remove dead WATCHER_STALE_GRACE assignment in fm-watch.sh

* no-mistakes(document): Update FM_WATCHER_STALE_GRACE default note for poll-derived grace

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Brings adibirzu/firstmate current with kunchenguid/firstmate at 891dc51.
The two branches share merge-base 6d396da, so this is an ordinary merge,
not a replay: no rebase, no cherry-pick, no force, nothing discarded.

Upstream commits absorbed (8):
  891dc51  derive watcher beacon staleness grace from poll cadence (kunchenguid#3946)
  72bfdd0  carry attribution-off policy in every claude launch (kunchenguid#3945)
  98b37d4  use system stat for Darwin BSD formats (kunchenguid#3305)
  3af74fe  make every counted wake queue row presentable or retired (kunchenguid#3950)
  36fd955  harden the export-DOM render step, record Pi 0.85.1 evidence (kunchenguid#3952)
  ffd2c89  gate secondmate wake-loop stall alerts on real queue no-progress (kunchenguid#3943)
  0b9f518  resolve extension-registered providers in the supervision branch (kunchenguid#3871)
  d4eb228  bound stale alarms for backlog captain holds (kunchenguid#3842)

Five files conflicted; each was resolved on its own merits rather than by a
blanket side preference.

bin/fm-wake-lib.sh - union. The only conflicting production file. Both sides
appended an independent function at the same insertion point: the fork's
fm_lock_resolve_path/fm_lock_paths_equal (symlink-divergent lock identity) and
upstream's fm_poll_derived_grace (poll-cadence guard grace). Different names,
different conditions, no contested behavior, and both have live callers, so
both are kept.

tests/fm-watch-triage.test.sh - kept the fork's watch_marker_key helper. This
is the d4eb228 cherry-pick artifact: upstream's inline `tr ':/.' '___'` is the
v1 marker-key format, and the fork has since moved production to the v2
`v2-<hex>` format in bin/fm-marker-lib.sh. Taking upstream's line would have
made the fixture compute a key the merged watcher never writes.

tests/fm-calm-pi-extension.test.sh - took upstream. 36fd955 extracts the
inline Chrome dump-dom block into render_export_dom with bounded retries and a
diagnostic report; the fork's side is the superseded inline form. The local
declaration hunk follows the same body, dropping the four inline-only Chrome
variables and adding chrome_report.

docs/configuration.md - took upstream. ffd2c89 changes
FM_SECONDMATE_WAKE_STALL_SECS from a 60s row-age floor to a 180s no-progress
interval, and the auto-merged bin/fm-watch.sh already enforces 180, so the
fork's line would have documented a default the code no longer has.

docs/fm-test-portable-shards.md - kept the fork's newer measurement pass, then
regenerated the derived figures against the merged runner, which neither side
could have had right: the serial lane is now 183 scripts over 175 hints with 8
unhinted, shards 36/37/37/36/37 at ~16.81 min and 17 ms imbalance.
bin/fm-test-run.sh --check-coverage passes on the merged tree.

Verified: no content from d4eb228 is applied twice, and no upstream
production change was lost to an auto-merge.
@adibirzu
adibirzu merged commit 5e5d8b8 into main Sep 8, 2026
10 of 16 checks passed
@adibirzu
adibirzu deleted the fm/upstream-sync-r7 branch September 10, 2026 12:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants