Skip to content

Fork renewal 2026-09-14: 13 upstream commits - #26

Merged
sanis merged 20 commits into
mainfrom
chore/upstream-renewal-2026-09-13
Sep 14, 2026
Merged

sanis merged 20 commits into
mainfrom
chore/upstream-renewal-2026-09-13

Conversation

@sanis

@sanis sanis commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Fork renewal 2026-09-14

Continues this PR (opened 2026-09-13) rather than stacking a new one: its base
(origin/main) had not moved since it branched, so per the standing renewal
policy this session updated the existing branch instead of opening a second PR.

Drift measured at session start (before this update):
HEAD..upstream/main = 632 commits, upstream/main..HEAD = 1 commit.
That large number is origin/main's own drift, not this branch's: origin/main
is stale relative to this already-in-flight renewal, which had previously
merged upstream through 10 commits.
Compared directly against this branch's own prior head, the real remaining
drift was 3 new upstream commits (HEAD..upstream/main = 0 after this
session's merge).

Branch: chore/upstream-renewal-2026-09-13 (unchanged)

Merge

git merge upstream/main (commit 3c00b0d) pulled in the 3 new upstream
commits since the branch's prior state:

Conflicting files: none. Git's ort merge strategy auto-merged every
touched file cleanly (CONTRIBUTING.md, bin/fm-afk-return.sh,
docs/configuration.md, tests/fm-afk-return.test.sh) with no conflict
markers, so no manual resolution judgment was required this round.

Fork-only behavior verified to survive:

  • The fork's dedicated documentation-audiences case-pattern family in
    bin/fm-test-run.sh (established in a previous renewal round, ac8ffe6)
    remains the sole owner of fm-documentation-audiences.test.sh — confirmed
    no duplicate entry reappeared in the shared pure-contract-unit family this
    round (the known trap this task warned about).
  • docs/fm-test-portable-shards.md already documents the "derived from a live
    command" model — bin/fm-test-run.sh --check-coverage is the current
    account of lane size, shard composition, and balance rather than a copied
    table — so there was no static table to recompute; confirmed that model is
    still intact post-merge.

Local verification (before pushing through the gate)

  • bin/fm-lint.sh: clean (ShellCheck 0.11.0 pinned, actionlint 1.7.12 pinned, 3 workflow files valid)
  • CI=true bin/fm-lint.sh: clean (full ShellCheck extended analysis enabled)
  • bin/fm-test-run.sh --check-coverage: FM_TEST_COVERAGE ok total=209 parallel=24 parallel_max_ms=417163 parallel_imbalance_ms=2894 parallel_unhinted=0 serial=169 serial_shards=5 serial_unhinted=17 herdr=16
  • bin/fm-doc-audience-check.sh: ok surfaces=105 local_links=414
  • Targeted tests for every file this merge touched (tests/fm-afk-return.test.sh, tests/fm-claude-trust.test.sh, tests/fm-procevent-when.test.sh, tests/fm-update.test.sh): all passed, FM_TEST_SUMMARY total=4 failed=0

no-mistakes gate (run 01M2F13DMYRBVTW63YKM0CS1CN)

  • intent: skipped
  • rebase: completed
  • review: completed after 2 fix rounds. 7 findings surfaced (2 auto-fix, 2 ask-user, 3 no-op/info):
    • Applied (auto-fix): bin/fm-claude-trust.sh — a worktree's own recorded
      declined-external-imports flag could be silently overridden by the primary
      project's carried-forward consent; now also checked against the target
      (worktree) entry itself.
    • Applied (auto-fix, corrected mid-review): bin/fm-update.sh — a failed
      fm-procevent-when.sh rebind-all was silently swallowed (|| true). The
      first fix round made the whole script exit non-zero on that failure, but
      the pipeline's own re-review caught that this would break an existing,
      unmodified caller — bin/fm-remote-secondmate-control.sh's cmd_update
      treats any non-zero exit as a fatal "remote code root update failed" and
      skips the following sync, even though the actual git update had already
      landed. The fix was narrowed to surface a clear warning: line on stderr
      without changing the script's exit status or existing callers' contract.
    • Accepted as-is (ask-user, decided by judgment — no live user available
      in this scheduled/autonomous run): claude-trust-project-entry-widening —
      every claude worktree-mode spawn now also pre-trusts the primary
      checkout's own ~/.claude.json entry. The finding itself notes this is
      already documented and tested as intentional upstream behavior, not
      something this merge's resolution introduced, so it was accepted rather
      than second-guessing an upstream product/security decision during a
      routine sync.
    • Accepted as-is (ask-user, decided by judgment): afk-return-acked-gap-suppression
      — a plausible corner case where a second real watcher outage during an
      already-acked away window could go unreported. This is upstream's own
      newly-merged logic, not introduced by this merge. Patching it would mean
      unilaterally diverging from upstream in safety-relevant supervision code,
      outside a routine renewal's scope. Flagging for review: worth deciding
      whether to report this upstream or accept the risk as-is; not judged
      clearly safe enough to silently patch in a fork renewal.
    • No action (no-op/informational): nudge-pid1-numeric-match,
      pr-merge-github-no-defense-in-depth-auto-guard,
      pr-merge-authority-persist-failure-exits-before-outcome-report.
  • test: completed
  • document: completed (added header documentation and new test coverage for both review fixes)
  • lint: completed
  • push: completed — the gate pushed 550227f0 to origin/chore/upstream-renewal-2026-09-13
  • pr / ci: skipped — gh CLI is not installed in this environment (run.automatic_skips)
  • Outcome: passed-with-skips — skipped only for the gh-CLI reason above; not itself a code failure or proof of CI readiness

What I could not do

  • Could not open or update this PR through no-mistakes's own pr step, or
    check CI through its ci step, because gh is not installed in this
    session — updated this PR by hand via the GitHub API instead, per the task
    instructions for this environment.
  • Did not wait for GitHub-side CI to finish before writing this report; please
    confirm all checks are green on this PR before merging.
  • The two ask-user findings above were decided by my own judgment rather
    than a human's, since no user was available to consult in this scheduled
    run — please double-check those two if you'd like a second opinion,
    especially the afk-return-acked-gap-suppression one.

Generated by Claude Code

kunchenguid and others added 10 commits September 12, 2026 00:02
* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.
…henguid#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
Upstream brings the Antigravity (agy) worker/scout adapter, synchronous
merge authority under the away posture, bounded per-item backlog reads,
superseded-check-run handling, and secondmate Claude trust pre-registration.

Conflicts resolved:

- bin/fm-test-run.sh: union of both sides' lane classification. The fork
  keeps its dedicated documentation-audiences shard, so upstream's listing
  of fm-documentation-audiences.test.sh under pure-contract-unit is dropped
  (a leading case arm would shadow the fork's own arm). The fork's
  fm-download-lib.test.sh entry stays and upstream's new
  fm-agy-harness.test.sh is added.

- docs/fm-test-portable-shards.md: upstream removed the copied per-shard
  table and the derived hint counts in favour of pointing at
  `bin/fm-test-run.sh --check-coverage` as the live account. That newer
  model is taken wholesale rather than recomputing a table upstream no
  longer maintains; the fork's paragraph recording the provenance of its
  four fork-only test measurements is re-applied on top.

Verified: bin/fm-test-run.sh --check-coverage passes
(total=208 parallel=24 serial=168 serial_shards=5 serial_unhinted=16
herdr=16); changed-file lint clean on pinned ShellCheck 0.11.0 and
actionlint 1.7.12.
…uid#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes kunchenguid#2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 6d720020-d832-415c-9ea4-2192879cc008


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

sanis commented Sep 13, 2026

Copy link
Copy Markdown
Owner Author

Blocker: PR must be raised via no-mistakes — cannot be fixed from this session

Failing check: PR must be raised via no-mistakes (job), on head d9ab5b37.

This is a real failure, not an infra flake, and it is not something a push from here can clear.

.github/workflows/no-mistakes-required.yml accepts exactly two proofs, and this branch carries neither:

  1. A no-mistakes signature in the PR body (Updates from [git push no-mistakes]) with a valid step attestation bound to the current head. The body has no signature.
  2. Any commit in the PR whose subject matches ^no-mistakes\([a-z-]+\): . The PR's nine commits are the upstream renewal merge plus the eight upstream commits it brings; none carries that subject prefix.

Both proofs are, in the workflow's own words, "deterministic pipeline output" — they exist only because git push no-mistakes actually ran the review, test, and document steps against this head. That pipeline runs on the maintainer's machine and is not reachable from this cloud session. Fabricating either proof would assert that those steps ran when they did not, so I have not done it, and I have not re-run the job: the result is deterministic and a re-run would return the same failure.

What unblocks it — either raise this branch through git push no-mistakes locally so the attestation is written against d9ab5b37, or land the branch by another route of your choosing.

Context on the rest of the run

This branch is the 2026-09-13 fork renewal: 8 upstream commits merged through 35761484, against 245 fork-only commits. Two conflicts, both in test-lane bookkeeping rather than product code, and every fork-only behavior was verified present in code and prose afterwards.

Already green on this run: Test coverage guard and Repo invariants. The remaining jobs (five serial shards, two parallel shards, lint, Herdr, macOS Bash) were still in flight when this was written, and matter here: 14 scripts could not be exercised honestly in the cloud container that produced the merge (it runs as root, under a POSIX locale, on a slow runner), and all 14 were shown to fail identically on pristine upstream/main under those same conditions. This CI run is the first honest verdict on them. I will report on any that fail here.


Generated by Claude Code

sanis commented Sep 13, 2026

Copy link
Copy Markdown
Owner Author

CI has finished on d9ab5b37. 13 of 14 checks green; the only failure is the no-mistakes attestation gate covered in my previous comment, which is unchanged and still needs your local pipeline.

Green: Test coverage guard · Repo invariants · Lint · Behavior portable parallel 1 and 2 · Behavior portable serial 1, 2, 3, 4 and 5 · Behavior tests (Herdr) · Stock macOS Bash snapshot compatibility · Behavior timing aggregate.

That settles the open question from my previous comment. The 14 scripts the cloud container could not exercise — it runs as root, under a POSIX locale, on a slow runner — all pass on real CI. Nothing in this merge was hiding behind that environment, and no test was skipped, weakened, or deleted to get here.

No merge conflict: base is still f0ffdce. The PR is mergeable on its content and blocked only on the attestation.


Generated by Claude Code

NewAiCoder and others added 3 commits September 12, 2026 23:26
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…wal-2026-09-13

# Conflicts:
#	bin/fm-test-run.sh
…ttern

The merge conflict resolution had placed it in pure-contract-unit
alongside upstream's other new additions, but the fork already owns a
dedicated documentation-audiences family for it further down, making
the pure-contract-unit copy dead (shellcheck SC2221/SC2222).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013u2pvUP884FjBKE33jQiPA
@claude claude Bot changed the title Add agy (Antigravity CLI) crewmate/scout adapter support Fork renewal 2026-09-13: 10 upstream commits Sep 13, 2026
@claude
claude Bot marked this pull request as draft September 13, 2026 07:58

sanis commented Sep 13, 2026

Copy link
Copy Markdown
Owner Author

Status update — the head moved, so my earlier report (which named d9ab5b37) is superseded.

Current head ac8ffe6d. Since that report, the branch gained a second upstream merge bringing two further commits (feat(afk): add quiet supervision mode for a present captain, fix(bin): let verified harness ancestry outrank retained markers) plus fix(bin): remove duplicate fm-documentation-audiences.test.sh case pattern.

I verified that last one, because it touches the conflict this renewal resolved by hand: it removes upstream's copy of fm-documentation-audiences.test.sh from the pure-contract-unit arm and keeps the fork's dedicated documentation-audiences shard. That is the correct resolution — the second merge had reintroduced the duplicate. The fork's shard is intact, and the fork-only material I checked before is still present after the second merge (ClickUp routing, the --no-description-check merge carve-out, the Herdr in-pane-daemon findings, the full-PR-URL rule).

CI on ac8ffe6d: all 12 behavior, lint, invariant and coverage checks green. Coverage guard now reports 209 scripts (up from 208), and it passes locally on this head too.

The only failure remains PR must be raised via no-mistakes, unchanged and for the reason already given above — it has now run twice on this head, both deterministic. Nothing new to add there.

One note for whoever lands this: commit ac8ffe6d carries a Co-Authored-By: Claude Sonnet 5 trailer, which AGENTS.md prohibits ("Never add an agent name as a commit co-author"). I have not rewritten it, since the branch has more than one writer.


Generated by Claude Code

NewAiCoder-bot and others added 7 commits September 13, 2026 03:20
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.
…kunchenguid#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
@claude claude Bot changed the title Fork renewal 2026-09-13: 10 upstream commits Fork renewal 2026-09-14: 13 upstream commits Sep 14, 2026
@claude
claude Bot marked this pull request as ready for review September 14, 2026 04:53
@claude

claude Bot commented Sep 14, 2026

Copy link
Copy Markdown

Final status: all 14 checks green, including PR must be raised via no-mistakes, which now carries a real pipeline run (no-mistakes(review):, no-mistakes(document): commits on head 550227f0). No merge conflict, no open review threads, base still f0ffdce.

For context on the earlier comments: I ran this fork-renewal branch from a container that runs as root, where no-mistakes'''s review step could not launch its claude reviewer (claude itself refuses --dangerously-skip-permissions under root) and my own session could not spawn a nested agent to work around that. I documented that blocker in the PR body and in my earlier comments rather than fabricate the gate'''s proof. It looks like that gap has since been closed from an environment that could actually run the pipeline.

Marking this ready for review since it is green and mergeable.


Generated by Claude Code

@sanis
sanis merged commit 599ea9e into main Sep 14, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants